Is it reasonable to take over production before disaster recovery has been proven?

0
0
Asked By VelvetMango42 On

We're taking over a production Java application from a SaaS provider, whose service expires at the end of the month. The application is already running in our own environment and needs to operate indefinitely after the transition.

The problem is that we haven't completed a production database restore from backup, and our long-term backup and recovery process is not fully established. Organizational constraints currently allow deployments only to the test environment, not production.

We also depend on another team for DevOps and infrastructure work. They're capable, but most of the systems they support use Python or Node, so the Java deployment and operational requirements are relatively new for them. The transition is being handled with a fairly rigid, deadline-driven approach.

At this point, we can demonstrate that the application runs, but we have not proven that we can recover it after a serious failure. Until we can restore the database, clean or validate the recovered data, and bring the application back into working condition, I don't feel we have a real recovery plan.

Is this kind of migration common? What would you consider the absolute minimum to complete, test, document, and formally accept before the SaaS environment is shut down? It feels like we're building the production system and its safety net simultaneously, under an immovable deadline.

3 Answers

Answered By PaperOrbit_6 On

First establish who is accountable for keeping the application running in production and who has authority to accept the remaining risks. If that is you, management needs to explicitly understand that you are being asked to take responsibility without having production deployment access or a validated recovery path.

The application working today is not evidence that the transition is safe. A written go/no-go checklist, an owner for each operational task, tested rollback and restore procedures, monitoring, and a support escalation path would make the risk visible and actionable.

VelvetMango42 -

That’s the difficult part: I’m effectively responsible for the application, and I was hired because the organization lacked Java and cloud-architecture experience. We’ve fixed many of the provider’s problems and everything currently works, but the deadline is forcing us to finish operational readiness at the same time as the migration.

NorthStarPine3 -

The fact that the work is being pushed through despite the deadline and unresolved recovery testing is itself a management risk. Make sure the decision-makers understand the gap and formally accept it if they choose to proceed.

Answered By CedarFox_81 On

This needs to be treated as a management and project-risk issue, not something you should quietly absorb as an individual contributor. Put the concerns in writing to your manager and the people accountable for the transition. Be specific about what has and hasn’t been tested, the consequences of failure, and what decision or resources you need.

Answered By QuietHarbor7 On

At an absolute minimum, perform and document a backup restore before the old service disappears. Verify that the recovered database is usable, that the application can connect to it, and that the resulting system supports the critical business workflows. You also need enough retained backups to recover from accidental deletion, corruption, ransomware, or a bad deployment—not just a backup job that reports success.

A full disaster-recovery exercise may be appropriate for a business-critical system, but the immediate non-negotiable is proving that your backups can actually produce a functioning service. Record the steps, timings, dependencies, credentials or access requirements, and who is responsible for each part.

Related Questions

LEAVE A REPLY

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.