Evaluating Flexday AI
Reliability and business continuity
What the infrastructure code and the platform define for backups, redundancy, durable work, scaling, health checks and releases, and what to ask Flexday for.
- Technical
Last reviewed
This page states only what the platform's code and its AWS infrastructure definitions set: the backups, the redundancy, how long-running work survives a failure, how the services scale, how health is checked and how a release rolls out and back. Flexday SaaS runs on that AWS shape (see Deployment options and data residency). Commitments about the hosted service, such as its availability, its recovery objectives and the results of its recovery tests, are facts about Flexday as a company, so they are listed at the end as things to request.
At a glance
- Backups are managed. The database keeps automated backups for 7 days, or 30 in production, and takes a final snapshot before it is ever deleted.
- Builds and Agent turns survive restarts. Builds, Flows, document indexing and Agent turns run as durable jobs; a build or an Agent turn interrupted by a crash resumes from its record of finished steps.
- Outside effects are deduplicated. Emails, ticket writes and other outside changes are claimed with an idempotency key before they run, so replaying a step that completed does not repeat them.
- Releases check themselves. Database changes run first under a lock, and a service whose new copy does not become healthy is rolled back to the previous one automatically.
- Redundancy is a setting. As committed, each environment runs one database instance and one cache node; adding a second of each is configuration, not a redesign.
Backups and recovery points
| Store | What the infrastructure defines |
|---|---|
| Database (Aurora PostgreSQL) | Automated backups kept 7 days in Flexday's development, test and internal environments and 30 days in production. A final snapshot is taken if the database cluster is deleted, and production clusters carry deletion protection. A cluster can be created from a snapshot. |
| Object storage (S3: apps, files, documents, export bundles) | Versioning on, so an overwritten or deleted object keeps its previous version for 30 days; incomplete uploads are cleaned up after 7 days. No copy to another region is defined. |
| Shared file system (EFS: older app drafts and sample data, and temporary upload space) | AWS automatic backups switched on. |
| Cache and job queues (ElastiCache for Redis) | A daily snapshot, kept for one day. |
| Inside the platform | Every applied version of a resource is kept as an immutable snapshot, and a File Store keeps earlier versions of each file until its own retention settings remove them. See Data lifecycle and portability. |
Redundancy
| Component | As configured in every environment | What can be changed |
|---|---|---|
| Network | Private subnets in two availability zones; outbound traffic leaves through one NAT gateway, or one per zone in production | A NAT gateway per zone is one setting |
| Database | One Aurora Serverless instance (a writer with no standby) | A second instance, which Aurora can fail over to, is a count in the configuration |
| Cache and queues | One Redis node | Multi-zone with automatic failover is one setting. Without it, replacing the node loses jobs still waiting in the queue; the platform's sweep marks such jobs failed rather than leaving them waiting for ever. |
| Services | Every service except the Studio API and the documentation site runs as one to three copies; the Studio API runs as one copy, and the documentation site, a separate application, as one or two | The scaling range per service |
Durable work
Builds, Flow runs, document indexing and Agent turns never run inside a web request. They run as durable jobs on the worker:
- One record, two places. A job sits in a Redis queue and has a durable record in PostgreSQL, so a closed browser can reconnect and replay its progress.
- Heartbeats and a sweep. A running job reports a heartbeat every 15 seconds. A sweep that runs on a schedule fails a job whose heartbeat has been silent for 90 seconds, so nothing waits for ever.
- Resume, not restart. If the worker dies part way through a Builder turn or an Agent turn, the job is delivered again and resumes from its ledger of tool calls already completed.
- At least once, with idempotency. Flow steps can run more than once after a failure, so every step with an outside effect (sending an email or a Teams message, a write over HTTP, sending or fetching a file, writing to a business application) claims an idempotency key first. A replay returns the recorded result instead of repeating the effect. If the process stops after an effect but before its result is recorded, a replay within fifteen minutes is held as "in doubt" rather than sent again. An effect that fails with an error releases its key so the step's retry can try again; for email, where a lost acknowledgement looks like a failure, that can rarely mean a message arrives twice.
- One live job per resource. The database refuses a second live job of the same kind on the same resource, so two clicks cannot start the same work twice.
See Background jobs and resilience.
Scaling
- Services. The gateway, the worker and the web applications scale between one and three copies on CPU and memory, each with a 70% target (the documentation site between one and two).
- Database. Aurora Serverless capacity moves between a minimum and a maximum set per environment.
- Work. Jobs wait in the queue until a worker is free, so a burst of builds or Flow runs queues rather than failing.
Health and readiness
| Check | What it does |
|---|---|
| Liveness | The API's health endpoint answers as long as the process runs, with no dependencies. |
| Readiness | The API's readiness endpoint answers only when the service is not shutting down, the database is reachable, every database migration is applied and, where sign-in is enforced on a deployment that serves more than one workspace, the database connection does not bypass row-level security; the load balancer sends traffic only to copies that pass. |
| Load balancer health | A copy is taken out of service after five failed checks ten seconds apart. |
| Alarms | CloudWatch alarms on the database, each service, the load balancer, the cache, the content delivery network and email, sent to critical and warning notification topics. |
Releases and rollback
- Database changes first. When the API image changed, a one-off migrate task applies the database changes, holding a lock so only one copy runs. If it fails, the release stops before any service changes.
- Rolling deployment. Each changed service is rolled to its new version and the pipeline waits for it to become stable; most services start new copies before stopping old ones, while the Studio API restarts with a short pause. If a service does not become stable, it is rolled back to its previous version automatically.
- Smoke tests. After a successful roll, the pipeline probes the API's health and readiness and the Studio. A failed smoke test is reported but does not roll the release back.
- Graceful shutdown. A stopping API copy stops taking new work, ends live progress streams with a final message, and waits up to 15 seconds for requests in flight. The worker stops taking jobs and has up to 60 seconds to finish the ones it holds; after that, a Builder or Agent turn still running is delivered again and resumes, and any other job is marked failed. The gateway does not wait, so a request in flight at that moment can be cut.
- Forward-only database changes. Migrations are not rolled back. A change that would break the previous release is split into steps, with the old shape kept until a later release, and the final step refuses to run until the data has been moved.
What to request from Flexday
These are commitments, not settings, so no page here can state them. Ask for:
- the availability commitment for the hosted service, and how availability is measured and reported;
- the recovery objectives Flexday commits to, and the date and outcome of the last restore test;
- the business continuity and disaster recovery plans;
- how incidents and maintenance are communicated, and the support hours and response targets.