Production Delivery for Forward Deployed Engineers
Updated · Tech checked
Shipping an FDE project to a customer's production means surviving their security review, deploying through their change process, instrumenting what you shipped, and planning rollback before you need it - a checklist-driven discipline, not heroics.
The direct answer
Production delivery is the FDE's defining discipline: the work isn't done when it runs on your laptop, or even in staging - it's done when it runs reliably in the customer's environment, is observable, can be rolled back, and the customer's team can operate it without you. This guide is the checklist we use.
1. Get through security review without dying
Customer security review kills more FDE deployments than bugs do. Prepare early:
- Data flow map. What data leaves their systems, where it transits, where it rests, for how long, encrypted how.
- Auth model. Least-privilege service accounts; secrets in their vault, never in your code or notebook.
- PII inventory. Which fields are personal data? What's masked or excluded? Who can query what?
- Pen-test posture. Know their scan windows; expect findings and budget fix time.
2. Environment strategy
- Match staging to production as closely as their infrastructure allows; where impossible, document the delta and test compensating paths.
- Config via environment, never hardcoded; secrets injected at runtime.
- Fly the deployment path once: build → artifact → deploy → verify → rollback, in staging, before the real window.
3. Rollout patterns
| Pattern | When to use | Key risk |
|---|---|---|
| Big bang | Low-traffic internal tools | All errors hit everyone |
| Canary (slice of traffic/data) | High-volume services | Two code paths to maintain |
| Shadow run (produce, don't act) | Writes to critical systems | Divergence between shadow and real |
| Feature flag | Uncertain adoption | Flag debt; someone must clean up |
For customer data writes, prefer shadow or dry-run modes first: produce the new outputs, diff against the old system, and let the customer eyeball discrepancies before you switch over.
4. Observability you ship, not promise
- Structured logs with correlation IDs; error rates and latency histograms on a dashboard the customer can view.
- Alerting on symptoms (error rate, freshness lag), not causes.
- Cost meters for anything AI: tokens, requests, per-workflow cost - surprises here end projects.
5. Rollback: design it before you need it
- Migrations are expand/contract: additive first, backfill, then switch reads, then remove old paths.
- Every deploy has a one-command rollback, tested.
- Data writes are idempotent and reconcilable; keep a repair runbook.
6. Handover
The most forgotten step - and the one that defines whether the customer calls you again:
- Runbook: what it does, how it fails, how to tell, what to do.
- Architecture diagram and data-flow map.
- Ownership table: every component has a name attached.
- A 60-minute walkthrough session, recorded.
7. The delivery checklist
- [ ] Security review passed, findings closed or waived in writing
- [ ] Secrets in customer vault; repo scanned for leaks
- [ ] Staging deployment verified end-to-end
- [ ] Rollback tested in staging
- [ ] Dashboards + alerts live, customer has access
- [ ] Runbook written and reviewed by the customer's ops team
- [ ] Success metrics defined before go-live, reviewed 2 weeks after
- [ ] Handover session held, recorded, ownership table signed off
"It worked in staging" is a phase, not a milestone. The milestone is: the customer's engineer can debug it without calling you.
For practice, the public capstone briefs offer this checklist for self-review; external project review is not available. The production articles cover individual techniques in depth.
Straight answers
Frequently asked questions
What does production delivery mean for an FDE?
Taking a customer solution through security review, staged rollout, observability, tested rollback, and documented handover so the customer's team can operate it independently.
How do FDEs handle customer security reviews?
Prepare a data-flow map, PII inventory, least-privilege auth model, and scan-window plan early; expect findings and budget time to fix them before the deployment window.
What rollout strategy should FDEs use for data writes?
Shadow or dry-run first: produce outputs without acting, diff against the legacy system, let the customer review, then switch over with tested rollback and reconcilable writes.
Sources
- Google SRE Book - Production Readiness Review (checked 2026-09-26)
- 12factor.net - Config (checked 2026-09-26)