Testing AI-Generated Code Before Shipping: A Production Workflow
AI-generated code should enter production through the same accountable engineering system as human code, with additional controls for agent scope, untrusted execution and plausible-but-unverified output. The safe workflow is not βgenerate, glance, ship.β It is bound, inspect, test, review, stage, observe and retain rollback.
Use this guide with the Coding Agents hub and AI Testing & Evaluation.
1. Define the change boundary
Before generation, record:
- the exact objective and acceptance criteria;
- allowed repositories, paths and services;
- files or systems that must not change;
- whether dependencies or migrations are permitted;
- required build and test commands;
- whether commit, push or deployment is authorized;
- the human owner of the result.
Run the agent in an isolated branch or worktree with the minimum permissions needed. Do not give repository-writing, cloud and production-data access merely because the task may eventually deploy.
2. Treat the diff as untrusted input
Inspect the complete change before executing project-controlled scripts with secrets. Agent code can modify package scripts, CI workflows, build hooks and configuration as easily as application logic.
Check for:
- unexpected files, generated artifacts and mass formatting;
- new dependencies or lockfile changes;
- modified tests, snapshots or quality thresholds;
- workflow permissions and secret access;
- migrations, deletion logic and external calls;
- embedded credentials or sensitive logs;
- unrelated βhelpfulβ refactors.
If scope is wrong, correct it before deeper testing. Passing tests do not authorize an overbroad change.
3. Run fast deterministic checks
Start with inexpensive gates:
- syntax and type checking;
- formatting and linting;
- generated-code or schema validation;
- unit tests for changed logic;
- dependency and secret scanning;
- diff-integrity checks.
Unit tests should include success, boundary and failure behavior. Confirm new tests fail against the broken or previous implementation when feasible; otherwise they may simply document what the generated code already does.
4. Test integration boundaries
Agent-generated code often looks coherent inside one file while violating a real contract. Exercise:
- database reads, writes and migrations;
- authentication and authorization;
- queues, webhooks and retries;
- third-party APIs through test doubles or sandboxes;
- serialization and structured outputs;
- backwards compatibility;
- cancellation and partial failure.
Keep model, tool and external-service behavior controlled when the purpose is deterministic application testing. Evaluate variable AI output separately through the LLM regression pipeline.
5. Sandbox consequential execution
Use ephemeral environments, disposable data and least-privilege credentials. Block or fake irreversible actions such as payments, emails, account deletion and production deployment.
For coding agents, sandboxing should limit:
- filesystem paths;
- network destinations;
- secret availability;
- cloud and package-registry permissions;
- runtime duration and resource use;
- ability to invoke nested automation.
A sandbox reduces blast radius; it does not prove code correctness. Inspect outputs and final state after execution.
6. Validate user and system workflows
Run end-to-end tests for the critical paths affected by the change. Assert observable outcomes, not implementation details:
- the user reaches the intended state;
- authorization fails for the wrong identity;
- retries do not duplicate side effects;
- errors remain recoverable;
- monitoring receives the expected identifiers;
- the application can revert or continue safely.
Use browser automation only for behavior that genuinely crosses the UI. Playwright vs Cypress for AI applications and AI-generated Playwright tests explain how to keep those tests deterministic and reviewed.
7. Protect the CI pipeline
Bounded case: GLMβs inference infrastructure
Z.aiβs infrastructure account describes engineers using GLM agents with defined goals, correctness tests, traces and benchmark feedback. Its useful engineering lesson is the feedback loop: reject numerically incorrect kernels, inspect system traces, and measure the integrated serving path rather than accepting a faster isolated microbenchmark as a production win. These are vendor-reported results, not our reproduction. Human engineers still set boundaries and validate changes; the account should not be treated as demonstrated recursive self-improvement.
For your own agent-generated optimization, preserve a baseline, test representative shapes and failure cases, compare end-to-end latency/throughput under the same workload, and require an independent release gate. Dense local feedback can improve iteration without replacing human acceptance criteria.
Separate untrusted pull-request checks from privileged deployment jobs. The validation job should run with minimum token permissions and no production secrets. Deployment credentials should become available only after protected checks and required human approval.
A practical gate sequence is:
pull request
β scope and diff review
β build, lint, unit and security checks
β isolated integration/E2E tests
β required human approval
β protected staging deployment
β smoke and migration verification
β production approval
Use branch protection or repository rules so required checks and approvals cannot be skipped casually. See GitHub Actions for AI applications.
8. Deploy progressively
Do not make βmergedβ equivalent to βsafe everywhere.β Depending on risk:
- deploy first to staging or an ephemeral preview;
- smoke-test the changed behavior;
- verify migrations and configuration;
- use a canary or feature flag for user-facing changes;
- monitor error, latency and business signals;
- expand only when predefined gates pass.
For model, prompt and RAG changes, use canary deployments for LLM features. AI Operations covers ownership, observability and rollback.
9. Make rollback real
Before production, document:
- the last known-good commit and artifact;
- whether schema changes are backward compatible;
- how feature flags or routing return traffic to stable behavior;
- what data created by the new version requires reconciliation;
- who can trigger rollback;
- which signals trigger it automatically or manually.
Test rollback in a representative environment. A plan that depends on an irreversible migration or an unavailable operator is not a rollback plan.
10. Preserve evidence
Attach to the pull request or release:
- task and scope;
- final diff;
- build and test results;
- security and dependency findings;
- reviewer approval;
- deployment and smoke results;
- known limitations;
- rollback target and owner.
The distinction between AI code review and AI testing is important here: review explains why the change appears acceptable, while tests and deployment checks provide execution evidence.
AI coding speed is valuable only when the release system can absorb it. A bounded, testable and reversible change is production work. An impressive diff without independent evidence is still a draft.