How to Audit an AI-Generated Database Migration Before It Touches Production

An open notebook with handwritten notes on a desk

An AI coding assistant can write a schema migration in about four seconds that would have taken a developer twenty minutes to draft by hand. The SQL is usually syntactically correct. It usually passes the test suite. And it can still be the single most dangerous file in a pull request, because the assistant has no idea how many rows are in your orders table, whether your ORM holds a connection open for the duration of the migration, or that the column it just suggested dropping is still read by a reporting job nobody remembers writing. Migrations are the one category of AI-generated code where "it compiles and the tests pass" tells you almost nothing about whether it's safe to run.

Why Database Migrations Are a Different Risk Than Other AI-Generated Code

Most AI-generated code gets a forgiving safety net. A bad function gets caught by a unit test, a bad component gets caught by a visual regression, a bad API handler gets caught by an integration test hitting a staging environment. Migrations don't get that same cushion, because the thing that makes a migration dangerous usually isn't the SQL being wrong. It's the SQL being correct, but expensive, blocking, or irreversible at a scale the assistant never saw.

A migration that adds a NOT NULL column with a default value is a one-line diff. On a table with a few thousand rows it runs in milliseconds. On a table with forty million rows, depending on the database engine and version, that same line can rewrite the entire table and hold a lock that blocks every write for the duration. The assistant produces identical code in both cases, because it has no visibility into table size, traffic patterns, or what else depends on that table staying responsive during business hours.

What AI Coding Assistants Consistently Get Wrong About Schema Changes

The failure pattern repeats across tools and models, not because any particular assistant is poorly built, but because the context it's missing is structural. It doesn't know your table sizes unless you paste them in. It doesn't know your database engine's specific locking behavior for a given operation unless you ask about that engine by name. It doesn't know which of your services still read a column you're about to drop, because that knowledge lives in application code the assistant wasn't asked to search.

The result is migrations that are locally correct and globally risky. A DROP COLUMN the assistant suggests to clean up dead code looks obviously safe in isolation. Whether it's actually safe depends entirely on whether anything still reads that column, which is a question about the rest of your codebase, not about the migration file itself.

The Column Drop That Looks Safe Until It Isn't

Dropping a column is the single most common place this goes wrong in practice. An assistant asked to remove an unused field will happily generate the migration, and it will be correct SQL. What it won't do is grep your codebase for every place that column is still selected, serialized, or referenced in a background job that only runs on the first of the month.

A terminal window showing monospace code close up
Photo by Ec lipse on Pexels

The safer pattern, and one worth asking the assistant to follow explicitly rather than assuming it will, is to stop writing to a column first, deploy that change, confirm nothing errors for a full cycle of whatever batch jobs touch the table, and only then schedule the actual drop as a separate migration. An AI assistant asked for "a migration to remove this column" will give you the one-step version unless you specifically ask for the staged one.

Why Passing Tests Doesn't Mean a Migration Is Safe to Run

Most test suites run against a small, fast, freshly-seeded test database. That environment is close to useless for catching the two failure modes that actually hurt in production: lock duration under real data volume, and lock contention against real concurrent traffic. A migration can pass a CI pipeline cleanly in eight seconds and then hold an exclusive lock for eleven minutes against a production table that has four orders of magnitude more rows than the test fixture.

This is also why the assistant's own confidence is a poor signal here. It will describe a migration as safe based on syntax validity and successful test execution, which is a genuinely different claim than "this won't block production traffic." Treat any migration the assistant produced as untested for the specific thing that matters most until you've checked it against something closer to real data volume.

A Pre-Flight Checklist Before You Run Anything the AI Wrote

Before an AI-generated migration goes anywhere near production, a short checklist catches most of the dangerous cases. Confirm the table's approximate row count and whether the operation is one your database engine performs as a fast metadata-only change or a full table rewrite, since the same statement can behave very differently depending on engine and version. Confirm whether the migration acquires a lock that blocks reads, blocks writes, or both, and for roughly how long given the table's real size.

Check whether the migration runs inside the deploy's normal traffic window or whether it needs to be scheduled for a quieter period. Confirm a rollback path exists and has actually been tested, not just written. Flyway and Liquibase both document engine-specific locking behavior for common operations, and checking the relevant doc page for your specific database and version takes a few minutes against a migration that could otherwise cost hours of incident response.

Expand-Contract: Running Migrations in Two Steps Even When the AI Suggests One

The expand-contract pattern, sometimes called parallel change, splits a risky migration into two deploys instead of one. The expand step adds the new structure, column, table, or index, without removing or renaming anything the application still depends on. Application code gets updated to use the new structure. Only once that's deployed and confirmed stable does the contract step remove the old structure.

An AI assistant asked for "a migration to rename this column" will typically produce a single-step ALTER TABLE ... RENAME COLUMN, which is fast on most engines but still means any code still referencing the old name breaks the instant the migration runs, with no window to catch it. Asking explicitly for the expand-contract version, add the new column, backfill it, switch reads over, then drop the old column in a later migration, turns an atomic risk into a sequence of independently reversible steps.

Index Builds and Lock Behavior the Assistant Won't Warn You About

Adding an index sounds like one of the safer migration categories, and on a small table it usually is. On a large table, a standard index build can lock writes for the entire duration of the build, which on a sufficiently large table can run for minutes or longer. Most mainstream database engines offer a concurrent or online index-build mode specifically to avoid this, but it's slower, has different failure semantics, and an assistant won't default to it unless asked.

A server rack with neatly organized cables in a data center
Photo by Vladimir Srajber on Pexels

The same caution applies to adding a foreign key constraint on an existing table, which on some engines requires a full table scan to validate existing data against the new constraint, another operation that behaves very differently at ten thousand rows versus forty million. Ask for the non-blocking or concurrent variant explicitly when the table in question is one that matters during business hours.

The Human Review Step That Still Has to Happen

None of this means AI-generated migrations shouldn't be used. It means the review step for a migration needs to ask different questions than the review step for a typical pull request. A reviewer looking at application code is mostly checking logic and style. A reviewer looking at a migration needs to be checking table size, lock behavior, rollback path, and whether anything downstream still depends on the structure being changed.

"The AI will write syntactically perfect SQL that still takes your production database down for ten minutes under load. It doesn't know your table is forty million rows or that your ORM holds a connection open for the duration of the migration window." - Dennis Traina, founder of 137Foundry

That review doesn't need to be a full-time DBA on every team. It needs to be someone who has specifically been asked to check those four things, with the authority to send a migration back for restructuring rather than rubber-stamping something that passed CI.

Writing the Rollback Plan Before You Need It

A rollback plan written after a migration has already started causing problems is written under pressure, by someone who's also trying to diagnose what's actually wrong in real time. That's the worst possible condition for writing correct recovery SQL. The rollback plan belongs in the same pull request as the migration itself, reviewed with the same scrutiny, before anything runs against production.

For additive changes, the rollback is often straightforward, drop the column or table that was just added. For anything destructive, a dropped column, a changed data type, a removed constraint, the honest rollback plan is frequently "restore from a pre-migration backup," which is a very different operational reality than a one-line reverse migration, and worth knowing explicitly before you run the forward migration rather than discovering it during an incident.

Watching the Migration After It Ships

Even a correctly reviewed migration deserves active monitoring while it runs against production, not a fire-and-forget deploy. Watch query latency and error rates on the affected table during and immediately after the migration window, not just the migration job's own exit code. A migration that completes successfully by its own accounting can still have degraded the table's performance in a way that only shows up in downstream query times.

A network operations center with several monitors showing data
Photo by Brett Sayles on Pexels

This is also where having existing data quality and performance monitoring in place pays off well beyond the migration itself. 137Foundry's data integration work typically includes this kind of production visibility as a baseline, specifically because migrations are one of the more common places a quiet performance regression gets introduced without anyone noticing until days later.

Making This Part of the Standard Workflow, Not a One-Off Precaution

Teams that handle this well treat the migration-specific review checklist as a standing part of their process, not something pulled out only after a bad incident. New engineers get onboarded to the checklist the same way they get onboarded to the normal code review process. The checklist itself gets revisited and expanded every time a migration causes a problem the checklist didn't already catch, which keeps it a living document rather than a one-time artifact that quietly goes stale.

The underlying concept of a schema migration, and the operational discipline around running one safely, has a well-established body of practice behind it that predates AI-assisted coding entirely; schema migration as a discipline has always required this kind of care, and the same principles around ACID guarantees that governed manual migrations for decades still govern AI-generated ones. The assistant changed how fast the SQL gets written. It didn't change what the database actually does when that SQL runs.

137Foundry treats migration review as a standard part of AI automation engagements rather than an afterthought bolted on after a team has already had an incident. If your team is shipping AI-generated schema changes today without a checklist like this one in place, that's worth fixing before the next migration, not after it. Our services page covers how we approach this kind of production reliability work, and our about page has more on the engineering philosophy behind it.

Need help with your next project?

137Foundry builds custom software, AI integrations, and automation systems for businesses that need real solutions.

Book a Free Consultation View Services