AI test generation starts with context, not prompts
This is the first in a short series on moving a QA function from classic automation to AI-assisted automation — not as a rip-and-replace, but as a series of changes each of which has to earn its place. I'm starting where most teams start when they try to "let AI write the tests," because it's also where most of them quietly fail: context.
Here's the failure mode. Someone pastes a Jira ticket into a model and asks for test cases. What comes back looks great — well-formatted, plausible, confident. Then a senior QA reads it and every problem is the same shape: it tests what the ticket says, not what the feature is. It misses the edge case two developers argued about in a pull request. It re-creates a case that already exists in TestRail under a different title. It writes coverage for a behavior that a Slack thread killed three weeks ago. The output isn't wrong because the model is weak. It's wrong because it was working from a fraction of the context a human would have gathered first.
Classic automation quietly ran on context you already had
Think about what actually happened when you wrote a test the "classic" way. You read the ticket, sure — but you also remembered that this area already has regression cases, so you didn't duplicate them. You'd seen the PR, so you knew the code added a locale fallback the ticket never mentioned. You were in the channel where the team decided the feature ships dev/stage only. You knew where the preconditions come from. None of that was in the ticket. It was in your head, assembled from five different systems you'd been marinating in for months.
That's the part people skip when they picture "AI writing tests." The transition to AI-assisted automation isn't replacing the tester's fingers with a model. It's taking the context the tester carried implicitly and making it available explicitly — to a machine that has no months of marination to draw on. Get that right and generation becomes genuinely useful. Skip it and you've built a very fast way to produce plausible, shallow tests.
The truth is scattered, and each source answers a different question
When I map where the real specification of a feature actually lives, it's never one place:
- TestRail (or your test management system) — what's already covered. Without this, an AI cheerfully generates duplicates of cases you've had for a year.
- The Jira ticket and its acceptance criteria — the requirement, and, more importantly, the comments where the AC quietly changed. The title says one thing; comment 14 says the real thing.
- The GitHub PR — the diff is what the code actually does, and the review comments are where the edge cases got argued into existence. This is the single richest source of "what could break," and it's the one prompt-a-ticket workflows never touch.
- Slack (or wherever decisions happen) — the "we decided not to support X" that never made it back to the ticket. Negative requirements almost only live in chat.
- Test data & config — the preconditions, fixtures, and feature flags a case needs to even run.
A model fed only the ticket is missing four of these five. A senior QA reconstructs all five without thinking about it. The whole game is closing that gap.
The pattern: pull everything into one grounded context store first
So the move isn't a better prompt. It's an aggregation step that happens before any generation. Pull each source — read-only — into a single local store, keep the full raw payload so nothing is lost or paraphrased away, and make it searchable both by keyword and by meaning. Now, when the AI drafts cases for a feature, it isn't guessing from a title. It retrieves the actual acceptance criteria, the actual PR diff and its review thread, the existing cases that are semantically close, and the decision from chat.
The semantic layer matters here for the same reason it does in retrieval for QA: the model needs to find the related existing case even when it's worded completely differently, or it'll duplicate it. Aggregation without meaning-aware retrieval just gives you a bigger pile to keyword-miss through.
What actually changes in the output
Make the context complete and the drafts stop being generic. A worked example, at the pattern level:
A ticket says "add a promo banner." Fed just that, a model writes five reasonable-looking banner tests. Fed the real context, it sees that the PR review flagged a locale-fallback edge case, that a Slack thread scoped the feature to dev/stage only, and that the suite already has three banner cases. Now the useful draft is different: two new cases — the locale fallback and the environment guard — plus a flag that the three existing cases need a small update, not three more near-duplicates.
That's the shift. The context-blind model optimizes for "produce test-shaped text." The context-fed one starts to make the calls a QA would: what's genuinely new, what already exists, what the team explicitly excluded. It's still not a senior QA — but it's now drafting from the same material one would use, instead of from a headline.
Generation is a draft, never a write
One hard line, because it's the line that makes any of this safe to run near a real team: better context raises the quality of the drafts, but it never removes the review gate. The AI produces a draft; nothing reaches TestRail without going through a dry-run, a diffable and hashed plan, and an explicit human confirmation. Context makes the drafts good. The guarded-write chain keeps them from doing damage when the retrieval was subtly off — because sometimes it will be, and a slightly-too-loose match that fires straight into your test suite is exactly the incident you don't want to explain later.
I'll go deep on that dry-run → plan-hash → confirm mechanism in a later post, because it deserves its own. For now the point is just that it sits downstream of everything above: context is what makes AI test generation worth doing; the guard is what makes it responsible to do.
When it's worth building
This is overkill for a small suite one person holds in their head. It starts paying off the moment the real specification of a feature is smeared across five systems and only a long-tenured engineer can reassemble it — which is precisely the situation where a team is also tempted to "just let AI write the tests" and gets burned by shallow output.
When I move a team from classic to AI-assisted automation, this aggregation layer is usually the first thing I stand up, before any generation — it's the foundation the rest of the series builds on: bringing chat context in, feeding PR diffs into drafting, and the guarded-write mechanics that keep all of it safe. If that's where your team is — real coverage scattered across tools, growing pressure to use AI, and shallow results so far — this is where I'd start.
Dealing with something similar on your team? Let's talk.