How We Turned Cursor into a True Senior QA Assistant
ByAlan Gershman
QA Manager at Papaya
Ask any QA professional what their least favorite part of the job is, and you’ll probably get the same answer: the "gap" where a feature moves from a product spec to a test plan. As a QA Manager at Papaya, I’ve seen many talented teams burn out against this wall of manual friction. It’s the moment where the excitement of “Hey, we have a new feature!” meets the gray reality: QA engineers sifting through mountains of requirements, trying to decipher what the developer actually wrote in the code (versus what they meant to do), and manually documenting endless test cases.
This process isn't just slow, it’s highly prone to human error. When you’re operating at a scale of tens of millions of users, a single missed edge case can snowball into a production disaster. I found myself wondering: if our developers are already leveraging AI to streamline their coding, why should QA stay in the dark ages? When we started building the Papaya QA Assistant, our goal was simple: stop being "data entry clerks" and start being investigators.
What began as a small time-saving utility quickly evolved into a sophisticated AI agent deeply integrated into our development core.
An Agent That Owns the Entire Workflow
One of the biggest hurdles with AI is its tendency toward shallow answers or hallucinations. Developing a production-ready AI agent requires more than just stochastic guessing; it demands deep integration with the source of truth.
We decided to give our Assistant "the keys to the kingdom." Using MCP (Model Context Protocol), we connected the AI to our entire ecosystem: our Test Management System (Qase), requirement docs and task tracking (Jira/Confluence), and the codebase itself (GitHub). The goal was to move the AI from guessing to knowing. It was a game-changer.
By shifting to a "Personal Assistant" model, we changed the rules of the game. The tool no longer just generates text, it manages an entire end-to-end workflow:
1. One-Click Test Case Generation
In our new workflow, the manual heavy lifting is gone. The QA engineer simply pastes a link to a PRD (Product Requirement Document) or a Jira task into Cursor, and the QA Assistant automatically extracts the relevant context to build a full test plan - identifying features, prerequisites, and expected results.
But the real magic is the style. The AI doesn’t write like a generic robot, it speaks our language. It understands our specific technical terminology, how we structure suites, and our internal documentation standards. The output looks like it was written by a veteran team member, not an algorithm. Once the engineer approves the plan, it’s pushed to Qase automatically. No more copy-pasting, no more formatting errors.
2. Deep PR Analysis: Diving into GitHub
This is where code-level intelligence shines. Instead of relying solely on a potentially outdated spec, the Assistant scans the actual GitHub Diffs. It analyzes method changes, identifies impacted files, and even "listens" to the conversation in Code Reviews. If a developer comments to a peer, "Careful, this might crash on a slow connection," the Assistant catches it and automatically generates a corresponding test case.
3. Impact Analysis: Smart Regression
One of the hardest questions in QA is: "What do I need to test for this tiny change?" Instead of running a "blind" regression on the entire system - wasting days of work - the AI analyzes the "blast radius" of the change. It can flag: "This change touched the payment logic, let's trigger these 15 specific critical tests." This surgical approach allows us to ship faster and with higher confidence.
4. The Developer Handoff
To truly shift quality left, the Assistant generates a pre-deployment checklist for our developers. By synthesizing the requirements with the actual code changes, it provides a tailored set of verification tests for the developer to run before the handoff. This eliminates the exhausting 'ping-pong' between Dev and QA, ensuring high-integrity code reaches us the first time.
5. Automation: When the Tool Knows the Framework
Because the Assistant lives in the codebase, it "knows" our existing automation suites. When analyzing a new PR, it doesn't just suggest manual tests - it identifies automation gaps. If it sees new functionality without coverage, it can generate end-to-end tests based on our internal framework. No one has to write boilerplate scripts from scratch anymore.
6. Test Case Health Check
To maintain a high standard of quality over time, we’ve implemented a Health Checker for our entire test repository. When managing a repository of thousands of test cases, they naturally tend to "decay" - duplicate tests accumulate, priorities drift out of date, and hollow test cases (those missing clear steps or expected results) begin to creep in.
Today, a single command in Cursor triggers the Assistant to perform a deep scan of our entire testing system, identifying quality gaps across multiple categories and proposing instant, automated fixes. It’s like sending a forensic auditor to perform a comprehensive sweep of every single line in your testing infrastructure - except it finishes in 30 seconds and never misses a detail.
How It Works in Practice
The execution is seamless:
Input: Paste a PRD link or Jira ID into Cursor.
Analysis: The Assistant analyzes the spec and checks Qase to prevent duplicate test cases.
Correlation: It locates the code changes in GitHub and cross-references them with the requirements.
Generation: It builds a prioritized test plan with automated tagging.
Approval: The QA engineer reviews the output. With one click, it’s live in Qase.
Human-in-the-Loop: Quality Over Hype
To be clear: we don't let the AI run wild. We operate on a Clean Output philosophy. The Assistant doesn't clutter the UI with technical logs, it presents a clean, actionable test plan.
Every AI-generated output requires human validation. The AI does 80%–90% of the manual labor, but the final "green light" to push tests or approve a release remains in the hands of our professionals. It’s the perfect marriage of machine speed and human judgment.
We are already testing our next milestone: Root Cause Analysis. Today, you can describe a bug in plain English, and the Assistant will dive into GitHub, find the specific commit and line of code responsible, and even propose a fix.
Reclaiming Time for Strategy and Research
When we built this, we hoped to save time. The actual result was a shift in team culture. Tasks that used to take half a day of deciphering and documentation now take minutes of review. This has freed our team to focus on what actually matters: high-level strategy, hunting for complex logical bugs in unexpected places, and perfecting the user experience. The AI took over the "robotic" work, leaving us with the parts that require creativity and human intelligence. The tests might be starting to write themselves, but the value of our QA engineers has never been more apparent. We’ve simply given them the ultimate toolkit to do what they do best.
Ask any QA professional what their least favorite part of the job is, and you’ll probably get the same answer: the "gap" where a feature moves from a product spec to a test plan. As a QA Manager at Papaya, I’ve seen many talented teams burn out against this wall of manual friction. It’s the moment where the excitement of “Hey, we have a new feature!” meets the gray reality: QA engineers sifting through mountains of requirements, trying to decipher what the developer actually wrote in the code (versus what they meant to do), and manually documenting endless test cases.
This process isn't just slow, it’s highly prone to human error. When you’re operating at a scale of tens of millions of users, a single missed edge case can snowball into a production disaster. I found myself wondering: if our developers are already leveraging AI to streamline their coding, why should QA stay in the dark ages? When we started building the Papaya QA Assistant, our goal was simple: stop being "data entry clerks" and start being investigators.
What began as a small time-saving utility quickly evolved into a sophisticated AI agent deeply integrated into our development core.
An Agent That Owns the Entire Workflow
One of the biggest hurdles with AI is its tendency toward shallow answers or hallucinations. Developing a production-ready AI agent requires more than just stochastic guessing; it demands deep integration with the source of truth.
We decided to give our Assistant "the keys to the kingdom." Using MCP (Model Context Protocol), we connected the AI to our entire ecosystem: our Test Management System (Qase), requirement docs and task tracking (Jira/Confluence), and the codebase itself (GitHub). The goal was to move the AI from guessing to knowing. It was a game-changer.
By shifting to a "Personal Assistant" model, we changed the rules of the game. The tool no longer just generates text, it manages an entire end-to-end workflow:
1. One-Click Test Case Generation
In our new workflow, the manual heavy lifting is gone. The QA engineer simply pastes a link to a PRD (Product Requirement Document) or a Jira task into Cursor, and the QA Assistant automatically extracts the relevant context to build a full test plan - identifying features, prerequisites, and expected results.
But the real magic is the style. The AI doesn’t write like a generic robot, it speaks our language. It understands our specific technical terminology, how we structure suites, and our internal documentation standards. The output looks like it was written by a veteran team member, not an algorithm. Once the engineer approves the plan, it’s pushed to Qase automatically. No more copy-pasting, no more formatting errors.
2. Deep PR Analysis: Diving into GitHub
This is where code-level intelligence shines. Instead of relying solely on a potentially outdated spec, the Assistant scans the actual GitHub Diffs. It analyzes method changes, identifies impacted files, and even "listens" to the conversation in Code Reviews. If a developer comments to a peer, "Careful, this might crash on a slow connection," the Assistant catches it and automatically generates a corresponding test case.
3. Impact Analysis: Smart Regression
One of the hardest questions in QA is: "What do I need to test for this tiny change?" Instead of running a "blind" regression on the entire system - wasting days of work - the AI analyzes the "blast radius" of the change. It can flag: "This change touched the payment logic, let's trigger these 15 specific critical tests." This surgical approach allows us to ship faster and with higher confidence.
4. The Developer Handoff
To truly shift quality left, the Assistant generates a pre-deployment checklist for our developers. By synthesizing the requirements with the actual code changes, it provides a tailored set of verification tests for the developer to run before the handoff. This eliminates the exhausting 'ping-pong' between Dev and QA, ensuring high-integrity code reaches us the first time.
5. Automation: When the Tool Knows the Framework
Because the Assistant lives in the codebase, it "knows" our existing automation suites. When analyzing a new PR, it doesn't just suggest manual tests - it identifies automation gaps. If it sees new functionality without coverage, it can generate end-to-end tests based on our internal framework. No one has to write boilerplate scripts from scratch anymore.
6. Test Case Health Check
To maintain a high standard of quality over time, we’ve implemented a Health Checker for our entire test repository. When managing a repository of thousands of test cases, they naturally tend to "decay" - duplicate tests accumulate, priorities drift out of date, and hollow test cases (those missing clear steps or expected results) begin to creep in.
Today, a single command in Cursor triggers the Assistant to perform a deep scan of our entire testing system, identifying quality gaps across multiple categories and proposing instant, automated fixes. It’s like sending a forensic auditor to perform a comprehensive sweep of every single line in your testing infrastructure - except it finishes in 30 seconds and never misses a detail.
How It Works in Practice
The execution is seamless:
Human-in-the-Loop: Quality Over Hype
To be clear: we don't let the AI run wild. We operate on a Clean Output philosophy. The Assistant doesn't clutter the UI with technical logs, it presents a clean, actionable test plan.
Every AI-generated output requires human validation. The AI does 80%–90% of the manual labor, but the final "green light" to push tests or approve a release remains in the hands of our professionals. It’s the perfect marriage of machine speed and human judgment.
We are already testing our next milestone: Root Cause Analysis. Today, you can describe a bug in plain English, and the Assistant will dive into GitHub, find the specific commit and line of code responsible, and even propose a fix.
Reclaiming Time for Strategy and Research
When we built this, we hoped to save time. The actual result was a shift in team culture. Tasks that used to take half a day of deciphering and documentation now take minutes of review.
This has freed our team to focus on what actually matters: high-level strategy, hunting for complex logical bugs in unexpected places, and perfecting the user experience. The AI took over the "robotic" work, leaving us with the parts that require creativity and human intelligence.
The tests might be starting to write themselves, but the value of our QA engineers has never been more apparent. We’ve simply given them the ultimate toolkit to do what they do best.