QA Agents: test-case automation that reads your product, writes the tests, runs them live and heals itself
A crew of specialised agents on a living product knowledge graph, covering the full lifecycle from requirement to self-healing automation.
Offering
QA Agents - Agentic Quality Engineering
Scope
Requirement to self-healing automation
By
QK AI Labs, QualityKiosk Technologies
Executive summary
QA Agents, by QK AI Labs 🔍
The magnifier: I read your product, write the tests you are missing, run them live, and heal them when the screen changes.
QA Agents turn quality engineering from a manual, release-bound activity into an agentic, continuous one. A crew of specialised agents reads your requirements, source and existing test assets, builds a living product knowledge graph, and from it authors, optimises, scripts, executes and self-heals your tests.
The agents do not replace your framework or your people. They sit on top of what you already run, learn your product from day zero and from history, and keep a human in the loop at every gate that matters. The result is more coverage, less flakiness and faster regression, with full auditability.
155 to 53
raw cases optimised
30-70%
faster generation
<3%
flakiness, live customer
12h+
explorer session length
Where QA effort sits today, and where QA Agents move it
Figures are drawn from QK AI Labs engagements and are directional. First-pass gains are typically 30 to 50 percent and rise to 50 to 70 percent as the knowledge graph matures.
The crew
Seven agents, one knowledge graph 🧰🔍🗼
QA Agents is not a single bot. It is a crew of specialised agents, each with a job, all reading from and writing back to the same product knowledge graph. More sources connected means a better graph, which means better tests.
Requirement Analyst 🧰: ingests a BRD, user story or even a short write-up, grooms it, and flags gaps, missing prerequisites and a recommended structure before anything is built.
Graph Builder 🧰: turns structured and unstructured data, PDFs, Excel, images, source and configs, into a product trail that maps modules, coverage and downstream impact.
Test Designer 🔍: generates epics and stories for automation, detailed test plans, smoke, regression, security, concurrency and data-integrity suites, with exit criteria and a risk register.
Optimizer 🔍: merges and consolidates raw cases, removing duplication while preserving coverage.
Script Generator 🧰: writes automation in Playwright, Selenium, Robot or a hybrid, in whatever framework format your team uses.
Explorer, the AI Lens 🔍: navigates the live application, learns its business rules, captures stable locators, and spots fields and flows the requirement never mentioned.
Execution and Heal 🗼: runs the suite, triages failures, auto-heals broken scripts, and keeps flakiness low across environments.
One living product graph; a crew of specialised QA agents on top of it
How it works
From requirement to self-healing automation 🧰
The toolbox: I take whatever you have, a BRD, Jira, source, old Excel cases, and turn it into a structured product graph the testers can build on.
The pipeline runs left to right, but it is a loop. Every run feeds learning back into the graph, so the next cycle starts smarter. A requirement enters, the graph makes sense of it, agents author and script the tests, the Explorer validates them against the live product, and the execution agents run and heal.
How one requirement travels from input to self-healing automation
The core idea
The knowledge graph is the heart. Agents are deterministic where rules work and probabilistic only where they must be. Each is evaluated continuously across 30 to 40 parameters so it becomes more predictable over time, not less.
Use cases
Where QA Agents earn their place 🔍
The same crew covers the full test-case automation lifecycle. These are the recurring use cases QK AI Labs delivers against.
1. BRD and story to automation 🧰
A requirement in any format is groomed, gaps are surfaced, and it is structured into epics, stories and a recommended build. From there the agents generate test plans and cases ready for scripting.
2. Raw test cases to an optimised, automated suite 🔍
Legacy Excel cases are consolidated, de-duplicated and converted into your framework format, then turned into automation scripts. In one engagement 155 raw cases became 53 optimised cases at 87 percent accuracy.
3. Coverage expansion, horizontal and deep 🔍
The agents add boundary values, corner and edge cases, negative and combinatorial scenarios, and generate the test data to drive them, lifting coverage well beyond the manual baseline.
4. Environment and tenant delta 🔍
One core suite is kept in a shared delta package. The Explorer learns the differences between environments, versions and tenants and runs only the relevant cases, so there is no rewrite and no duplicate suites to maintain.
5. Incident to coverage, with RCA 🗼
When a defect appears, the agents investigate, produce a root-cause analysis, write the missed test scenario back into the repository and build the automation for it, so the same gap cannot recur.
6. Framework migration and modernisation 🧰
A migration agent reads an existing framework, narrates what it understood for confirmation, then converts gradually, Selenium to Playwright or the reverse, growing scope as accuracy is proven.
7. Backend, batch and file-processing testing 🗼
For long-running cron and batch jobs, a dedicated testing agent drives the workflow, waits, varies file formats and fields, and checks logs, APIs and the database, while a separate analyzer agent wakes on a schedule to produce insight.
Agents move fast; humans stay in the loop at the gates that matter
Handling variation
Deterministic where it works, agentic where it must 🔍
Most environment and tenant differences, perhaps 70 to 80 percent, can be handled with deterministic inputs. The remainder, the minute config-driven behaviour that rules cannot capture, is exactly where the agents add a brain.
Keep what rules do well; let agents cover the minute changes rules cannot
The Explorer takes the base suite into a new environment or tenant, validates each case against what the live product actually does, captures the delta and feeds it back. The common suite stays single and maintainable; the agents resolve the difference at run time.
The AI Lens
The Explorer: a brain for automation 🔍
The magnifier: I do not watch a recording. I navigate your live app, understand each page, and bring back the rules, the locators and the fields nobody documented.
Classic automation is a black box with no sense of what the application looks like. The AI Lens closes that gap. It navigates the live product in real time, builds a narrative of what it understood, and captures stable locators, business rules and edge boundaries.
Stable locators: it prefers reusable, parameterised locators over brittle spans, so scripts survive change.
Business rules: it infers and records rules, for example a date picker that disables past dates, and links them back to the matching cases.
Guardrails: it stops short of destructive actions such as changing a password or creating users, captures what it found and asks before proceeding.
Reviewer agents: a second agent checks the first, catching missed rules and enforcing locator standards.
Self-healing: when a locator or path is missing it finds a fallback, keeping flakiness under 3 percent on a live customer.
Platform
How it is built and deployed 🗼
The lighthouse: single tenant, access-controlled, and governed, so nothing sensitive leaves and every action is auditable.
QA Agents is built on a retrieval-augmented knowledge graph with its own small model for native context, calling larger models only for reasoning tasks. A governance layer ensures only metadata is sent out, never raw data as-is.
Deployment: start on a managed cloud instance for a fast first result, then move into your VPC for direct, native connections once you are satisfied. Both are supported.
Connectors: native, bidirectional links to Jira, source control, release notes and more; every change pushes an update into the graph.
Access control: role-based access throughout, with a credentials repository and backups so agents can log in and run reliably.
Governance: a single-tenant, living repository; models are chosen by context; prompts and best practices are configurable by your team.
Evaluation: every agent is evaluated across 30 to 40 parameters with continuous observability to guard against drift and hallucination.
Adoption
Earned autonomy, compounding value 🔍
The honest expectation: a 30 to 50 percent gain in the first pass, rising to 50 to 70 percent as the graph picks up context. An SME is factored in early to baseline the graph; from there it gets easier.
Value compounds as the product graph deepens; autonomy is earned, not assumed
What to do next
Pick one application. Point the crew at it, baseline the graph with a short SME window, and measure the first-pass gain. Then expand application by application and use case by use case.
Let us run it on your application
Start with one application
Share one application and its existing assets. We will baseline the knowledge graph, generate and run the first suite, and show the gain before you commit further.
Shakthi
General Manager - QK AI Labs, QualityKiosk Technologies