← All posts

Answer · AI Agents & Automation

How Are AI Agents Used in Software Development?

Code generation, review, testing, and operations: where agents genuinely help engineering teams, and the review gate that keeps them safe.

Asaasin EngineeringPublished September 22, 20267 min read

In short

AI agents draft code, tests, and documentation faster than a human alone, but a named engineer still reviews and owns every line before it ships. In practice, agents handle generation, review support, test-writing, and narrow production tasks like transcription or document scanning, while a pull request, typed contracts, and CI tests remain the gate to production.

Where agents actually speed up code generation

Inside our pods, agent tooling does the first draft of a function, a migration, a component, or a test file. The stack includes Claude Code, GitHub Copilot, Cursor, and LangGraph, and the capabilities we build with them cover agentic workflows, tool use, structured outputs, retrieval pipelines, evals, guardrails, and MCP servers connecting agents to internal tools. That is a lot of surface area, but the rule underneath it is simple: anything AI writes faster, AI writes. A named engineer owns and checks it. That is where the speed comes from, not from removing the person who is accountable for what merges.

This distinction matters because agent-generated code is not automatically wrong, and it is not automatically right either. It is a draft. The value it adds is in how much of the boilerplate, the repetitive CRUD, the test scaffolding, and the first pass at a migration it takes off an engineer's desk, so that engineer spends the day reviewing and shaping instead of typing from a blank file.

The production review gate that keeps agent output safe

AI-assisted code goes through the same gate as any other code in a client's repository. Concretely, that means:

  1. A pull request in the client's own repo. Nothing merges outside the repository the client owns.
  2. Review by the engineer who owns that surface. Not a rotating reviewer, not an automated approval, a specific named person.
  3. Typed contracts. Function signatures and data shapes are checked, not inferred at runtime.
  4. Tests in CI. The change does not merge if the test suite fails.
  5. Schema changes as reviewed migrations. A database change is a migration file in the PR, reviewed like any other code, never a manual change applied outside version control.

This is the same gate whether a junior engineer, a senior engineer, or an agent wrote the first draft. Our how it works page walks through the same discipline applied to the full delivery cycle, from the first prototype through weekly ship cadence to handover.

Stronger verification patterns for regulated builds

Some builds carry more risk than a typical CRUD app, and the review gate gets stiffer to match. Two patterns from regulated work show what that looks like in practice.

On a HIPAA-grade compounding-pharmacy platform, we built spec-first: each phase of the build was checked against numbered requirements by an adversarial verifier before it was allowed to merge. The verifier's job was to try to find a requirement the code did not actually satisfy, not to confirm what the code already claimed to do. That build shipped eleven epics behind the spec gate, with a seven-year immutable audit log as one of the requirements it had to hold up against. We wrote up the full pattern in building HIPAA-grade pharmacy routing with a seven-year audit log.

On a campaign-intelligence platform handling 25 million voter records and $2.365B in matched federal contributions, new states onboard through a pipeline gated by an automated Playwright verification gate: every page is verified end to end before it goes live, not spot-checked after the fact. That pipeline is described in scoring 25 million voter records. Both patterns exist because a human reviewer alone, working through a large surface area under a deadline, misses things an automated adversarial check or an end-to-end verification pass will not.

Testing evidence, not testing claims

Generic testing claims are cheap. What holds up is the count and the kind of test actually shipped. On the compounding-pharmacy build, that means 490+ unit tests, a strict TypeScript typecheck running clean, and patient-facing screens passing WCAG 2.1 AA, all proven in code rather than asserted in a sales deck. Vitest and Playwright sit in the stack as dedicated testing tools, not incidental libraries: Vitest for unit and integration coverage, Playwright for the end-to-end verification pass described above. That is the standard we hold agent-assisted output to as well, since agent-written tests run through the same CI pipeline as everything else.

Where agents run in production, beyond writing code

Agents also do real work inside shipped products, not just inside the build process that creates them:

  • Voice-to-chart note drafting. On a developmental-dentistry practice network, GPT-4o runs live in production: a provider dictates a note during a visit and a model drafts a structured SOAP entry into the patient chart.
  • Radiograph vision analysis. The same platform runs a vision model reading radiographs as part of an imaging pipeline that also moves CBCT scans through analysis, alongside provider and patient portals built on 80+ REST endpoints.
  • Offline, air-gapped fraud detection. On a public-sector spend auditor, eight fraud detectors run over an ingest-enrich-detect-score pipeline entirely offline, backed by Ollama, making zero external calls. This design was chosen because the accounts-payable data cannot leave the premises. Full detail is in building an air-gapped fraud detection pipeline. The system is designed to surface duplicate payments, contract-splitting, and shell vendors from modeled data, not to claim it has found real fraud on live client accounts.
  • AI-graded compliance scanning. A white-label privacy-compliance platform extracts a site's privacy policy with a headless browser and grades it with GPT-4, returning a status and specific fixes an agency can resell under its own brand. That build is walked through in building a white-label AI compliance scanner.

For more on how a HIPAA-aligned scribe pattern gets built end to end, see how we built a HIPAA-aligned AI scribe, and for the broader landscape of agents working together rather than as a single call, see multi-agent systems, explained.

How we handle model dependency risk

Every one of the production uses above sits behind a single interface abstraction inside the codebase, rather than calling a specific model's API directly from a dozen places. When a model is deprecated or a client wants to switch providers, the change is a config change and a re-test, not a rebuild. That is a deliberate architectural choice, not an accident: it means a platform is not locked to one vendor's roadmap, and everything still ships into the client's own repository and cloud account from week one, with full ownership of the code and no license-back to us. Our security page covers the related controls, including signed BAAs on request and HIPAA-aligned controls (never "HIPAA certified," since HIPAA has no certification to hold).

What the data does not show

No third-party market statistics or competitor benchmarks on AI-agent adoption in software development were available in the source material for this piece, so none are cited here. Everything above is a first-party account of how agents run inside our own client codebases, described with the actual test counts, detector counts, and endpoint counts from shipped work, not an industry-wide survey. If you are comparing that against how a pod is priced and staffed to do this kind of work, the pods page and our buyer's guide to AI agent development services lay out the structure without leaning on numbers we cannot back.

The short version

AI agents in software development speed up the drafting of code, tests, and documentation, and in production they also draft clinical notes, read radiographs, run offline fraud detectors, and grade compliance documents. None of that output ships unchecked: a named engineer reviews it in a pull request against typed contracts and CI tests, and regulated builds add a spec-first verifier or an automated end-to-end gate on top. The model behind any of it sits behind an interface so a deprecation is a config change, not a rewrite.

Frequently asked questions

Do AI agents write production code without a human checking it?
No. Agents draft code faster than a person typing alone, but a named engineer reviews and owns every change before it merges. The change goes through a pull request in the client's own repository, with typed contracts and tests in CI, the same gate any human-written code passes through.
What is a spec-first adversarial verifier?
It is a verification step used on a HIPAA-grade compounding-pharmacy build, where each phase of the code was checked against numbered requirements by a verifier whose job was to find gaps, not confirm success. The build only merged a phase once it passed that check, and it shipped with 490+ unit tests and a strict typecheck as further evidence.
Can AI agents work in an environment where data cannot leave the building?
Yes. On a public-sector spend-audit platform, eight fraud detectors run entirely offline in an air-gapped mode backed by Ollama, making zero external API calls, so sensitive accounts-payable data never leaves the premises.
What happens when the model behind an agent gets deprecated?
The model sits behind a single interface inside the codebase, so swapping it is a config change and a re-test, not a rebuild. This is a deliberate design choice across the production uses described above, meant to avoid locking a client's platform to one model provider's release schedule.

Sources

Get in touch.

Thirty minutes to map your problem to a plan and a timeline. You will leave the call with scope, price, and a start date.

What happens on the call
01You describe the outcome you need.
02We map it to scope, price, and a start date.
03You decide whether to proceed to a free prototype.
Schedule a 30-minute call