How to Choose an AI Coding Agent
Nous Research had put off a cleanup of its Hermes Agent codebase. It had more than one million non-test Python lines; one file alone had 34,847. Pulling engineers away from features for that refactor was hard to justify. So Teknium gave the job to Hermes Agent, which coordinated AI coding agents. In the team’s September 2026 account, it dispatched 1,393 subagents, peaked at 218 running at once, and finished the main run in about 19 active hours. Estimated model cost reached roughly $25,000 including follow-up runs, excluding human review.
Then reviewers found the part that the speed chart could not show: public plugin names removed because no internal caller used them, plus changed exception handling at roughly 65 sites. For a team choosing among Anthropic Claude Code, OpenAI Codex, Cursor Cloud Agents and GitHub Copilot cloud agent, the real test is accepted changes, reviewer time, regressions, access controls and full cost, not generation speed alone. The team fixed those regressions before merge and made further fixes afterward. Parallel agents can make a million-line repository feel small; review becomes the scarce resource. This case used Hermes Agent with Claude Fable 5.1. It was not a test of the four products below.
An AI coding agents comparison should start with three real tasks from your repository, not the prettiest demo function. Give those four products the same briefs. The 12 runs below are a screening exercise, not a statistical benchmark or a pre-announced winner.
First, compare products, not just model names
A coding agent is a working environment: it reads a repository, proposes or makes edits, runs commands, and hands work to a person. A model is one component inside it. Buying a team workflow on a standalone model leaderboard misses permission controls, context setup, branch isolation, test execution, pull-request handling and the cost of review.
OpenAI introduced GPT-6 Sol and GPT-6 Luna for Codex on 22 September; Anthropic introduced Claude Opus 5.5 for Claude Code the same day. The OpenAI announcement lists API rates of $2 input / $10 output per million tokens for Sol and $0.10 / $0.50 for Luna. Anthropic’s announcement lists $4 / $20 for Opus 5.5. Those are dated API token prices, not the all-in price of each coding product, and the products do not all expose identical models or billing rules. A cheaper token does not automatically produce a cheaper accepted pull request. See also: AI Agent Pricing in 2026: What Does an AI Agent Cost?.
Claude Code, Codex, Cursor and Copilot: what actually differs?
| Product surface | What the official product source describes | Where to probe in your own pilot |
|---|---|---|
| Anthropic Claude Code | Works in terminal, IDE and other supported surfaces; can inspect a codebase, edit files, run commands and use project guidance. | Local developer workflow, repository instructions, command approvals and how much setup the team must maintain. |
| OpenAI Codex | Coding work across Codex interfaces including local and delegated workflows. | Whether your task is faster as interactive work or a delegated branch, and how tests and human approvals are surfaced. |
| Cursor Cloud Agents | Runs delegated coding work in isolated cloud environments and can return changes for review. | Cloud environment setup, access to the right dependencies, branch/PR behavior and data boundaries. |
| GitHub Copilot cloud agent | Works from GitHub issues and pull requests in an ephemeral development environment; proposes changes on a branch for review. | Fit with your GitHub permissions, Actions environment, issue-to-PR flow and reviewer controls. |
These rows describe documented product capabilities, not what Powercode Group tested. Claude Code and Codex each have multiple interfaces; Cursor’s cloud agent is not simply its editor autocomplete; Copilot cloud agent is not the same task as in-editor completion. Match the specific surface to the work you expect to delegate.
If you only need faster suggestions at the keyboard, a cloud-agent comparison may be the wrong purchase exercise. If you need autonomous pull requests, an autocomplete demo is the wrong proof.
How do you evaluate AI coding agents? Start with 12 runs
Choose three representative tasks from one safe repository fixture, then try each with the four products. That is 3 × 4 = 12 runs. Keep the task brief, starting commit, permitted files, test commands, acceptance criteria and human reviewer consistent. Record the model and product settings that were actually available; do not pretend every product ran the same model.
- A bounded bug fix. Example: an API retries a failed payment request and can create a duplicate record. The accepted change must prevent the duplicate, preserve the public API and add a regression test.
- A small feature. Example: add a new export filter behind an existing permission check. The accepted change needs a migration or contract update only if the task genuinely requires one.
- A risky refactor. Example: rename an internal module while preserving public plugin names, exception behavior and backwards compatibility. This is the task where fast, broad edits can hide an expensive regression.

The examples are proposed test tasks, not Powercode Group customer results. Use a scrubbed or synthetic repository if production data or secrets would be exposed. If a product cannot support your normal development environment without a major exception, record that as a buying result, not as a reason to weaken the test.
Twelve runs are enough to uncover obvious workflow differences. They are not enough to declare a universal winner. Repeat the two finalists on more tasks and, where outputs vary, rerun a sample with the same starting state.
What should an AI coding agent scorecard measure?
For every run, retain the prompt, initial commit, environment, product/model version, diff, test log, reviewer notes, elapsed time, usage charge and final disposition. Then score the result in this order:
| Gate | Pass question | Why it matters |
|---|---|---|
| Correctness | Do the required tests pass, including the new regression test? | “Code produced” is not “problem solved.” |
| Contract and safety | Are public interfaces, permissions, secrets and error behavior intact? | Broad refactors can pass happy-path tests and still break consumers. |
| Reviewability | Can an engineer explain the diff and the agent’s assumptions? | A 20-minute generation followed by a three-hour review is not a 20-minute task. |
| Acceptance | Was the change merged, revised, rejected or abandoned? | Count accepted work, not merely attempted work. |
| Total cost | What did the run, reruns, setup and human review cost? | A low token rate can lose after failed attempts or heavy review. |
Do not collapse this into a single weighted score on day one. A security failure is not offset by a pretty diff. Set hard gates first, then compare time and cost among runs that passed. Our agent-evaluation guide covers broader test design; the AI code-review guide deals with a neighboring, narrower task. See also: AI Agent Security: A Practical Guide for Engineering Leads.
One useful measure is cost per accepted change: total product usage, setup, reruns and human review divided by accepted changes in the same task cohort. Keep median reviewer minutes and defect escapes beside it. This is a pilot metric, not a benchmark promise. It prevents the absurd outcome where a tool wins because it opened more pull requests that no one can merge. See also: AI Agent ROI: How to Calculate Value, Costs and Payback.
What the 1,393-agent story actually teaches
The Nous report is an extreme example, not a target for most teams. Its authors describe reducing more than a million non-test Python lines and a main run costing about $19,300; the approximately $25,000 total includes follow-up work and excludes human-review cost. They also document fixes after human inspection. The useful lesson is to design the acceptance gate before scaling the agents. Parallelism increases the number of changes available for review; it does not grant review capacity.
For the risky-refactor task, therefore, include at least one backwards-compatibility test on an external interface and one test for failure behavior. Ask a human who maintains the code to review the diff without seeing which tool produced it. A reviewer should be allowed to reject an elegant rewrite that changes a contract.
Do not publish the Nous figures as evidence that Claude Code, Codex, Cursor or Copilot would have done the same job. They were not the tools in that case. Do not cite vendor benchmark percentages as interchangeable with your repository results either.
Which AI coding agent fits your team?
Start with Claude Code if your developers want a terminal- or IDE-centered agent inside an established local workflow and can maintain repository guidance. Start with Codex if you want to compare interactive and delegated coding in OpenAI’s coding environment. Start with Cursor Cloud Agents when isolated cloud delegation and its editor-to-cloud workflow match your team. Start with Copilot cloud agent if issues, permissions, Actions and pull requests already center on GitHub. These are hypotheses from documented surfaces, not product endorsements. Your 12-run results decide.
For a single-tool deep dive, see our Claude Code workflow guide. If your goal is hiring signal rather than coding automation, developer assessment is a different decision and should not be conflated with agent output.
There is also a credible no-buy-yet answer. If the repository has unreliable tests, undocumented public contracts, unrestricted credentials or no reviewer time, an autonomous agent can multiply the uncertainty. First make a representative task reproducible, define access boundaries and assign a reviewer. Then run the screen.
The buyer’s question is not “Which agent writes code fastest?” It is “Which workflow gives our team the most accepted, explainable changes without shifting hidden cost and risk to reviewers?” If you want to design that pilot around your repository and delivery process, Powercode Group’s custom software team can help scope the tasks, acceptance gates and rollout boundary.
Sources and editorial notes
Official product documentation and launch posts are linked beside changing claims. The Nous Research account is the team’s own report, not independent validation or a four-product benchmark. Research cut-off: 23 September 2026; model availability, plan terms, product behavior and quoted prices may change. This article does not claim Powercode Group tested or ranked the four products.