← Tous les articles

AI Agents · September 28, 2026 · 4 min de lecture

Anthropic Ran 950 AI Agents for 21 Hours. The Interesting Part Is Who Checked Their Work.

Anthropic just ran 950 Claude agents for 21 hours on a real scientific search, and the agents did not make the discovery. Here is what that says about where AI agents actually create value.

Somewhere in most leadership teams there is a quiet question nobody says out loud: if we let an AI agent run for a full day, unsupervised, on a real task, what actually comes back. A finished answer, or a mess that takes longer to check than the task would have taken by hand. That question is not hypothetical anymore. Anthropic just answered it, in public, with real numbers.

On September 25, Anthropic published results from an experiment in which about 950 Claude agents ran for 21 hours, using roughly 210 million tokens, searching a public DNA sequence database for a specific type of enzyme called reverse transcriptase. The agents were not writing an essay or summarizing a meeting. They were filtering. Out of a huge pool of candidate sequences, they narrowed the search down to a shortlist worth a closer look, and flagged one system, an enzyme paired with a repeating DNA sequence, that resembled the building blocks of CRISPR, the gene editing technology that reshaped biology over the last decade. Feng Zhang, the MIT and Broad Institute researcher whose lab helped develop CRISPR based tools, called it an exciting example of how AI agents can contribute to biological discovery.

That is the headline most people will read. Here is the part that matters more for anyone running a business. Anthropic is explicit that the agents did not make the discovery. Human scientists set the search parameters at the start, reviewed every candidate the agents flagged, decided which ones were worth pursuing, and did every single hour of lab work: protein expression, biochemical testing, structural analysis. The agents did the filtering. The humans did the judgment.

Most coverage of this story will treat it as proof that AI agents are now doing science. The more useful reading is narrower. A 21 hour agent run took a search space of roughly 200,000 candidates and turned it into a handful worth a scientist's time. That is not autonomy. That is volume reduction, at a scale and speed no team of humans could match on their own, followed by exactly the same verification step a rigorous lab would always require. The agents changed what got put in front of a human. They did not change who decided what mattered.

That distinction is the entire adoption question for most SMEs and leadership teams right now, whether the task is DNA sequences, inbound leads, support tickets, or CVs.

Four practical implications follow from this.

  1. Look for the filtering step, not the final answer. The highest value AI use case in most businesses is not "write the report," it is "cut the pile down to the part worth a human's time." Sales lead lists, applicant pools, support ticket backlogs, contract reviews: these are filtering problems before they are judgment problems.
  2. Verification capacity is your real constraint, not subscription cost. Anthropic could run 950 agents because it had scientists ready to review every shortlist item before anything moved forward. A team that deploys an agent to filter and has no one dedicated to checking the shortlist has only built half a system. Budget the review time before you budget the tool.
  3. The task needs a checkable output. The agents were looking for something with a defined shape: a gene plus a repeating sequence. Open ended tasks with no clear pass or fail line are much harder to filter reliably, and much easier to trust wrongly. In MAKIA terms, this is a Knowledge question before it is an Impact one: what a correct answer looks like has to be defined before AI touches the volume. Before handing a process to an agent, ask what a correct shortlist item actually looks like, in writing.
  4. This is what disciplined agent use looks like, and it is rare. Most companies experimenting with AI agents skip the part where someone owns verification. Anthropic's experiment is notable less for the size of the agent swarm and more for the fact that every output had a named human decision point before anything counted as real.

None of this requires a research budget or a biology lab. It requires knowing, for one process in your business, exactly where the filtering ends and the judgment begins, and making sure a person is still standing at that second point.

Try this week

Pick one recurring task where your team manually filters a large volume down to a short list: inbound leads, resumes, support tickets, invoices to review. Time how long the filtering itself takes this week, separately from the decision making that happens once the shortlist exists. Write down both numbers. That ratio is where an AI agent could help you this quarter, and where it should not replace anyone.

Sources

  • Claude discovers a novel enzyme system, Anthropic
Partager cet articleLinkedInFil RSS