The Tiny Model That Embarrassed Frontier AI Labs

A compact glowing neural core outperforming large server racks, representing a 27B parameter AI agent beating frontier models
KEY POINTS
  • London lab Inherent released Faraday on August 22, 2026, an agent that reproduces the findings of published scientific papers without being told the answer in advance.
  • Faraday runs on Qwen 3.6, a 27-billion-parameter model, yet the company says it outperformed Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5 on the replication task.
  • Inherent was founded by Google DeepMind alumni and emerged from stealth in May 2026 with a $50 million seed round; it has twelve employees and plans to reach about 20 to 25 by year end.
  • Rather than building its own coding tool, Faraday calls OpenAI’s GPT-5.5 Codex, the way a scientist uses existing software instead of writing everything from scratch.

Twenty-seven billion parameters. That is the entire size of the model behind Faraday, the research agent a twelve-person London startup released on August 22, 2026 – and the number that makes the rest of the story awkward for everyone else. Inherent, founded by Google DeepMind alumni, says Faraday beat Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5 at independently replicating the results of published scientific papers. Both of those are frontier-scale systems. Faraday is not. Cofounder and chief scientist Edward Hughes told TechCrunch the result itself was not even the interesting part.

What Faraday Actually Does

Replication, not recall

The task Faraday was measured on is narrow and unforgiving: take a published scientific paper, and independently reproduce its findings without being handed the answer in advance. This is not summarization and it is not retrieval. The agent has to reconstruct the experimental setup, run it, and arrive at the same conclusion on its own. If it cheats by recalling the paper’s stated result, the exercise is worthless.

Inherent chose the benchmark deliberately, because it mirrors how human researchers are trained. “Many PhD students actually start by doing this,” Hughes said. Replication is the apprenticeship of science – the step where a researcher learns not just what the field knows, but how the field came to know it.

The parameter gap

Faraday is built on Qwen 3.6, a model with 27 billion parameters. Parameter count is a rough proxy for a model’s size and, typically, its training cost. Claude Opus 4.8 and GPT-5.5 sit in an entirely different weight class. The gap is the whole point: if a small open-weight base model plus the right training loop can beat frontier systems on a hard reasoning task, then raw scale is not the only lever that matters.

Trend Insight — The interesting claim here is not that a small model won. It is that a task-specific training loop closed a gap that hundreds of billions of parameters were supposed to guarantee. For any company evaluating build-versus-buy on AI, that reframes the question: the moat may be in the training environment, not the base model.


Teaching a Model to Have Taste

Reinforcement learning over rules

Inherent’s bar for success went beyond accuracy. The team wanted Faraday to demonstrate what Hughes calls “research taste” – an instinct for which experiments are worth running and how to design them well. Taste is not the kind of thing you can write down as a rule set, which is why the company leaned on reinforcement learning, a training method that rewards a system for good outcomes rather than specifying the steps to reach them.

Critically, Inherent did not train its agents primarily on the study of how science is conducted. It bet that a reward-based approach would generalize better toward the longer-term target: agents capable of contributing across many scientific fields, not just reproducing results in one. “We’re always guided by that north star of building an AI scientist agent and imbuing our agents with taste,” Hughes said.

Buy the tools, build the scientist

That focus also shaped what Inherent refused to build. Instead of developing its own coding tool, the team had Faraday use OpenAI’s GPT-5.5 Codex – the same way working scientists rely on existing software rather than writing every utility themselves. A competitor’s product sits inside the stack of an agent that then beat that competitor. It is a pointed architectural choice.

Trend Insight — Inherent is drawing a line between the layer it wants to own and the layer it is happy to rent. Owning the reasoning loop while renting the coding tool is the same calculus most enterprises face when they decide which parts of their AI stack are actually differentiating.


Twelve People in King’s Cross

The anti-sycophancy design goal

Inherent is also trying to avoid the failure mode where an agent simply tells the user what they want to hear. Hughes described the target behavior in terms of his favorite kind of human colleague – the one who comes back and says: “I got curious about this, and I went off and I did these experiments. What do you think of these results?” That is an agent that initiates, disagrees, and brings back evidence, rather than one that ratifies whatever the operator already believed.

A small team and a hiring window

All twelve Inherent employees work in person from an office in King’s Cross, the London neighborhood that Google DeepMind’s presence helped turn into a major AI hub. The company emerged from stealth in May 2026 with a $50 million seed round and plans to grow headcount to “about 20 to 25” by the end of the year.

Hughes has publicly called for the end of “garden leave” – the British practice of barring departing employees from joining or founding a rival for months after they resign, a restriction American researchers generally do not face. “This is a personal view rather than a company view, but I was affected by the garden leave problem,” he told TechCrunch. With Demis Hassabis’s new role leaving some DeepMind staff unsettled, Inherent’s hiring push arrives at a convenient moment.

Trend Insight — Worth keeping the caveat in view: this is a vendor-reported result on a self-selected task, not an independent evaluation. The finding to watch is whether the same training recipe holds up when someone else runs it, and whether replication skill actually transfers to discovering things nobody has published yet.


Related

Sources

  1. TechCrunch – Inherent, founded by DeepMind alumni, says its AI ‘teammate’ just outperformed Anthropic and OpenAI at replicating research (Aug 22, 2026)
  2. Inherent Labs – Training to Replicate (Faraday research page)
  3. Anthropic Newsroom

AI Biz Insider · AI Trends EN · aibizinsider.com


AI Biz Insider에서 더 알아보기

구독을 신청하면 최신 게시물을 이메일로 받아볼 수 있습니다.

코멘트

댓글 남기기

AI Biz Insider에서 더 알아보기

지금 구독하여 계속 읽고 전체 아카이브에 액세스하세요.

계속 읽기

AI Biz Insider에서 더 알아보기

지금 구독하여 계속 읽고 전체 아카이브에 액세스하세요.

계속 읽기