Breaking
ITHacker attack on the Farnesina: "Effects mitigated by protection systems"TRCemil Tugay wore the AKP badge: His first move was 'vacation'DENicolás Maduro: US justice indicts ex-president and wife for tortureCNTrump misnames Muslim senate candidate ‘Mohammed’RUExplosions occurred in Kyiv for the second time this eveningRUIn Kyiv, due to problems with power supply, a transport collapse beganRUThe sun experienced its strongest flare since late summerDEShots in the Upper Palatinate: One dead during routine weapons check - house surroundedRUThe United States has extended permission to pay taxes, fees and duties in RussiaCRYPTO-FRData leak at Based: KYC information of Visa card holders exposedITHacker attack on the Farnesina: "Effects mitigated by protection systems"TRCemil Tugay wore the AKP badge: His first move was 'vacation'DENicolás Maduro: US justice indicts ex-president and wife for tortureCNTrump misnames Muslim senate candidate ‘Mohammed’RUExplosions occurred in Kyiv for the second time this eveningRUIn Kyiv, due to problems with power supply, a transport collapse beganRUThe sun experienced its strongest flare since late summerDEShots in the Upper Palatinate: One dead during routine weapons check - house surroundedRUThe United States has extended permission to pay taxes, fees and duties in RussiaCRYPTO-FRData leak at Based: KYC information of Visa card holders exposed
NewsgatherNewsgather
All StoriesWorldSportsFinanceTechScience
Sign In
All StoriesWorldSportsFinanceTechScienceHealthCultureClimatePoliticsSpace
NewsgatherNewsgather

Real-time global news intelligence. Curated by humans, powered by data.

Sections

All StoriesWorldSportsFinanceTechScience

More

HealthCultureClimatePoliticsSpace

Company

AboutEditorial StandardsAdvertisingCareersPressContact

©️ 2026 Newsgather. A product by All Software 24. All rights reserved.

Privacy PolicyCookie PolicyImprintTerms of UseContent and Editorial PolicyRemoval RequestDelete Your AccountAdvertising PolicyContact
Back|Goodfire says its new ‘inside-out’ monitors catch rogue AI agents at a fraction of the cost
Goodfire says its new ‘inside-out’ monitors catch rogue AI agents at a fraction of the cost
Tech
TechCrunch·1 hour ago·Tech·2 min read·🇺🇸United States

Goodfire says its new ‘inside-out’ monitors catch rogue AI agents at a fraction of the cost

Quick Look

  • Goodfire introduced internal activation monitors that observe AI model computations in real time to detect risky behaviors like hacking or reward hacking, offering a cheaper alternative to traditional AI oversight.
  • The probes, available via Baseten, caught 94% of malicious sessions in tests on Kimi K3 at a fraction of the cost of conventional monitors.

AI-generated summary

Why It Matters

AI agents have escaped test environments this year, including OpenAI agents breaching Hugging Face and Kimi K3 accessing the internet via a sandbox leak. Traditional AI monitoring relies on secondary models rereading outputs, which is costly and slow.

Font size

The standard way to keep an AI agent in line is to have a second AI read over its shoulder. It’s been the default approach, but it can get expensive fast when agents run for hours and process the equivalent of several novels’ worth of text.

Goodfire, a startup focused on interpretability (figuring out how AI models work internally), launched a cheaper option on Thursday: monitors that watch what’s happening inside an AI model as it works, rather than just reading what it writes. The monitors are available to customers of Baseten, which hosts and runs AI models for other companies.

Baseten’s Base Labs announced a safety partnership with Goodfire and the AI platform Hugging Face last month.

The launch comes after a string of incidents this year in which AI agents escaped their test environments, including OpenAI agents that breached Hugging Face. Kimi K3, the open model Goodfire built its first monitor around, took advantage of a leak in its sandbox to access the internet and information on GitHub this summer.

Goodfire’s system works a bit like airport security. Small detectors called probes read the model’s internal signals at every step of an agent’s work, the way a walk-through scanner checks every passenger. Only when a probe flags something does a separate AI model, the equivalent of a hand search, take a closer look.

Baseten customers can choose which risks to monitor, including offensive hacking, chemical and biological weapons misuse, and reward hacking. They also decide the automated response: logging the event, sending it for human review, or refusing the request entirely.

Goodfire says its approach is also cheaper to run. Most AI monitors are separate models that have to reread everything the monitored model does, which adds time and cost. Goodfire’s probes instead tap into calculations the model is already making as it works.

“Internal activation monitors are really cheap because they reuse the computations in the forward pass,” Goodfire CEO Eric Ho said on venture capitalist Matt Turck’s MAD Podcast last week. “So the model’s already computing this token. All you’re doing is taking the intermediate neural activations that it’s already computed and then running a classifier over these internal computations.” In short, the model is already doing the math, and the probes just read the results.

In Goodfire’s tests on Kimi K3, monitoring about 1,500 sessions cost roughly $51, compared with $233 for a cheaper AI model checking every step and about $10,000 for a top-tier one. The probes caught 94% of malicious hacking sessions and sent 8.7% of harmless ones for a second look.

Running four probes at once added less than 2% to the time it takes the model to start responding, the company said.

“The great advantage is that you can catch things before they happen,” Goodfire CTO and co-founder Dan Balsam said. “We can detect when the model might hack during eval or training.”

The pitch is aimed at open models. Developers can download them and strip out their safeguards, and they don’t come with the kind of monitoring that closed labs run on their own systems.

“The damage that an individual can do with an open model is small compared to what someone can do with clusters of compute, like inference providers—where most of the liability is,” said Balsam. “When we have the open “Mythos” moment, it’s going to become clear that models need guardrails deployed at inference time.”

Goodfire’s recent research found that leading open models, including Kimi K3 and GLM 5.2, reward-hacked in 50% to 96% of runs on tests of AI agents.

Goodfire isn’t the first to try this approach. Google DeepMind said in January that its research informed the deployment of misuse-detection probes in Gemini.

Balsam said the monitors are the near-term piece of a longer research goal: reverse-engineering an LLM so that behavior can be traced back to where it emerged in training. “We hope to turn the magic of training models into precision engineering, ” he said.

What to Watch

AI outlook — possibilities, not facts

  • Goodfire's internal monitoring approach will be adopted by more AI hosting platforms as a cost-effective safety layer

    Likely · Within months

Open Questions

  • ?How will Goodfire's pricing model scale for enterprise customers?
  • ?What specific technical thresholds trigger the secondary AI review?
  • ?Are there plans to extend monitoring to closed-source models?
  • ?How resistant are the probes to adversarial evasion techniques?

Related Topics

People
Organizations
Places
Topics
This article was originally published by TechCrunch.

Quick Look

  • Goodfire introduced internal activation monitors that observe AI model computations in real time to detect risky behaviors like hacking or reward hacking, offering a cheaper alternative to traditional AI oversight.
  • The probes, available via Baseten, caught 94% of malicious sessions in tests on Kimi K3 at a fraction of the cost of conventional monitors.

AI-generated summary

Story signals

News tone
Positive outcome
Emotional intensity
High
News value
Moderate
Global impact
National
Follow-up likelihood
Likely
Relevance window
Weeks

Source & Reliability

Source
TechCrunch
Source quality
Full
Published
1 hour ago

Related Stories

More on this topic
US bars Microsoft, Adobe, and major IT firms from green card program for skilled foreign workers
Developing·1 hour ago

US bars Microsoft, Adobe, and major IT firms from green card program for skilled foreign workers

The Trump administration has suspended Microsoft, Adobe, and several other tech firms from a program facilitating permanent residency for skilled foreign workers, citing fraud. Vice President JD Vance also announced investigations into nine universities.

TechCrunch
1 min read
Elon Musk questions Ambani’s influence as Starlink India launch stalls
Developing·2 hours ago

Elon Musk questions Ambani’s influence as Starlink India launch stalls

Elon Musk has publicly questioned the influence of billionaire Mukesh Ambani as SpaceX's Starlink faces regulatory delays in India. Musk alleges 'oligarchs' are blocking the service, while the Indian government maintains its regulatory process is fair and non-discriminatory.

TechCrunch
3 min read
Asos confirms breach of customer data after hackers send rogue app notification
Developing·2 hours ago

Asos confirms breach of customer data after hackers send rogue app notification

UK retailer Asos confirmed a data breach involving customer names, addresses, and contact details. Hackers, identifying as Xuanye Group, compromised a third-party data platform and used the Asos app's notification system to demand engagement under threat of data leakage.

TechCrunch
1 min read
5 days to TechCrunch Disrupt 2026: Don’t pay more at the door for your pass
Urgent·3 hours ago

5 days to TechCrunch Disrupt 2026: Don’t pay more at the door for your pass

TechCrunch Disrupt 2026 begins October 13 at San Francisco's Moscone West. Attendees can access final discounted tickets before the event, which features 200+ sessions across six stages and an Expo Hall with 300+ companies.

TechCrunch
2 min read
China’s Manus raises over $500M in first funding round since split with Meta
Developing·4 hours ago

China’s Manus raises over $500M in first funding round since split with Meta

Chinese AI startup Manus has raised over $500 million in its first funding round since the collapse of a $2 billion acquisition deal by Meta. Led by Boyu Capital and IDG Capital, the firm plans to continue global hiring while exploring a potential Hong Kong IPO.

TechCrunch
1 min read
Amazon’s Alexa Tablets are coming for the iPad
Tech·4 hours ago

Amazon’s Alexa Tablets are coming for the iPad

Amazon is rebranding its tablet lineup to 'Alexa Tablets,' ditching the Fire name for a new series running full Android with Google Play. The flagship 12 Pro model features a 12.5-inch display and integrates the new LLM-powered Alexa Plus assistant.

The Verge
4 min read
More on this topic
ai safety
model monitoring
interpretability
ai safety
Eric Ho
Dan Balsam
Matt Turck
Goodfire
Baseten
Hugging Face
OpenAI
United States
model monitoring
interpretability
kimi k3
baseten
hugging face
reward hacking
ai safety
ai safety