Eigen RadarAI
Analysis

Goodfire watches inside AI agents before escalating suspicious steps

Goodfire has opened a monitoring service through Baseten that checks an AI agent’s internal activity during a task. Small detectors flag steps for deeper review, rather than sending every step to another model. Customers choose which risks to watch and whether an alert prompts logging, human review or refusal. Kimi K3 is the first supported model.

Artificial Intelligence··Night
Two people stand at an unmarked display in a bright workspace, with another worker seated behind them.

Internal signals select steps for deeper review

Goodfire, a company developing tools to interpret AI models, opened its agent-monitoring service through the model-hosting platform Baseten on 8 October. The first supported model is Kimi K3. Small detectors examine activity inside the model while it carries out a task. When a step looks suspicious, the system sends it to another model for a more detailed assessment.[1], [2]

Customers choose both risk categories and responses

Customers can ask the monitors to flag offensive hacking, activity linked to chemical or biological weapons, or reward hacking: attempts to exploit a scoring mechanism instead of completing the assigned task. They also configure the response. An alert can be logged, sent to a person for review or used to refuse an action. Detection and the application’s decision about permission therefore happen at separate stages.[1]

Access begins with one model on one hosting platform

The design reserves heavier external review for selected steps rather than applying it continuously to every output. Goodfire executives Eric Ho and Dan Balsam describe internal model activity as the basis for that selection. The company supplied its own tests and cost comparisons; they are developer measurements. The newly available service covers a specific model and hosting channel. Its release does not establish equivalent protection for every model or every agent deployment.[1]

References

  1. News sourceTechCrunchGoodfire opens monitors that inspect agents’ internal activity↩1↩2↩3
  2. News sourceCrypto BriefingGoodfire launches internal AI monitors through Baseten↩