Judgment Labs is an applied-research company focused on the continuous-improvement infrastructure for AI agents. The company builds Agent Behavior Monitoring (ABM) technology designed to help AI-native teams analyze production data and identify behavioral anomalies - such as instruction drifts and context retrieval loss - in deployed agents at scale.
The company's product suite includes Agent Search for behavioral-level trajectory querying, Agent Judge for trajectory-level evaluation, Behavior Discovery for surfacing failure modes from unlabeled production data, and AutoRubrics for automatically constructing evaluation rubrics from verifiable signals. Together, these tools form a stack aimed at turning production telemetry into actionable improvements for agent reliability.
Judgment Labs has raised $32 million across combined Seed and Series A funding. The company operates as an applied-research lab targeting last-mile agent reliability, serving AI-native teams and applied-research organisations.






