From Black Box to Glass Box: AI Governance, Evaluations, and Controls

  • Brief

September 17, 2026

We can read agent’s minds, but they can’t read ours

Recent headlines are anthropomorphizing AI agents, indicating they could turn on us like a military traitor or a rogue employee. But there is a profound difference between an AI agent and a human: you can read an agent’s mind and follow its steps and actions as they happen.

Agents tell you their plan and provide you with step-by-step updates on their actions, roadblocks, and decisions. The problem isn’t a lack of data, it’s a lack of systematic monitoring and scoring of these traces. On the flip side, agents can’t read human minds. They are frequently given incomplete instructions or unfettered reward functions. Coders complain about AI slop, referring to illogical code generated by agents. But there is also human communication slop, which is a major reason for unexpected agent actions. A fully developed evaluation and benchmarking system can help here.

For compliance officers, CROs, CTOs, and risk leaders in financial services, such a system would turn an agent’s telemetry into an air traffic control system that tracks deployments, process adherence, identifies gaps, and tests agent actions against regulatory frameworks. For users and data scientists, evaluations and benchmarks would point out weaknesses in instructions, irrelevant steps, and unclear prompting.

The real risk isn't autonomy, it's addressing the blind spot.

Clearly, agents need guardrails. But there is a downside to placing too many constraints on agents. Agents have “agency,” which is what separates them from software. Too many or the wrong guardrails would constrain the upside of agents. A carefully measured amount of agency can unleash great returns. We’ve developed agents that can run reconciliations, identify root causes, triage exceptions, and flag anomalies at scale. Letting agents learn from their past decisions, while leveraging deep domain skills in complex asset classes and documents, has moved time savings from around 25 percent to over 70 percent.

The glass box: continuous transparency and evaluation

The principle underlying our approach to agent evaluations is simple. Every action an agent takes is visible, attributable, and defensible to a regulator. Complex, domain-specific work is measured for accuracy, speed, and cost. Scoring predictions (e.g., calling the right root cause) gives feedback to agents that helps them improve. Team members see where agents are strong and weak, enabling them to add human oversight in the right places.

There are several features that separate governed agents from unmonitored ones:

  • Domain-specific task evaluations. Agents are tested first in sandboxed simulations for their ability to perform work, read documents, and take actions with carefully orchestrated evaluations.
  • Root cause tracking. You can see the exceptions, root causes, and track the percent of exceptions that are auto-closed.
  • Centralized oversight dashboards. Your analysts see what every agent is doing, in real-time, and in one place.
  • Performance analytics per exception type. Accuracy, resolution time, and basis-point impact tracked for each type of work.

This is the difference between "our agents are running" and "we know what our agents are doing." One is a status update. The other is a control framework.

The proof points of an evaluation and control system

Nothing compares to the experience of running and fine-tuning agents in daily production. This is because agents confront many more real-world issues in production. Our control system is borne from nearly a decade of running daily production AI systems. Here are some key aspects of our production experience:

  • 120+ agents in daily production, across $13T in AUM coverage. These agents run live inside the operations of some of the world's largest and most complex asset managers and fund administrators.
  • North of 90% root cause accuracy, with every decision logged.
  • 1,000+ exceptions triaged daily by agents, with full compliance tracking on each.
  • ISO 42001 certified. Governance built into the architecture, not added under regulatory pressure.
  • Agents learning on day one in production, improving accuracy daily on live data while remaining fully traceable.

Consider a $1T+ global asset manager. Reconciliations ran manually across legacy systems, generating high volumes of false positives and offering almost no visibility into root causes. After moving to OnCorps, analysts oversaw the work through dashboards instead of chasing breaks by hand. More than 30,000 breaks were processed and reviewed, agents identified accurate root causes 92% of the time, and the client captured 80%+ efficiency gains.

You can’t manage what you can’t measure

Having full confidence in agents is a high bar. It doesn't mean the agent works most of the time. It means you can defend every action it takes to your board, your auditors, and your regulator. Ron Allen, OnCorps’ CEO, defines the standard plainly: "Production-ready AI means clients can allow agents to take action on their behalf with full confidence."

Measurement and continuous feedback bring much more than safety. Tech forward teams have already used this approach to power-up agentic coding. For example, Claude Code is the fastest growing software in history. It owes its success to benchmarks and evals. Before its release, LLMs like Opus and Sonnet could code, but with very mixed results. This “single-turn” way of problem solving is now pervasive among teams trying to solve financial operations problems, like reading financial documents. (Single-turn means the LLM only tries once on a given task.) Humans must try new things repeatedly to get things right. They learn from experience. Benchmarks and evals provide agents a procedural learning loop of observing, acting, measuring, and retrying. This has allowed our team to develop and run complex procedures while building expertise. Benchmarks, coupled with a custom harness, allow agents to learn from experience.

A foundational AI governance and scoring system can do for financial processes what Claude Code (built on a benchmark and eval framework) did for coding. Our evaluations are built on a foundation of standard operating procedures and document descriptions by asset class. This allows us to scale new use cases much more rapidly, since the ability to change parameters and measure them quickly in simulation is a part of the platform.

The takeaway

As firms race to deploy agents, very few teams have invested in evaluation-based support models. “Evals” and agent observability are primarily used by data scientists pre-production.

OnCorps built a "glass box" architecture so the claim "we know what our agents are doing" holds up under scrutiny. Since benchmarks and evaluations are critical to learning new processes, faster scale-up is also a benefit to a glass box architecture. This type of system may be standard in several years, but for many firms, several years is too late. It is possible to do this now and we encourage leaders to investigate its benefits.

Let your team do their best work.

Give operational teams the leverage to do their best work.

Please rotate your device