Crosby Journal

Crosby Engineering

How we do evals at Crosby

Our goal at Crosby Where the billable hour goes to die. is to simulate complex legal reasoning.

We are an AI company with our own law firm. We believe agents should do as much legal work as possible. Our lawyers and engineers sit side by side building agents to deliver completed legal work in hours instead of weeks.

As it turns out, contracts are an unusually good place to start.

They are high volume, common across a wide range of industries, and captured in a written artifact. On the surface, a lot of the work is grounded in concrete context: the contract itself, the client's preferred positions, the negotiation's history, and commercial risk. But contracts are deceptively complex and the nuance in weighing all of these factors cannot be derived from the documents alone.

Two lawyers can look at the same information and produce a different set of valid Redline A marked-up copy of a contract showing proposed changes, with insertions underlined and deletions struck through. To "redline" a contract is to propose edits this way.. In one experiment, we had two skilled attorneys create ideal redlines (golden set) for the same set of 20 NDA (Non-Disclosure Agreement) A contract in which one or both parties agree to keep information they share with each other confidential.. Out of 500 total edits, they agreed on only 45% of them. That makes contract review a fascinating puzzle that is simultaneously measurable and deeply subjective.

Unlike traditional software engineering, we don't have a compiler or unit tests that can decisively tell us whether an AI-generated review is actually something a human would have also sent out to a client. Transforming this non-verifiable problem into something that we can verify at scale has become a core engineering goal for us at Crosby.

To solve this, we've followed three major guiding principles that have allowed us to build a robust eval framework.

Every change we make today at Crosby runs through this system. It has let us ship hundreds of agentic features with high confidence and improved our in-house lawyer efficiency up to 5x.

Verifying deterministically wherever possible

While building our evals, we asked many of the same questions lawyers ask:

Is this redline too aggressive? Does it preserve the Commercial intent The business deal the parties are actually trying to make. A good redline protects the client without undermining it.? Would this client accept the risk? Is this wording better than the alternative?

These questions cannot be readily broken down for a "compiler" and require human judgment. But we found that a surprisingly large portion of this evaluation can be broken down into discrete units of work.

For example:

  • Did the edit land on the paragraph the agent was targeting?
  • Did it make the most surgical edit?
  • Is the comment attached to the actual text it refers to?
  • Did it introduce an undefined term?

Two redlines can express the same legal position but one can be a clean, three-word edit while the other rewrites an entire paragraph, breaks a Defined term A word or phrase given a specific meaning in the contract, usually capitalized (e.g. "Confidential Information"). Misusing one, or using one that's never defined, can change what a clause means., and adds a comment to the wrong sentence. Calling both outputs "semantically equivalent" is not good enough. The second redline is not something a lawyer would trust.

Full rewrite

The Receiving Party shall protect the Confidential Information using reasonable efforts. The Receiving Party shall notify the Disclosing Party of any unauthorized disclosure. Recipient must inform the Discloser in writing no later than ten (10) days after discovering any unauthorized disclosure.

Surgical edit

The Receiving Party shall protect the Confidential Information using reasonable efforts. The Receiving Party shall notify the Disclosing Party of any unauthorized disclosure within ten days.

Deterministic checks are the first line in our evaluation system that establishes that the output is structurally valid before we attempt to evaluate its legal quality.

When something breaks, we know exactly which requirement failed. The same output produces the same result every time and we have a fast, cheap feedback loop to immediately detect whether our agents have reintroduced errors we already know how to catch. One recent fix to our agent's comment logic that looked clean in review quietly introduced 40 malformed comments, and our checks caught it before it shipped. This gives us a reliable foundation at scale and lets us focus our human attorneys' time and attention on the questions that require more nuanced legal judgment.

Evaluating judgment through real comparisons

Once we got past structural correctness, we still needed a way to measure legal judgment.

A common approach here is making golden datasets (i.e. ideal model outputs alongside a set of rubrics for evaluating those outputs). However, golden datasets are incredibly difficult to build for contract review. Good redlines depend on a large set of factors such as commercial intent, client preference, Counterparty The other side of a contract or negotiation: the party across the table from the client. leverage, deal urgency, tone etc. Choosing when and why to apply each variable matters across every different client situation and translating that into a detailed enough rubric for an LLM is an enormous exercise in itself.

We found a more practical starting point by looking at how our lawyers already evaluate legal work today.

It is difficult to have lawyers consistently agree on an answer to: How good is this redline on a scale from 1 to 10?

It is much easier to answer: Which of these two redlines would you rather send to the client?

Pairwise comparisons reduce the calibration problem and are much closer to how lawyers naturally evaluate work. Instead of inventing an absolute definition of quality, they allow us to capture the nuance that comes up with legal work inherently.

Internally, we run labeling exercises that distribute hundreds of comparisons per reviewer, comparing human to human, model to human, and model to model. Collecting pairwise judgments is much faster and easier across a large number of examples while also reducing the chance of noise from smaller sample sets.

While reviewing, our lawyers don't see who or what produced each output and we randomize the order of every comparison. They choose the redline they prefer and separately indicate whether they would actually send the version as written to the client. Just preferring one redline over another doesn't establish that the redline is good enough to send back to our client and might still require an additional human touch.

Our core automation workflow already gives us a starting point for these comparisons. At Crosby, we have a natural data engine. Every day, our AI agents produce a contract review, a licensed attorney evaluates and modifies it where necessary, and the final output is sent back to the client. This gives us both the agent's original proposal and the lawyer-reviewed version sent to the client.

These pairs naturally give us concrete examples of what our agents could be doing better and are a rich dataset for pairwise comparisons. Can our judges recognize why the lawyer-reviewed version is better than the agent's original proposal?

Importantly, not every difference reflects an agent failure in the same way. Sometimes the agent had the right information but failed to synthesize it. Other times, the lawyer is drawing on context the agent never had access to in the first place like verbal preferences from a phone call. Distinguishing between these kinds of gaps matters to know where we need to invest next.

Pairs where we already know the winner also let us test the judges themselves. A stronger model should generally beat a much weaker model. A run with the correct Playbook A client's written guide to its preferred contract positions: what it will accept, what it won't, and the fallbacks it's willing to offer in a negotiation. should beat one with the playbook removed or with the wrong client's playbook. These comparisons don't prove that our judges are sensitive to the subtle differences that matter most, but they do establish a baseline: if a judge can't get these obvious cases right, it isn't ready to evaluate closer calls.

Building a shared system around the actual work

Coming up with the right method of evaluation is only half of the problem. Deciding how we need to actually design the evaluation system itself is another battle entirely.

Two things became clear early on at Crosby.

  • Our strongest evaluators were our licensed attorneys. As engineers, we were able to ramp up on contract reviews and even took a stab at defining rubrics ourselves, but nothing could substitute for a lawyer's judgment on what constitutes good legal work.
  • Any evals platform therefore had to serve both engineers and lawyers alike.

For lawyers, this meant meeting them where they already work. If our eval surface looked like a foreign tool, our lawyers would spend cognitive time wrestling with the interface instead of focusing on substantive evaluation. So we designed our labeling workbench to feel like they were still in Microsoft Word.

The workbench renders the actual docx file, and places every agent-generated redline beside the expected lawyer edit. We also focused on providing knobs for the pieces that the lawyer actually would interact with during a review, like any relevant client context, while keeping the system prompt and tooling isolated to our engineering layer. This allows our attorneys to become context engineers themselves, iterating on inputs into the system and immediately seeing how that translates to output fidelity in an isolated sandbox.

This creates a very powerful feedback loop. Lawyers define what good legal work looks like, automated judges apply that judgment at scale to assess agentic legal work, and then engineers can measure and improve these systems before rolling them out at greater scale.

After incorporating our lawyers into the eval process, we saw our lawyers complete reviews up to 5x faster. In controlled time trials, review time fell by 74% for one client's NDAs and around 80% for another's MSA (Master Services Agreement) A framework contract that sets the standard terms (payment, liability, IP and so on) for all future work between two parties. Project specifics are added later in statements of work.. Our gains largely came from our lawyers quickly pointing out gaps and creating an incredibly fast iteration loop for any change. The moment any change was being tested or rolled out, we immediately had a way to gather lawyer feedback at scale and pair with them on engineering stronger agents. The end result was our agent's first pass rapidly began requiring far fewer corrections.

Client 1

Avg. NDA Review Time

Down 74%

Client 1: Avg. NDA Review Time
Period Average (minutes)
Baseline 17.9
Trial 1 6.5
Trial 2 8
Trial 3 4.7

Client 2

Avg. MSA Review Time

Down 80%

Client 2: Avg. MSA Review Time
Period Average (minutes)
Baseline 91
Trial 1 36.8
Trial 2 33.8
Trial 3 18
Measured in controlled time trials over two months, across ~1,000 edits in 30 contracts. Baseline: lawyers using our agent before any tuning

At Crosby, being an AI law firm gives us a very unusual vantage point: we participate in the very work we are trying to automate. Not just simulated scenarios but actual legal work for actual clients, where mistakes have very real consequences. For every review, we see what our agents proposed, what the lawyer changed, and what ultimately went back to the client. Building this flywheel into our evaluation system is crucial to ensuring every review can help us improve the next one.

These lessons extend beyond just legal. Any firm delivering completed work through AI agents faces the problem of teaching agents to take the first pass at high-quality services work, and creating evals for potentially non-verifiable outputs. To start, a problem does not need to be fully verifiable for parts of it to be measured. Begin by verifying the parts you can check exactly and deterministically. For everything else, you don't need a perfect definition of your output to get started.

Starting with what human experts are already familiar with, such as pairwise comparisons, inherently captures the nuance of human judgment and can be a powerful way to test whether automated judges can recognize those same preferences. Putting this all together into artifacts and workflows where your experts already work only tightens this loop and has allowed us to turn every decision into concrete improvements at scale.

We still have a lot left to learn about how to both simulate and evaluate legal reasoning.

If you find any of these problems interesting, we're hiring at crosby.ai/careers 

Acknowledgements: Thank you to Ariba Khan, Raymond Lin, and Ryan Tanenholz for their work on the technical implementation, and to Tricia Matibag and Ross Weiser for shaping what good legal work looks like in our evals.

Crosby Engineering Back to Journal