← News

Ivo open-sources Sage, a legal model post-trained for long-horizon contract work

Ivo said at its Inscribe customer event that it released Ivo Sage — described as the first free open-source model from a legal AI company post-trained for long-horizon contract work — so attorneys, researchers, and developers can download, run, and fine-tune it, while also previewing the Ivo-micro1 Contract Bench.

Contract AI has mostly been closed software you rent, graded on the edits a model makes. By open-sourcing a model trained for long contract jobs, and previewing a test that scores when to leave language alone and when to escalate, Ivo is betting the hard part is judgment and the context of the deal.

On Thursday, 1 October 2026, Ivo said it released Ivo Sage, a free open-source model post-trained for long-horizon contract work. The GlobeNewswire page is titled “Ivo Becomes The First Legal AI Company to Publish a Free Open-Source Model Post-Trained for Long-Horizon Contract Work.” The page stamps October 01, 2026, 09:00 ET. The dateline is San Francisco. The source line is Ivo. The release says the announcement was made at Ivo Inscribe, which it calls the company’s first user conference and its first customer event for in-house legal and contracting teams. Long-horizon, in this release, means contract work that runs across many steps, rather than one short reply. Post-trained means the model was trained further, on top of a model that already existed. Open-source, here, means a person can download the model and build on it. Those lines are the wire’s.

What Ivo says is new about the model. The release says Ivo became the first legal AI company to release an open-source model post-trained for long-horizon contract work. A line under the headline says that, unlike other legal AI companies, Ivo Sage is the first model attorneys, researchers, and developers can download, run, and build on for free. That “first” is Ivo’s wording in the release. Free, in that sentence, is the price of the model. The release says Sage gives developers, researchers, and legal teams technology they can test, train, and build on. It says legal work looks different from one organization to the next, from a Fortune 500 company to a small business with a narrower set of agreements. Opening the model, Ivo says, lets practitioners, researchers, and developers fine-tune it on their own data, adjust its rules and safeguards for a specific use, and try new ideas quickly. Fine-tune means train it further on examples you choose. Ivo says it is sharing the work so others can test it, check it, and build on it. Those lines are Ivo’s.

Min-Kyu Jung, co-founder and chief executive, is quoted on the release. “The next leap for AI in contract and legal work won't come from bigger models. It will come from giving models the context of the work: the documents, the playbooks and the way a legal team operates,” he said. “That's the hard, unsolved part of legal AI, and it's where we're focused on next. We're opening up our model so others can test it, adapt it and build on it. That's how we move closer to contract intelligence the whole industry can rely on.” A playbook, in that sentence, is a legal team’s written rules for how a contract should be marked up. Context, here, is the documents and the way that team works. That quotation is his, in the release.

How the model was built, as the release states it. Ivo Sage was built in partnership with River AI by post-training DeepSeek V4 Flash on long-horizon contract work. The training used public data and synthetic data generated by real attorneys. Synthetic data, here, means examples written for training, by attorneys, rather than only contracts taken from a client file. DeepSeek V4 Flash is the existing model Sage was trained on top of. River provided technical support and the training infrastructure, which is the computing and the setup used to train the model. After reinforcement learning, the release says, the model went from meeting 70 percent of the pass criteria on the Legal Agent Benchmark (LAB) Contracts to 91 percent. Reinforcement learning is training in which the model is scored on how well it finished a task, then adjusted. A pass criterion is one check on a task. Seventy percent is about 7 checks in 10. Ninety-one percent is about 9 in 10. The release says the model reached quality comparable to much larger frontier models, with higher token efficiency, at a fraction of the cost. A frontier model, in that sentence, is one of the largest, most capable models. A token is a small piece of text the model reads or writes. Higher token efficiency means similar work with fewer of those pieces. Those scores, and the cost line, are Ivo’s, in the release. The release does not say an outside lab reran the test, and it does not print a dollar price for a task.

The test Ivo previewed at the same event. The release says other legal benchmarks grade only the edits a model makes, while most of the judgment in a review happens before any edit. A benchmark is a shared test, so models can be scored the same way. A good reviewer, the release says, has to know when to propose a change, when to leave acceptable language alone, and when a decision needs a person’s judgment. At Inscribe, Ivo previewed the Ivo-micro1 Contract Bench, built with data lab and research partner micro1. The bench is set up to measure five parts of that work: which issues to prioritize, when to show restraint, how to adapt to the specific deal, when to escalate, and how the model follows the playbook. Restraint means leaving acceptable language alone. Escalation means sending the decision to a person with authority, rather than the model deciding alone. A public set of results, the release says, will be available in the coming weeks. The page does not print that public set. Those lines are Ivo’s.

Four early findings in the release are Ivo’s research. The release says frontier models can only take contract review so far, and that strong judgment needs context. It says leading models handle some contract decisions well and miss others, especially when to push back and when to escalate. The first finding is about accepting language versus pushing back. When the right response to the other side’s change is to accept it or leave it, the release says models get it right 82 percent of the time. Eighty-two percent is about 4 times in 5. When the right response is to counter or reject, they get it right 46 percent of the time, a little under half. They add a needed new point of their own only 23 percent of the time, a little under 1 in 4. The release’s heading for that finding is that models concede easily and push back poorly. Those percentages are Ivo’s research, in the release.

The second finding is about escalation. The release says models meet only 23 percent of the escalation criteria the company’s expert attorneys set, on average. That is about 1 check in 4. The release says this is so even with explicit instructions and a dedicated tool for when and how to escalate. It reads the result as models making edit decisions on their own, rather than asking for permission. It names GPT-6 Astra as the outlier, at 71 percent, about 7 in 10 of those escalation checks. It says Astra gets there by escalating nearly everything, so it meets only 31 percent of its edit criteria, about 3 in 10, and finishes last overall. Those figures are Ivo’s research, in the release. The wire does not print the full ranking.

The third finding is about how big the edits are. The release says attorneys make 64 percent of their changes inline, averaging 92 characters. Inline means a change inside the existing sentence, rather than a rewrite of a whole block. Sixty-four percent is about 2 in 3. Ninety-two characters is about the length of a short clause. On the same contracts, the release says, models make 14 percent to 37 percent of their changes inline, and their edits rewrite whole blocks of text at 207 to 369 characters. Fourteen percent is about 1 in 7. Thirty-seven percent is a bit over 1 in 3. Two hundred to about 370 characters is a short paragraph. Those counts are Ivo’s research, in the release.

The fourth finding is what happens when the deal changes the answer. On average, the release says, models meet 51 percent of the criteria for applying a playbook’s standard positions. Fifty-one percent is about half. When the right move depends on the facts of the deal, that falls to 38 percent, a bit under 2 in 5. Every model drops by 9 to 15 points. The release says the playbook often still supplies the rule, and that recognizing when the deal calls for that rule is where models fall short. Those figures are Ivo’s research, in the release.

The same release also says Ivo Collaborate is generally available. General availability means customers can use it, not only watch a preview. Collaborate, as the release describes it, is an end-to-end contract intelligence platform for Fortune 500 legal teams, announced on the same page as the open model. It reads incoming contracts and manages the workflow from intake through signature. Intake is the moment a contract arrives. Routine agreements move by rules the legal team set. Issues that need a person’s judgment get flagged. A search the release calls Intelligence lets teams look through signed agreements to see what the business has promised and what it is owed, and to see patterns across the portfolio. The release says Ivo’s wider aim is to make contracts a source of intelligence for legal, procurement, finance, and sales. The release does not print a price for Collaborate, and it does not print a price for Sage.

Who the about box says Ivo is, and where the weights sit. Ivo calls itself an AI-native contract intelligence platform for the enterprise. It says the product helps legal and business teams move faster, reduce risk, and grow revenue by turning contracts into business intelligence. The about box says the company was founded in New Zealand and is headquartered in San Francisco. It says customers include leading global enterprises and high-growth companies across multiple industries. It does not name those customers. Media inquiries on the release go to LaunchSquad for Ivo, at ivo@launchsquad.com. The weights are on Hugging Face at ivo-ai-labs/Ivo-Sage. The model card lists an MIT license and names DeepSeek V4 Flash as the base, the model Sage was trained on top of. MIT is a license that allows download, use, and changes, including commercial use, if the license notice stays with the weights. The card is the place to download the model. The scores in the release remain Ivo’s measurements.

The picture is the Hugging Face social card for ivo-ai-labs/Ivo-Sage. A teal band at the top fades through blue into orange on the left and purple on the right, and the lower right of the frame goes dark. Thin black bars sit above and below the card. At the upper left, a small square mark sits beside the name ivo-ai-labs. Large white type under that reads /Ivo-Sage. At the lower left, the yellow Hugging Face mark sits beside huggingface.co. The card does not print a calendar date. It is the model-card graphic. It does not show a contract, and it does not show the benchmark scores.

In plain terms, Ivo said on Thursday, at its Inscribe event in San Francisco, that it published Ivo Sage, a free model attorneys, researchers, and developers can download, run, and train further. Ivo says it built Sage with River AI by further training DeepSeek V4 Flash on public data and on examples attorneys wrote. Ivo says that after reinforcement learning the model met 91 percent of the pass checks on the contracts portion of the Legal Agent Benchmark, up from 70 percent, at a quality it compares with much larger models and with fewer tokens. Those figures are the company’s. At the same event Ivo previewed the Ivo-micro1 Contract Bench, a test of which issues to prioritize, when to leave language alone, how to adapt to the deal, when to escalate, and whether the model follows the playbook. Ivo said a public set of results is coming in the weeks ahead. Ivo’s early research, in the same release, says models are better at accepting language than at pushing back, meet about 23 percent of the escalation checks attorneys set, edit in larger blocks than lawyers do, and score lower when the facts of the deal change the playbook answer. The same release says Ivo Collaborate is generally available. It does not name a customer, and it does not print a price.

RELATED

ONLINE…

Comments

guidelines

Loading…

Loading…

Sources