← News

AppZen ships ZenLM Plus finance-specialized LLMs that beat frontier models on expense audit

AppZen said Thursday it introduced ZenLM Plus, a family of finance-specialized large language models for Travel & Expense audit that led five of six control comparisons against seven frontier models and will be generally available in December 2026 after select-customer access now.

Expense audit is where autonomous finance either sticks or stays a demo. AppZen’s bet is that a model trained on real receipts and policy calls can beat a general frontier model on the messy controls that leak money, then hand the rest of the stack to Mastermind so agents act with an audit trail — and it published the head-to-head numbers.

On Thursday, 1 October 2026, AppZen introduced ZenLM Plus, a family of finance-specialized large language models built for the decisions at the center of finance operations. A large language model is the kind of software that reads documents and writes a judgment from them. Finance-specialized, in AppZen’s phrase, means these models were built for expense rules and receipts, not for general chat. Travel and expense, which the release shortens to T&E, is the spend an employee puts on a trip, a hotel, or a meal. The dateline is San Jose, California. PR Newswire carries the release. The source line is AppZen. Those lines are the release.

How AppZen says it scored the models. It compared ZenLM Plus with seven frontier models: Opus 5, Sonnet 5, Gemini 3.1 Pro, Gemini 3.5 Flash, GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.6 Luna. Frontier, here, is AppZen’s word for the general models it lined up against its own. For each audit control, every system received the same expense data, the same supporting documents, and the same customer configuration. AppZen scored two things. Precision is the share of flagged expenses that were true violations. Recall is the share of true violations the system caught. It combined the two into an F1 score from 0 to 100, where 100 means perfect accuracy: every real violation caught, and nothing extra flagged. Those definitions are AppZen’s. The release does not print how many expense lines were in the test.

ZenLM Plus led five of the six individual audit-control comparisons. On targeted policy-category cases it scored an F1 of 97.4, against 88.3 for GPT-5.6 Sol, which AppZen calls the strongest frontier result in that test. 97.4 is close to a perfect score on AppZen’s scale. 88.3 is lower: more misses, more false flags, or both. The subhead on the same release says 97.4 is 8.7 points above the strongest frontier model tested. 97.4 minus 88.3 is 9.1. Both numbers are printed on the page. The release does not explain the difference between 8.7 and 9.1. AppZen also says ZenLM Plus led all four consolidated policy groups, including advantages of 10.8 points in premium travel and upgrades, 4.2 points in electronic devices and gifts, and 2.5 points in policy and documentation exceptions. Those three gaps are the ones it prints. It says it led the fourth group too, and it does not print that fourth margin. Those scores are AppZen’s.

The widest gap on a single control was non-conforming receipt detection. ZenLM Plus scored 92.4 F1, 7.2 points ahead of Opus 5. AppZen says this job is more than recognizing a document. A payment slip can prove a card was charged and still fail as a merchant receipt. An emailed order confirmation can be valid even when it does not look like a paper receipt. ZenLM Plus also led on catching duplicates across reports, on receipt itemization verification, and on merchant category matching. Receipt verification was the one control a frontier model won. ZenLM Plus scored 93.3 F1. Gemini 3.1 Pro and Sonnet 5 scored 94.1. 94.1 minus 93.3 is 0.8 points, on the numbers AppZen prints. Those lines are AppZen’s.

On cost, AppZen says ZenLM Plus had the lowest modeled inference cost of the systems it evaluated. Inference, here, is the cost of running the model on a batch of expenses, not the cost of training it. Modeled means AppZen calculated the cost. It is not a customer invoice. Even GPT-5.6 Luna, which AppZen calls the least expensive frontier model in the test, was about twice as costly per 1,000 audited expense lines. Opus 5 was about 50 times as costly. Twice means Luna costs about two units on AppZen’s model for each unit ZenLM Plus costs. Fifty times means Opus 5 costs about fifty of those units. The release does not print a dollar price. Those lines are AppZen’s.

What the models are supposed to check, as the release lists it for Travel and Expense teams inside AppZen Expense Audit. On policy and category, they apply the customer’s own rules to premium travel and upgrades, electronic devices and gifts, personal expenses, merchant categories, fuel-service restrictions, and missing-receipt affidavits. An affidavit, here, is the form an employee signs when the receipt is gone. On receipt validity, they decide whether the evidence counts as a real merchant receipt, and whether the merchant, date, amount, expense type, and currency match the expense that was submitted. On documentation, they check that hotel and meal receipts have the required line-item or daily detail, and that the expense arrived inside the company’s time limit. On duplicates, they look across that employee’s other reports and across other employees’ reports for the same receipt or the same purchase submitted twice. Depending on the control, the model can read what the employee typed, what was pulled off the receipt, the payment type, the company’s categories and thresholds, merchant information, card data, and related expense reports. AppZen says the judgment follows how that customer defines risk, not a general guess about what a reasonable company would do. Those lines are AppZen’s.

Anant Kale, chief executive and co-founder, said autonomous finance requires more than a powerful model. He said it requires the complete stack: finance-specialized models, proprietary data, intelligent workflows, and enterprise governance, working together so the system can decide and act with accuracy, transparency, and control. Proprietary, here, means the company’s own data, not a public sample. He said ZenLM Plus now outperforms the frontier models AppZen tested on expense audit tasks, at a lower cost, and that combined with AppZen’s agent-first platform it lets the company’s AI agents do more of the work on their own, at enterprise scale. An agent, in that sentence, is software that takes the next step in the audit, not only a chat reply. That quotation is his, in the release.

Kunal Verma, chief technology officer and co-founder, said the models are trained on real financial documents, policy decisions, and audit outcomes, then evaluated task by task before they reach production. He said that produced stronger results on most of the controls tested, at a fraction of the modeled inference cost. He also said that when a frontier model is better for a specific task, AppZen can use it. The goal he states is the best outcome for the customer. That quotation is his, in the release. The receipt-verification result, where Gemini 3.1 Pro and Sonnet 5 scored 94.1 against 93.3, is the control the same release says a frontier model won.

Where the models run. AppZen says ZenLM Plus sits inside the Mastermind Platform. Mastermind is the layer that holds the workflows and the controls. It routes each finance task to the method that fits: a specialized ZenLM Plus model, deterministic logic, or a frontier model. Deterministic logic is a fixed rule. The same input gives the same answer, with no model guess. AppZen says that split lets a finance team automate more work at a lower cost and still keep the accuracy, the governance, the audit trail, and the evidence for every decision. An audit trail is the record a reviewer can check later. Those lines are AppZen’s. The Mastermind product page describes the same idea in product language. It says ZenLM models are purpose-built for accounts payable and travel-and-expense work, including reading hotel folios, and that each agent records its reasoning so a person can read the logic in an audit log. A folio is the hotel’s itemized bill. The scores, the model names, and the December date are not on that page. They stay with the release.

Who can use it, and when. ZenLM Plus is available now for select customers inside AppZen Expense Audit. Select means a limited set of customers, not every account. General availability is planned for December 2026. Generally available means the ordinary release, the one a customer can take without a special arrangement. The release points readers to appzen.com/mastermind-ai-automation-platform. It does not print a price.

Who AppZen says it is, in the about box on the same release. It calls itself the leader in autonomous finance operations, software that carries accounts payable, expense auditing, and compliance without adding headcount. The box says Fortune 500 companies trust it, and it names Amazon, Boeing, Salesforce, Novartis, JPMorgan Chase, VMware, and ServiceNow. Those names are the company’s list of customers for the business. The release does not say each of them is a ZenLM Plus customer. The same box says chief financial officers reduce operating costs by up to 50 percent, reach automation rates of 80 percent or more, and see a measurable return within weeks. Up to 50 percent is the ceiling the company states, not a result it says every customer hits. Those percentages describe the company. The release does not attach them to Thursday’s model scores. The release does not print a funding round, a valuation, or a device clearance. Press contact on the page is Megan Botta at Pitch Public Relations.

The picture is a crop of AppZen’s Mastermind product page around the ZenLM section. The top bar shows Products, Solutions, Customers, Resources, and Sign In. Under it, a row of links reads Mastermind Platform, AI Agent Studio, DIY Low Code, AI Analytics, and Pre-built Apps, with the left edge of the frame cutting into the first label. On the right, a product card sits over a photo of an office: Hotel Itemization, Folio Validation, and a small badge marked 03. A black bar along the bottom reads “AppZen ZenLM Plus — finance-specialized LLMs,” and the page heading “ZenLM: AI models purpose-built for” still shows through behind that line. It is a picture of the product screen. It does not print a calendar date.

In plain terms, AppZen said on Thursday that ZenLM Plus, its finance-specialized models for expense audit, beat seven frontier models on five of six controls, including a 97.4 F1 on policy-category cases against 88.3 for GPT-5.6 Sol. Receipt verification was the exception: Gemini 3.1 Pro and Sonnet 5 scored 94.1, and ZenLM Plus scored 93.3. AppZen also says the specialized models were the cheapest to run in its cost model, with Opus 5 about 50 times as costly per 1,000 expense lines. The models are in Mastermind now for select Expense Audit customers, with general availability planned for December 2026. The scores are AppZen’s. The release does not name a customer whose audit ran on ZenLM Plus, and it does not print a price.

RELATED

ONLINE…

Comments

guidelines

Loading…

Loading…

Sources