
30 Sep 2026
China Telecom open-sources TeleOCR, a 1.2B model topping document-parsing benches
China Telecom AI said Wednesday its TeleOCR document-parsing model, with about 1.2 billion parameters, set a new state-of-the-art on OmniDocBench v1.6 and ranked first on related camera-document and scientific-figure benchmarks, with weights open-sourced on GitHub and Hugging Face.
Most document software still treats a clean PDF and a wrinkled phone photo as two different jobs, and a giant general model is an expensive way to do a narrow one. A 1.2 billion-parameter specialist that China Telecom says beats frontier models on parsing tests, with the weights published, is the bet that a focused model can turn contracts and charts into structure.
On Wednesday, 30 September 2026, China Telecom Artificial Intelligence Technology Co., Ltd. announced TeleOCR. The release shortens the company name to China Telecom AI. The model comes from Xingchen AGI Lab, the lab inside that company. TeleOCR is a document-parsing model the lab built itself. A vision-language model, the VLM in the headline, reads both the picture of a page and the words on it. GlobeNewswire carries the release. The page stamps it September 30, 2026, at 2:19 a.m. Eastern. The dateline is Beijing. The source line is China Telecom AI. Those lines are the company’s.
The score the release leads with. China Telecom AI says TeleOCR set a new state of the art on OmniDocBench v1.6, with an overall score of 96.87 out of 100. State of the art, here, means the best overall score the company reports among the models it compared. OmniDocBench is a public test of how well a model turns a document into structured text. The release calls 96.87 the highest of all models it evaluated. It says the test covers 10 document types, 11 layouts, and five languages. It says TeleOCR ranked first in reading the text, in rebuilding tables, and in restoring the order a person would read the page. Those figures are China Telecom AI’s report of its own runs. The release does not attach an outside lab’s repeat of the same test.
Two more tests, and a contest, as the same release states them. The model has about 1.2 billion parameters. A parameter is one number the model learned while it was trained. About 1.2 billion is small next to the general models named in the comparison. On Wild-OmniDocBench v1.5, a test of documents a camera photographed, the release prints an overall score of 88.53. It says that is about one point ahead of the runner-up, and 4 to 10 points ahead of most end-to-end models. An end-to-end model, here, reads the page in one pass instead of handing it through a chain of separate tools. On PureDocBench the average is 78.41 across three tracks. On the track the release calls the hardest, “real degradation,” it says TeleOCR led second place by four points. Degradation, here, means the page is worn, blurry, or otherwise damaged. The release does not name the runner-up on those tests. Those figures are China Telecom AI’s.
The contest line. The release says TeleOCR took first place in the ICDAR 2026 Sci-ImageMiner Challenge. ICDAR is an international conference on reading documents. Sci-ImageMiner is its contest on scientific figures, the charts and diagrams in a paper. The task it names is scientific-figure-to-table: turn a figure into a table. The release says the TEDS score was more than two percentage points above the runner-up’s. TEDS is a score for how close a rebuilt table is to the original table. The body of the release does not print the two raw TEDS numbers. That placement is China Telecom AI’s.
Who the company says it beat, and what that comparison is about. China Telecom AI says TeleOCR outperformed larger specialized models, including MinerU 2.5-Pro and PaddleOCR-VL-1.6, and general-purpose models such as Gemini 3 Pro and GPT-5.2, on document parsing. A specialized model is built to read documents. A general-purpose model is built to answer many kinds of questions. The comparison in the release is about turning a page into text, tables, and formulas. It is not a claim that TeleOCR is better at open-ended reasoning. The release text does not print Gemini 3 Pro’s score or GPT-5.2’s score. A spokesperson for the lab said the 1.2-billion-parameter model beats models dozens of times its size on document parsing. “Dozens of times” is the spokesperson’s phrase, in the release. The release does not print those other models’ parameter counts next to that sentence. Those comparisons are the company’s.
Why the lab says one model can cover both a clean file and a phone photo. The release says existing systems are usually good at either a digital document or a camera photo, and not both. A pipeline, a chain of separate steps, handles a clean PDF and struggles when a photograph is bent or tilted. An end-to-end model holds up better on that bend and falls short on fine structure such as tables and formulas. TeleOCR, the lab says, puts three ideas in one model. Deformation-aware learning teaches the model to see the bend itself, so a separate tool does not have to flatten the photo first. Content-structure decoupling splits the shape of a table from the words in the cells, so tables, formulas, and charts can be rebuilt. Multi-model consensus labeling builds the training answers by taking a vote across different models, so one model’s habits do not write every answer. A training answer, here, is the correct reading used to teach the model. Those lines are China Telecom AI’s.
What comes out, and where a person can get the model. The release says TeleOCR turns digital files and camera photos into structured Markdown, tables, and formulas. Markdown is a plain-text format that keeps headings and tables. It lists eight jobs: finding the layout of a digital page, splitting the layout of a photographed page, reading text, reading formulas into LaTeX, reading tables into a computer-readable table, reading blocks of code, reading scientific figures, and reading seals. LaTeX is the markup scientists use for equations. A seal, here, is a stamp on a page. Code and weights are open-sourced on GitHub and Hugging Face. The weights are the learned numbers. Open-sourced means those files are published for others to download. A production API is on China Telecom’s Tianyi AI Open Platform. An API is the address a program calls so it can use the model without hosting the weights. The release’s download lines name GitHub at github.com/caipeng328/TeleOCR and Hugging Face at huggingface.co/StarDoc-AI/TeleOCR. A second Hugging Face page, XingChen-AGI/TeleOCR, uses the same model name. The contact on the release is Xing Chen, DataAiTech-service@chinatelecom.cn. Those lines are China Telecom AI’s.
What the lab says comes next. A spokesperson said TeleOCR shows that careful engineering and training aimed at one job can beat raw size. The spokesperson said a photo of a wrinkled contract, or a tilted whiteboard, can be parsed in one pass, with no outside plugin to flatten it first. The research team said the next phase is deeper use in financial documents, medical-record digitization, academic research workflows, and government archives. Digitization, here, means turning a paper record into text a computer can search. Those quotations and plans are in the release. They describe what the company says it will work on. The release does not name a bank, a hospital, a university, or an archive as a customer.
In plain terms, China Telecom AI said on Wednesday that a document model with about 1.2 billion parameters scored 96.87 out of 100 on OmniDocBench v1.6, placed first on a camera-document test and on PureDocBench, and took first in the scientific-figure-to-table task at ICDAR 2026. The company says those parsing results beat larger document models and general models including Gemini 3 Pro and GPT-5.2. The weights are on GitHub and Hugging Face. An API is on the company’s Tianyi AI platform. The scores are the company’s report. The release does not include an outside repeat of the same runs.
The picture is the graphic that ran with the GlobeNewswire release. On a dark field it names TeleOCR and a 1.2 billion-parameter document parser, with a trophy and leaderboard cards for OmniDocBench and the ICDAR scientific-figure contest. A line says the weights are open on GitHub and Hugging Face. The credit names China Telecom AI via GlobeNewswire. It is the company’s graphic for this announcement. It is not a photograph of a lab, and it is not a screenshot of the model reading a contract.
RELATED
Sources
- GlobeNewswire — China Telecom unveils TeleOCR, 30 Sep 2026
globenewswire.com
- GitHub — caipeng328/TeleOCR
github.com
- Hugging Face — StarDoc-AI/TeleOCR
huggingface.co
- Hugging Face — XingChen-AGI/TeleOCR
huggingface.co
- Tianyi AI Open Platform — teleai.com.cn
teleai.com.cn