
TeleOCR Targets Distorted Pages and Structured Document Parsing
Published by AINave Editorial
TeleOCR’s pitch is unusual for a compact document model: it aims to handle both born-digital files and camera photos with geometric distortion, while returning structured content such as tables and formulas. The project reports strong table scores on OmniDocBench v1.6, but those scores do not settle whether it will fit a particular production workload.
One model for digital files and camera photos
Maintained by XingChen-AGI, TeleOCR is an open-source vision-language model with approximately 1.2 billion parameters. Its documentation describes geometry-aware methods for camera-captured documents and says it parses distorted DocUNet and DIR300 pages without dewarping preprocessing or a separate rectification model. That is a useful distinction for pipelines receiving phone photos: correcting page geometry need not be a separate documented step in those examples, though the claim is not a guarantee across all image conditions. The model’s intended scope and reported approach
The project also targets structured output rather than plain text alone. Its examples represent content blocks with bounding boxes and table cells with text and row and column positions. That can matter when downstream software needs to preserve layout or table relationships, not just retrieve words. The documentation identifies Transformers, AutoProcessor and AutoModel, but the supplied example is incomplete, so it does not establish a ready-to-run integration path. The described output structures and loading components
Strong table scores, with metric-specific trade-offs
On OmniDocBench v1.6, TeleOCR reports an overall score of 96.87, Table TEDS of 97.05 and Table TEDS-S of 98.52. These are project-reported benchmark results, not independently reproduced findings in the supplied material. In the same comparison, OvisOCR2 scores better on formula CDM and read-order edit, where lower is better for the latter. So the headline overall result does not mean TeleOCR leads every task that may matter in a document pipeline. The reported OmniDocBench scores and metric comparisons
There is also a discrepancy in the reported PureDocBench results: the abstract gives an overall score of 78.41, while the benchmark table lists 86.90 for clean documents, 77.47 for digital-degraded documents and 70.85 for real-degraded documents. Those figures appear to describe different groupings or protocols, but the supplied material does not explain the difference. They should not be treated as directly interchangeable. The PureDocBench figures as presented
What the published details leave open
The approximately 1.2B parameter count makes TeleOCR a compact specialized model relative to the much larger general vision-language models in the comparison. It does not, by itself, show what hardware or operating cost a deployment needs. The documentation gives no VRAM requirement, inference speed, batch-size guidance or maximum image resolution. The stated model size and missing deployment measurements
For teams evaluating it, the useful question is whether its strengths match the documents they actually process. A system that preserves table structure or reads a warped page can be valuable, but the reported scores do not establish performance on every language, formula type or scan condition. TeleOCR gives a concrete candidate for those workloads; its published evidence still leaves deployment fit to be determined on the target corpus and hardware. The project’s benchmark scope and documented limitations






















