[KongchangAI]
· 1 min read· 792 words

LightOn OCR-3 Launches on OpenDocRouter: A New Benchmark for Open-Source OCR Value

LightOn OCR-3 Launches on OpenDocRouter: A New Benchmark for Open-Source OCR Value

LightOn OCR-3 hits OpenDocRouter at ~$3.19/1K pages, matching Gemini 3.8 flash low performance at 45% lower cost.

Open-weight OCR model LightOn OCR-3 is now available on OpenDocRouter at $0.28/1M input tokens and $1.40/1M output tokens, roughly $3.19 per thousand pages. Official benchmarks place it on the Pareto frontier of open-weight OCR models — delivering performance comparable to Gemini 3.8 flash low at about 45% lower cost. Its strengths lie in text grounding and layout accuracy, with decent table handling and passable chart support. For developers, the model offers a pay-per-use path requiring no inference infrastructure, while retaining the option to self-host the same weights for compliance and data privacy needs.

LightOn OCR-3 Officially Arrives on OpenDocRouter

The open-weight OCR model LightOn OCR-3 is now available on the OpenDocRouter platform, priced at $0.28 per million input tokens and $1.40 per million output tokens — roughly $3.19 per thousand pages (based on ParseBench estimates). For teams that need to process scanned documents, PDFs, and structured tables at scale, this pricing is genuinely compelling.

This launch is more than just adding another model to a roster. OCR has long been the entry point for document intelligence pipelines, and hosting an open-weight model at a competitive price on a routing platform means developers can call it on demand — without building their own inference infrastructure — while getting performance that approaches closed-source alternatives.

LightOn OCR-3 launches on OpenDocRouter

ParseBench Results: Sitting on the Pareto Frontier

According to official benchmark results on ParseBench, LightOn OCR-3 sits on the Pareto frontier of open-weight OCR models. That framing matters — it means there's no obvious alternative on the performance-cost curve that is simultaneously cheaper and more capable.

The official comparison target is Gemini 3.8 flash low. The conclusion: comparable performance, but LightOn OCR-3 costs roughly 45% less. In high-volume, budget-conscious document processing scenarios, a significant price reduction at near-equivalent accuracy often delivers more practical value than marginal accuracy gains.

Performance Across Document Types

The official capability breakdown is refreshingly candid, and worth paying attention to:

  • Tables: Decent — handles common structured data extraction reasonably well
  • Charts: Workable — clearly not a strong suit; complex charts may still require human review
  • Grounding / localization: Quite good — spatial recognition of text positions on the page is notably accurate

This capability profile — strong at grounding, decent on tables, passable on charts — makes it best suited for workflows centered on layout reconstruction, text extraction, and position annotation. Tasks that demand deep semantic understanding of charts should be evaluated more carefully.


Pareto frontier is a concept from economics and multi-objective optimization. It refers to the set of solutions where you cannot improve one metric without degrading another. In the context of OCR model performance vs. cost, being "on the Pareto frontier" means: no cheaper model achieves higher accuracy at the same budget, and no more accurate model costs less at the same performance level. This is a stronger claim than simply having good "price-to-performance" — it describes optimality along the entire tradeoff curve, not just dominance on a single axis. ParseBench is a standardized evaluation suite designed specifically for document parsing models, covering scanned documents, PDFs, tables, and more. It provides unified accuracy and speed metrics and has become a widely adopted benchmark for comparing OCR and document understanding models in the open-source community.


The Real-World Cost-Performance Tradeoff

Breaking down the pricing makes things clearer: at $0.28/1M input and $1.40/1M output, the output side is significantly more expensive — consistent with the general economics of generative models. The $3.19/thousand-page cost is the most intuitive metric for estimating actual OCR project spend, since page count is typically the primary billing variable in bulk archiving, contract parsing, invoice digitization, and similar use cases.

Compared to solutions that lock you into a closed-source API, open-weight models offer an additional layer of flexibility: while calling via OpenDocRouter on a pay-per-use basis is the most convenient option today, teams can in principle deploy the same weights in their own environment — giving them more control over compliance, data privacy, and long-term costs.


Open-weight models differ subtly but importantly from fully open-source models. Open-weight models release trained model parameters for download and deployment, but the training code, datasets, or commercial use terms may be restricted. Fully open-source models typically also release the training pipeline and data. LightOn OCR-3 falls into the former category, meaning developers can obtain the weight files and run them in private environments — satisfying data residency and compliance requirements — but deep customization or retraining may be subject to license constraints. Before deploying any open-weight model in production, always review the specific license agreement (e.g., Apache 2.0, CC BY-NC, etc.) as a necessary due-diligence step.


What This Means for Developers

For developers building document processing pipelines, LightOn OCR-3 offers a low-friction "try it first" option. You can test the API directly on OpenDocRouter and consult the full ParseBench leaderboard to compare model scores across different dimensions before committing to production.

A sensible deployment strategy: use OCR-3 for the bulk of text extraction where grounding accuracy matters and cost sensitivity is high, while reserving human review or higher-tier models as a fallback for chart-heavy pages or scenarios with extreme precision requirements. As open-source OCR continues to advance along the cost-performance curve, the per-unit cost of document intelligence has further room to fall.

Share:

Related articles