Public evidence record · reviewed August 20, 2026

Formula OCR Verification Record: Math PDF to Word

This record examines three archived Offline OCR screenshots from two desktop sessions. It shows what the artifacts support, exposes visible output problems, and explains why they do not justify an accuracy percentage.

Prepared by Lingyun Zheng, developer of Offline OCR.

Disclosure: this is a first-party product evidence record, not an independent laboratory review. Claims are deliberately limited to the archived artifacts and product implementation inspected on August 20, 2026.

Short conclusion

The archived artifacts support a limited workflow claim: Offline OCR processed a one-page PDF in one session, and a separate May 28 session rendered mixed prose and formulas, started a PDF export, and produced a DOCX that opened in Microsoft Word. They do not establish recognition accuracy because the original source page, expected transcription, exact app version, and correction log were not preserved together.

Supported by the artifacts

  • A May 26 desktop capture visibly identifies “DeepSeek OCR Recognizing” for page 1 of a one-page PDF; the inspected product implementation labels this engine “Local LLM”.
  • A May 28 desktop capture shows mixed prose, an integral, a Taylor series and a matrix rendered inside the Offline OCR editor while HD PDF generation is started.
  • The May 28 DOCX is visibly open in Microsoft Word with paragraphs and mathematical notation rendered on the page.

Not established by the artifacts

  • No symbol-, expression- or page-level accuracy percentage can be calculated without the original input and a reviewed expected output.
  • The May 28 processing mode and exact Offline OCR version were not captured.
  • A static Word screenshot does not prove that every equation object or layout element remains editable.

Evidence inventory

The dates and observations below come from visible information in the archived screenshots. The May 26 and May 28 captures are treated as separate sessions, not combined into a single end-to-end test.

Evidence reviewed
Three original product screenshots stored with the website
Session dates visible
May 26, 2026 and May 28, 2026
Platform visible
Windows desktop
May 26 input visible
A one-page PDF; the screenshot shows page 1 of 1
May 26 engine visible
DeepSeek OCR; the inspected product implementation identifies it as a local model
May 28 content visible
English prose mixed with integrals, series, matrices and Greek symbols
May 28 exports visible
HD PDF generation and a .docx opened in Microsoft Word
App version
Not preserved in the screenshots
Original source and expected output
Not preserved with this public record
Accuracy score
Not reported; the evidence is insufficient for a valid calculation

What each screenshot shows

Offline OCR desktop screen showing DeepSeek OCR recognizing page 1 of a one-page PDF on May 26, 2026

Artifact 1 — recognition in progress

The dialog reads “DeepSeek OCR Recognizing” and identifies a one-page PDF. The inspected implementation labels this engine “Local LLM”. The artifact does not reveal the input page, final result, or the device network state during the capture.

Offline OCR editor generating an HD PDF from a May 28 math document containing integrals, a Taylor series and matrix notation

Artifact 2 — editor and PDF export

The May 28 editor view contains mixed prose and structured formulas. The export dialog states “Generating HD PDF”. Visible merged words remain in the document, so this is evidence of workflow completion, not perfect recognition.

A May 28 Offline OCR DOCX opened in Microsoft Word with mixed text, integrals, a Taylor series and matrix notation

Artifact 3 — DOCX opened in Word

Microsoft Word displays the exported .docx with mathematical notation rendered on the page. A screenshot confirms the visible result, but not the internal editability of every equation object.

Visible limitations

The record keeps recognition errors visible

The May 28 screenshots visibly contain merged boundaries such as “RecognitionAbstractThis”, “CalculusIn”, “OperationsEvaluating” and “FoundationsIn”. Some headings and numbered sections also run into adjacent text. These are material editing issues and should not be hidden in a trustworthy example.

Text segmentation needs review

Several words and headings are joined together. A user should expect to restore spaces and paragraph boundaries before reusing the document.

Formula rendering is not the same as formula accuracy

The screenshots show integrals, fractions, matrices and Greek characters, but the missing source page prevents a symbol-by-symbol correctness check.

Export evidence is narrower than an editability guarantee

Opening a .docx in Word proves a usable document was generated. It does not, by itself, prove that every visible formula was stored as a native editable Word equation.

Reproducible testing

What the next public OCR test must preserve

A future result can support accuracy claims only when another reviewer can repeat the run and compare the same input with the same unedited output.

  1. 1Publish or hash the original input page and record its license or ownership.
  2. 2Record the Offline OCR version, operating system, device and processing mode.
  3. 3Save the raw OCR result before any manual correction.
  4. 4Publish the expected transcription and define how symbols, structures and layout are scored.
  5. 5List every correction separately and report the review date.
  6. 6Keep successful samples and failure cases in the same public test set.

Questions about this evidence record

Does this record prove Offline OCR accuracy?

No. It verifies parts of a real product workflow, but the public artifacts do not include the original source and reviewed expected output required to calculate accuracy.

Was local OCR used in these screenshots?

The May 26 screenshot explicitly shows DeepSeek OCR, and the inspected product implementation identifies that engine as a local model. The screenshot alone does not prove the device network state. The processing mode used for the separate May 28 export session is not visible and is therefore left unclaimed.

Does the Word screenshot prove every formula is editable?

No. It proves that a DOCX opened in Microsoft Word and displayed mathematical notation. Object-level editability requires inspection of the actual file, which is not part of this public record.

Why publish an example with visible errors?

Visible limitations help users judge correction effort and prevent a workflow screenshot from being mistaken for an independently scored benchmark.

Test your own document before relying on any OCR claim

Use a representative page, keep the raw result, and review every formula before exporting or publishing the document.

Evidence record reviewed: August 20, 2026.