Fetching latest headlines…

Dev

I told a friend a number was not in her document. It was on page 6, sideways.

Dev.toUnited States · NORTH AMERICA

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend everypage is a small tool that answers questions about a PDF using a model that runs on your own machine. It names the...

0 views0 likes0 comments

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

everypage is a small tool that answers questions about a PDF using a model that runs on your own machine. It names the page for everything it says, and when it could not read a page it tells you that instead of telling you "no".

I built it for a friend. She is not technical. She has a folder of family property papers, some printed from a computer, some scanned, some scanned lying on their side. She asks me the kind of question anyone asks about a contract: does it say this anywhere?

A while ago she asked me one of those, and I answered "no, that figure is not in the document". I had searched the text. The text I searched ended at page 4. The figure was on page 6, on a scanned page that my extraction had quietly skipped. Nothing warned me. A document with two missing pages looks exactly like a complete document that does not contain the thing.

She trusted my "no". That is the bug I wanted to fix, and it is not a model bug. It is a reading bug.

So the tool has three answers, never two:

verdict what it means
FOUND here is the sentence, and here is the page it is on
NOT FOUND every page was read, and it is not there
CANNOT SAY it was not on the pages I could read, but there are pages I could not read

Demo

The sample is an invented 6-page sale agreement (every name, place and amount is made up). Pages 1 to 3 have a text layer. Pages 4 and 5 are scans. Page 6 is a scan lying on its side, and it is the only page that mentions unpaid taxes.

First, the usual way: extract the text, give it to the model, ask.

$ python3 demo/naive.py samples/sale-agreement.pdf "Are there unpaid property taxes, how much, and who has to pay them?"
NAIVE: pdftotext returned 1449 characters and no warning
NAIVE ANSWER: The provided document does not contain information about unpaid property taxes,
their amount, or who is responsible for paying them.

That is a confident, fluent, wrong answer, and the model did nothing wrong. It was handed half a document.

Now the same model, the same question, through everypage:

$ python3 -m everypage read samples/sale-agreement.pdf
COVERAGE: read 6 of 6 pages
  page 1   text-layer   591 chars  legibility 0.54  OK
  page 2   text-layer   446 chars  legibility 0.56  OK
  page 3   text-layer   408 chars  legibility 0.67  OK
  page 4   ocr          384 chars  legibility 0.48  OK
  page 5   ocr          225 chars  legibility 0.41  OK
  page 6   ocr          459 chars  legibility 0.68  rotated 90°  OK

$ python3 -m everypage ask samples/sale-agreement.pdf "Are there unpaid property taxes, how much, and who has to pay them?"
COVERAGE: read 6 of 6 pages
ANSWER: Unpaid property taxes total $18,450. (page 6) The seller must pay this amount before closing. (page 6)
If the seller does not pay the unpaid taxes by closing, the buyer may deduct $18,450 from the price and pay
the county directly. (page 6)
  page 6 (ocr): "Property taxes for the years 2022, 2023 and 2024 remain unpaid in the total amount of $18,450."
  page 6 (ocr): "The Seller shall pay this amount in full before closing."
  page 6 (ocr): "If the Seller has not paid the unpaid taxes by closing, the Buyer may deduct $18,450 from the price and pay the county directly."
VERDICT: FOUND.

A question whose honest answer is no:

$ python3 -m everypage ask samples/sale-agreement.pdf "Does the agreement say anything about a broker commission?"
COVERAGE: read 6 of 6 pages
VERDICT: NOT FOUND. All 6 of 6 pages were read, so this is a real negative.

And the case that started all this. Same agreement, but page 6 is blurred past reading:

$ python3 -m everypage ask samples/sale-agreement-damaged.pdf "Are there unpaid property taxes, how much, and who has to pay them?"
COVERAGE: read 5 of 6 pages; could NOT read page(s) 6
VERDICT: CANNOT SAY. Nothing found on the pages that were read, but page(s) 6 could not be read. This is NOT a 'no'.

That last line is the whole project. It is the sentence I should have said to her.

Code

GitHub logo MariusGithub13 / everypage

Ask a PDF a question with a local model. It names the page, and says which pages it could not read.

everypage

Ask a question about a PDF and get an answer that names its page, from a model that runs on your own machine If a page could not be read, it says so instead of saying "no".

It was built for one person: a friend with a folder of property papers, part printed, part scanned, some scanned sideways. Her question is usually "does it say X anywhere?". The honest answers to that are three, not two.

verdict meaning exit code
FOUND here is the sentence, here is the page 0
NOT FOUND every page was read, and it is not there 1
CANNOT SAY it was not on the pages I could read, but some pages I could not read 2

What it does

python3 -m everypage read  FILE.pdf                # how each page was read, and whether it worked
python3 -m everypage find  FILE.pdf "18,450"       # exact text or
…

https://github.com/MariusGithub13/everypage (MIT). About 300 lines of Python, no framework. The tests run without a model.

How I Built It

The model is Gemma, running locally. gemma2:2b (1.6 GB) served by Ollama on localhost. OCR is Tesseract, PDF handling is Poppler. All open, all on the machine.

A 2-billion-parameter model is small, and that shaped the design. I did not ask it to be reliable. I asked it to do one easy thing, and I made the code responsible for everything that has to be true.

1. The code reads, and keeps the receipts. Each page is read on its own. If it has a text layer, that is used. If not, the page is rendered at 300 dpi and OCR'd. If the result does not look like language, it is retried at 90, 270 and 180 degrees and the best reading wins. Every page ends as OK or UNREADABLE, and the coverage line is printed before anything else.

2. The model sees every page, one at a time. There is no retrieval step choosing which chunks are "relevant". On a short legal document, the page a retriever skips is the addendum. It is slower. With this small model on a single CPU core it is roughly 15 to 20 seconds a page, so about two minutes for six pages, and my friend's question is worth two minutes.

3. The model must quote, and the code checks the quote. For each page the model returns up to three sentences copied from that page. The code looks for each one on the page, character for character after normalising spaces and case. A sentence that is not there is thrown away, whatever the model says about it. The final answer is written only from sentences that survived, each with its page.

4. The model is never allowed to say "it is not there". If nothing survived, the code decides: all pages read means NOT FOUND, anything less means CANNOT SAY. The exit codes are 0, 1 and 2, so a script cannot confuse them either.

The bug I shipped to myself on the way. My first legibility check counted words that contain a vowel. I ran it on the sideways page and it reported legibility 0.96 and status OK. The "text" it had approved was this:

pebueyoun urewel JuoweeIby oY} Jo SULI9} 1940 [[V

That is the page read upside down. Upside-down English is full of vowels. The blurred page scored a perfect 1.00 with ee ee mee ON me mere me. So the check that existed to catch unreadable pages was passing them with top marks, which is the original bug wearing a badge. The fix was to stop asking "does this look like words" and ask "does this contain the small everyday words real prose is made of": the, of, and, shall. Upside-down text scores 0.03 on that, real pages score 0.4 to 0.7, and both failures are now tests.

I used AI coding agents to help write and test this. The design rules and the failure they come from are mine.

Why Does Open Innovation Matter?

Three reasons, all practical.

The papers never leave the machine. Property papers, family papers and medical papers are the documents people most need help with and least want to upload to a server they do not control. With an open-weight model on localhost that question does not come up. It works with the network cable out.

I can put the rules outside the model. Because everything runs in my own process, the model's output is just a string my code can check against the page before anyone sees it. The guarantee does not depend on a provider's settings, a prompt that might be ignored, or a model version that changes under me next month.

It costs nothing to run, so "read every page" is affordable. Showing every page to a model one at a time is wasteful by API standards. Locally the only cost is a couple of minutes of CPU, so I can choose thorough over clever.

A closed model would have given a more polished paragraph. It would not have fixed my bug, because my bug was never the paragraph.

Limits, and one thing I will not do

She should not have to install anything, so I run it on her papers at my end and send her the answer with the page numbers. What she says about it stays between us. I am not going to quote her here.

So nobody is surprised: PDFs only. The legibility word list is English, so other languages need their own list (pages get marked unreadable otherwise, which is the safe way to be wrong). A page that is all numbers gets marked unreadable for the same reason. And a verified quote proves the words are on the page, not that the model understood them. Read the quote. This is a reading aid, not legal advice.

Prize Categories

Best Use of Gemma: gemma2:2b runs locally and is the only model in the project.

Comments (0)

Sign in to join the discussion

Be the first to comment!