Originally published on · Updated on .
We gave the same fictional PDF to ChatGPT Free, Gemini Free, and free NotebookLM and asked the same questions. In IANautaLab’s test on September 28, 2026, all three tools retrieved much of the document. The most useful difference for choosing an AI appeared in omissions and a chart comparison, rather than a podium of scores.
Here you can see how the test was conducted, what each tool got right or missed, and when each workflow may make sense. The PDF used in the experiment is available for checking the answers or trying it yourself. This round’s results do not measure the products’ universal performance.
Short answer: which stood out on this PDF?
In this round, with this document, these free versions, and these conditions, NotebookLM obtained the highest raw score on the 24 questions: 46/48. Gemini Free scored 45/48 and ChatGPT Free 43/48. The gap is small. In the independent summaries, Gemini and NotebookLM received 12/14; ChatGPT, 10/14. These numbers are points assigned according to this test’s assessment criteria, not accuracy percentages.
The decisive finding for readers is the pattern: all did well on the table and footnote; ChatGPT retrieved chart values but misidentified the interval with the greatest increase. All three summary responses omitted at least one important condition or qualification. Even the highest-scoring tool therefore needs checking against the original.
How we conducted the test
IANautaLab created Test Document v1.1, an entirely fictional ten-page operational report. It combines text, numbers, a table, a chart, a footnote, a local exception, a joint condition, and a final qualification across its sections. The scenario does not describe real people, projects, or spending. Before receiving the responses, we defined the PDF, two requests, 24 questions, the expected answers, and scoring criteria. All three tools received the same document and requests.
To keep the comparison consistent, we tested only free options available in the round: ChatGPT Free, Gemini Free, and free NotebookLM. The short summary was a separate task from the 24 questions. We checked the responses against the document and criteria established before the run. We prepared a way to assess responses without identifying the tools, but it was not fully applied in the final assessment. We did not test paid plans.
We assigned two points when an answer was correct and preserved the necessary conditions; one when it was partially correct but omitted something; and zero when it was wrong, absent, or invented. Thus, 43/48 is the sum of points in this assessment and does not measure a general accuracy rate on other PDFs. No relevant factual hallucination was identified in this round. That is not a guarantee for future uses.
Try it yourself: download the PDF used in the test and try it with the tools you already use. Ask for a short summary, then ask which consecutive interval had the largest increase in the chart and whether the project’s second round was approved. Check pages 8–10 before accepting the answer.
Overall test results
The table separates questions from summary. Table, chart, footnote, combination, and condition are subsets of the 24 questions; their points must not be added to the total again. The condition column refers specifically to question 15. Scroll the table horizontally on narrow screens.
| Tool | Questions | Summary | Table | Chart | Footnote | Combination | Condition (question 15) |
|---|---|---|---|---|---|---|---|
| ChatGPT Free | 43/48 | 10/14 | 4/4 | 2/4 | 2/2 | 2/2 | 2/2 |
| Gemini Free | 45/48 | 12/14 | 4/4 | 4/4 | 2/2 | 2/2 | 2/2 |
| Free NotebookLM | 46/48 | 12/14 | 4/4 | 4/4 | 2/2 | 2/2 | 2/2 |
“Table” and “chart” group questions based on those elements; “footnote” concerns the footnote, “combination” brings together information from different sections, and “condition” assesses the second-round requirements in question 15. The summary was scored separately on a 14-point scale. The discussion below explains differences by capability without turning a small gap into a universal winner. Question 12, about whether the second round had been approved, was assessed separately; its partial score does not replace question 15’s score.
Where the differences actually appeared
Chart: retrieving values was not enough
The chart showed 20 enrollments in January, 25 in February, 31 in March, 34 in April, 38 in May, and 40 in June. The greatest increase between consecutive months was February to March: +6. In this run, ChatGPT Free answered April to May: +4. In another question, it correctly retrieved values from the same chart. The finding is therefore specific: it read the values correctly but failed the comparison requiring it to identify the greatest increase. Gemini and NotebookLM got both chart questions right.
Conditions and qualifications: what a short summary can hide
The fictional project’s second round had not been approved. To propose it, all local targets had to be met and spending kept within the ceiling; the North did not meet its target. In question 12, which asked whether the second round had been approved, ChatGPT mentioned a joint condition without specifying its two requirements or the North’s shortfall and received 1/2. This is a different question from question 15 in the table, scored 2/2. It shows why “it was not approved” can be a correct conclusion while still omitting the logic needed for a decision.
In the question requesting a broader synthesis, all three tools were partially correct for different reasons. ChatGPT preserved the qualification about what enrollments can establish but omitted the second-round condition. Gemini and NotebookLM preserved the condition but omitted the final qualification: enrollments alone do not demonstrate learning, adoption of practices, retention, or average attendance. In the independent summaries, ChatGPT omitted both the condition and the qualification; Gemini and NotebookLM omitted the qualification. Anyone using AI for a decision should specifically ask about conditions, exceptions, and limits, as well as requesting an overview.
Table, footnote, and information across sections
All three tools received 4/4 on table-based questions, 2/2 on the footnote about duplicate enrollments, 2/2 on combining information from different sections, and 2/2 on the second-round condition in question 15. This describes what happened with this PDF. It does not prove that any of them will always read tables, footnotes, or charts correctly in other files.
The three tested tools: what to look for
ChatGPT Free: varied questions, with the comparison checked
ChatGPT Free answered much of the questionnaire and lets you continue asking about a file in a chat. OpenAI’s documentation confirms document uploads, including PDFs, on the free plan, subject to limits. The 43/48 result included the error identifying the largest chart increase and partial answers about local targets, the approval conditions in question 12, and the broader synthesis. If a decision depends on a calculation, comparison, or condition, open the PDF and check manually. The Visual Retrieval documentation describes its own availability; we do not assume that every PDF uploaded to ChatGPT Free receives complete visual reading.
Gemini Free: strong performance on this set
Gemini Free scored 45/48 on the questions and 12/14 on the summary. It answered both chart questions correctly in this run but was partially correct on local targets, second-round approval in question 12, and the broader synthesis. Gemini’s official help documents PDF uploads and analysis, with tighter limits without a subscription and account or organization conditions for Drive files. It can be a practical route when the document is already in that environment. Qualifications and permissions still need checking.
Free NotebookLM: organized sources and this round’s highest raw score
Free NotebookLM scored 46/48 and 12/14. Google also calls the product Gemini Notebook. Its help on sources documents PDFs as a source for questions and summaries. In this test, it was partially correct on second-round approval in question 12 and the broader synthesis; in the short summary, it omitted the final qualification. A notebook can be useful when you need to return to sources, but a citation or high score does not remove the need to read the original passage.
Other documented options, outside this test
The previous article also presented Claude, Adobe Acrobat AI Assistant, and Copilot in OneDrive. They did not participate in this test, so they have no scores in this table. They are alternative workflows, not places in an experimental ranking.
- Claude: its official help documents uploading PDFs and other files in conversations. Consider it if you already work in that chat; test your own document and check the passages.
- Adobe Acrobat AI Assistant: Adobe documents answers and references in the PDF reader. Free access is limited; full or continued use depends on the applicable plan or license. Confirm availability in your account and open the cited source.
- Copilot in OneDrive: Microsoft documents OneDrive summaries for accounts with an eligible subscription or license. This does not describe free Copilot chat. It may make sense when the files are already in OneDrive and the feature is enabled.
Which should you use for your task?
To ask about an individual PDF: ChatGPT Free or Gemini Free can be starting points if your account permits uploads. Both retrieved much of the document in this round; ChatGPT needed additional checking on the chart comparison. To work in a source-centered space: NotebookLM may feel more natural, especially when returning to the original passage is part of the process. For a PDF already open in Acrobat or files in OneDrive: consider those workflows only with the necessary access, without attributing this test’s results to them.
If you want to learn how to build a good summary and check study material, read the method guide for PDFs and study materials. For comparing assistants on broader tasks, see the ChatGPT, Gemini, and Claude comparison. Here the focus is evidence from the same PDF given to three free versions.
Test limitations and care with your PDF
This was one round using one PDF created by IANautaLab, on September 28, 2026. Products can change, and responses may vary between runs. We did not assess speed, interface, intensive OCR, complex scanned PDFs, hundreds of pages, multiple documents, PDF editing, enterprise integrations, or legal aspects of privacy. Paid versions were not tested. Do not use these totals to infer performance in those scenarios.
Before sending a real document, check whether you may share it with the service and read the applicable policies. To check an answer, locate the original passage and confirm numbers, dates, units, conditions, and qualifications. A fluent sentence may omit exactly the detail that changes the conclusion. If the PDF is scanned, reading also depends on recognizing the content; accepting the file does not prove that every image or chart was interpreted.
How to repeat a useful test without turning a score into universal truth
- Use the same authorized PDF in the tools you want to compare; record the date, plan, and whether the upload worked.
- Make the same request: “Summarize the document in up to 80 words. Include targets, results, spending, the condition for a next round, and qualifications. Say when information is missing.”
- Ask a verifiable question: “What was the largest increase between consecutive months in the chart? Show the interval and the difference.”
- Return to the original: compare the answer with the chart, table, footnote, and qualifications section; record errors and omissions separately.
Test Document v1.1 lets you repeat this exercise. It is fictional and does not include an answer key in the published file. The best choice for you is the tool that meets the task, preserves important conditions, and lets you check the origin of claims with the least reasonable friction.

