Your document agent is only as good as the parse behind it!
I've built a Grounded Document Agent that answers questions across long PDFs, attaches citations to every response, and lets you inspect the exact source behind those citations.
A lot of document-agent failures look like retrieval problems, but the damage often happens much earlier.
If a PDF table is flattened incorrectly, a chart loses structure, or the reading order breaks, better embeddings cannot recover what was already lost.
For this project, I used LlamaParse to convert PDFs into layout-aware Markdown while preserving page metadata. From there, LlamaIndex handles indexing and retrieval, while Ollama runs the embeddings and the local "qwen3:4b-instruct" model.
For every question, the agent retrieves the most relevant sections, generates an answer with numbered citations, and lets you inspect the exact source text and page behind each citation.
That matters because a citation does not automatically make an answer correct.
If the underlying evidence was parsed incorrectly, the model can still cite the right page and give you the wrong answer.
The useful part is being able to trace the answer through parsing, retrieval, and generation instead of treating a confident response as the final truth.
PDF → LlamaParse → Local retrieval → Local LLM → Grounded answer with citations
Here’s the complete project:
I’ve also written a detailed article explaining how to build the complete grounded document agent step by step.