LQLQ ASSESS
CANDIDATE REPORT · OQ-124 · SUB-8130d14b
July 2026 cycle
PASS
pass rate this cycle: 24%, 11 of 45 cleared both gates

Shreya Poddar

OQ-124An agent that reads your files instead of RAG
RAG chunks, embeds, and indexes — and quietly misses things. read the full question
There's another way. A local agent that reads your files directly, the way a person would: it searches, opens, and follows cross-references. Build a thin version. Test it against a basic RAG baseline on a few real questions. Include one that spans several documents, and one needle-in-a-haystack. Report accuracy, latency, cost, and which you'd ship. A working POC measured against RAG beats arguing RAG is bad.
The machine bar Cleared · you finished above Claude's own run
Peer review✓ Cleared · 3 LQ members read it blind and the work held up
Top 20% · of the cohort
Your Build

A comparison setup with four components: a file-reading Agent, a RAG pipeline (chunking, embedding, ingesting), a corpus of legal documents, and an independent evaluator. A reports.html file consolidates results, and a manual_eval.py script extends evaluation to GitHub terms-of-service documents. Both the Agent and the RAG pipeline are configured to produce citations and to operate only over a whitelisted set of files.

How you approached it

Built both approaches against the same legal corpus and ran them through a shared evaluator over a question set covering single-document queries, a large cross-referenced document, absent-information queries, and a needle-in-a-haystack case. The evaluator generated reports.html, and an additional manual pass on GitHub T&C documents was used as a sanity check. Files were placed in an ASMT folder inside the working directory.

Where You Sit — The Field
cleared their bar · finished below it · the pass bar on that question · you
OQ-040OQ-037OQ-009OQ-122OQ-124YOUOQ-007OQ-030OQ-028OQ-026OQ-112OQ-042OQ-045OQ-115OQ-008OQ-130OQ-010OQ-126OQ-011OQ-046
0n = 45 candidates · 19 open questions, hardest bar first100

The bar isn't set by a committee. It's set by Claude's own attempt at each open question. Same budget, same rubric: where it landed on your question is where your line sits. The bar moves with the question, so it sits high wherever the solo run scored high. Everyone here is measured against the bar they actually faced, on the question they actually drew. 11 of 45 finished at or above their own this cycle.

What You Demonstrated
── your score · band likely range if regraded
SOLIDFUNCTIONALOrchestrationhow the work was structuredwell aboveToolingwhat was built to build withwell aboveSelf-awarenessknowing the work's limitsaboveQuestion fidelityanswering what was askedwell above
The Machine's Run

Claude Sonnet 5, working alone on this exact question: 131 turns over 2 self-review loops, no human input. The complete run is open source, every file it wrote, unedited:

github.com/LegalQuants/lq-assess-machine-baselines/tree/main/run/OQ-124
Among Those Who Beat It
YOU
each dot = one candidate who finished ahead of the machine, ordered by margin
System Design — You vs the Machine
System architecture as built, generated from the session
Fig. 1a — CANDIDATE · architecture as built, from the session · branches = verification loops
The machine's architecture on the same brief
Fig. 1b — CLAUDE ALONE · its architecture on the same brief, from its run
What LQ Members Said

What other LegalQuants members made of this work, reading it blind. Their reads decide the peer gate above. They do not change your rubric grade.

Good comparison and even did a manual test too.— LQ member, blind review
I like that this candidate was security-aware and was actively trying to hide his API key. The candidate found good tests too to consider success metrics.— LQ member, blind review
Sounds like the candidate enjoyed the tasks, that's always a great mindset.— LQ member, blind review
This candidate infused their legal experience throughout their prompting. I liked that they started with parallel agents and switched nimbly from open model to Anthropic when time seemed to be a concern. They were very adaptive.— LQ member, blind review
What Drove the Score · your own prompts, verbatim
07:00
your prompt · L107no evaluator should outsmart the others - lets use fable 5 for that so it can actually test the edge cases also, mix of court cases, us, uk, indian
The grader ran on a different model than the systems it graded, and three jurisdictions entered the corpus.
44:00
your prompt · L590the filename is github documents, this is not in the existing corpus but an independent check, should I give the questions and then you can ask both in parallel and prepare a table of their answers and I will manually check it and give my verdict.
A second corpus was built under manual_eval and both systems re-ran on documents the benchmark never saw.
46:00
your prompt · L627why do you have to get the files? should the agent not do this independently?
Neither system could ingest a document on its own, and that became the Limitations section of the report.
Others On Your Question

2 other candidates drew this open question. How each took it on, in the same words the blind reviewers read. The outcome shown is the same one the field chart above shows as dots.

○ finished below the barPartial assessment submission for OQ-124: a page-numbered text extraction of 8 HK wills/probate corpus files,…read more
Partial assessment submission for OQ-124: a page-numbered text extraction of 8 HK wills/probate corpus files, a written decision log (AMENDMENTS.md), a partial report README describing a planned comparison between a thin file-reading agent and a TF-IDF RAG baseline, plus a debrief interview record (answers.jsonl with 5 Q&A entries).
○ finished below the barA graphical knowledge-graph representation of Polish law that visualizes statutes, their context, and…read more
A graphical knowledge-graph representation of Polish law that visualizes statutes, their context, and cross-references between legal codes, adapted from a pre-existing personal project based on MCP codebase memory techniques.
Methodology

Two gates. You clear the assessment by clearing both. The machine bar is Claude Sonnet 5 working alone on your open question, on your effort budget. It was graded blind on the same rubric as you, across four dimensions and three independent gradings, from the working record of the session rather than the artifact alone. You clear it when your grade of record meets or exceeds its median. Its complete run is linked above. Peer review is members of the LegalQuants community reading the anonymized work and judging whether the session shows real thinking. You clear it unless a majority of your readers judge the work hollow. 11 members completed 179 blind reads across the cycle.

Take the assessment at assess.legalquants.com
LQ ASSESS · CANDIDATE REPORTJuly 2026 cycle · pass rate 24%