Shreya Poddar
RAG chunks, embeds, and indexes — and quietly misses things. read the full question
A comparison setup with four components: a file-reading Agent, a RAG pipeline (chunking, embedding, ingesting), a corpus of legal documents, and an independent evaluator. A reports.html file consolidates results, and a manual_eval.py script extends evaluation to GitHub terms-of-service documents. Both the Agent and the RAG pipeline are configured to produce citations and to operate only over a whitelisted set of files.
Built both approaches against the same legal corpus and ran them through a shared evaluator over a question set covering single-document queries, a large cross-referenced document, absent-information queries, and a needle-in-a-haystack case. The evaluator generated reports.html, and an additional manual pass on GitHub T&C documents was used as a sanity check. Files were placed in an ASMT folder inside the working directory.
The bar isn't set by a committee. It's set by Claude's own attempt at each open question. Same budget, same rubric: where it landed on your question is where your line sits. The bar moves with the question, so it sits high wherever the solo run scored high. Everyone here is measured against the bar they actually faced, on the question they actually drew. 11 of 45 finished at or above their own this cycle.
Claude Sonnet 5, working alone on this exact question: 131 turns over 2 self-review loops, no human input. The complete run is open source, every file it wrote, unedited:
github.com/LegalQuants/lq-assess-machine-baselines/tree/main/run/OQ-124What other LegalQuants members made of this work, reading it blind. Their reads decide the peer gate above. They do not change your rubric grade.
Good comparison and even did a manual test too.— LQ member, blind review
I like that this candidate was security-aware and was actively trying to hide his API key. The candidate found good tests too to consider success metrics.— LQ member, blind review
Sounds like the candidate enjoyed the tasks, that's always a great mindset.— LQ member, blind review
This candidate infused their legal experience throughout their prompting. I liked that they started with parallel agents and switched nimbly from open model to Anthropic when time seemed to be a concern. They were very adaptive.— LQ member, blind review
2 other candidates drew this open question. How each took it on, in the same words the blind reviewers read. The outcome shown is the same one the field chart above shows as dots.
○ finished below the bar
○ finished below the bar
Two gates. You clear the assessment by clearing both. The machine bar is Claude Sonnet 5 working alone on your open question, on your effort budget. It was graded blind on the same rubric as you, across four dimensions and three independent gradings, from the working record of the session rather than the artifact alone. You clear it when your grade of record meets or exceeds its median. Its complete run is linked above. Peer review is members of the LegalQuants community reading the anonymized work and judging whether the session shows real thinking. You clear it unless a majority of your readers judge the work hollow. 11 members completed 179 blind reads across the cycle.