What you get
Every paper scored against the same rubric, with the reason for each score and the evidence behind it — so two people asking the same question a month apart get comparable answers. This is what our highest-volume self-serve user does: a biotech founder running a custom evidence-grading prompt over hundreds of searches.Who it’s for
Any team with an internal standard for evidence quality — biotech, clinical, nutrition, or research operations — where “is this well-supported?” currently depends on who you ask.The prompt
Getting a rubric worth applying
1
Separate the axes
Design, execution, and relevance fail independently. A single 1–5 “quality” score hides which one is the problem and makes the grades incomparable across topics.
2
Anchor each level with an example
“Well-controlled” means different things to different readers. Naming a real paper at each level makes the scale reproducible.
3
Prohibit prestige as a proxy
Journal and author reputation are already available as
sjr_best_quartile and citation_count. Keep them as separate metadata, not smuggled into the quality score.4
Make NOT REPORTED a real value
Papers that omit a detail are common. Scoring them at the midpoint quietly inflates the evidence base.
Feeding it good candidates
study_type and takeaway on each result, which give the grader a head start on axis A. include_full_text_chunks=true is what makes axis B possible at all — execution quality is rarely visible in an abstract.
Audit requirements
Follow the grounding rules. Record the rubric version alongside every graded set: a rubric edit makes old grades incomparable, and teams discover this months later when two reports disagree.What to check before you trust it
- Check the score distribution. If nothing scores low, the rubric is being applied generously — real literature has a spread.
- Grade five papers by hand first and compare. Disagreements almost always point at an ambiguous rubric level, not a model failure.
- Watch for axis contamination. If design and execution scores move together on every paper, they are not being assessed separately.
- Execution scores from abstracts are unreliable. Without full text, mark axis B as low-confidence rather than guessing.
Related
Create a biomarker evidence report
Tiered grading applied to one biomarker question.
Check whether a finding has human evidence yet
Split preclinical from human evidence before grading it.
Best practices
Grounding rules and the five primitives.
All use cases
Browse the gallery by persona.