Skip to main content

What you get

Every paper scored against the same rubric, with the reason for each score and the evidence behind it — so two people asking the same question a month apart get comparable answers. This is what our highest-volume self-serve user does: a biotech founder running a custom evidence-grading prompt over hundreds of searches.

Who it’s for

Any team with an internal standard for evidence quality — biotech, clinical, nutrition, or research operations — where “is this well-supported?” currently depends on who you ask.

The prompt

Getting a rubric worth applying

1

Separate the axes

Design, execution, and relevance fail independently. A single 1–5 “quality” score hides which one is the problem and makes the grades incomparable across topics.
2

Anchor each level with an example

“Well-controlled” means different things to different readers. Naming a real paper at each level makes the scale reproducible.
3

Prohibit prestige as a proxy

Journal and author reputation are already available as sjr_best_quartile and citation_count. Keep them as separate metadata, not smuggled into the quality score.
4

Make NOT REPORTED a real value

Papers that omit a detail are common. Scoring them at the midpoint quietly inflates the evidence base.

Feeding it good candidates

Paid plans return study_type and takeaway on each result, which give the grader a head start on axis A. include_full_text_chunks=true is what makes axis B possible at all — execution quality is rarely visible in an abstract.

Audit requirements

Follow the grounding rules. Record the rubric version alongside every graded set: a rubric edit makes old grades incomparable, and teams discover this months later when two reports disagree.

What to check before you trust it

  • Check the score distribution. If nothing scores low, the rubric is being applied generously — real literature has a spread.
  • Grade five papers by hand first and compare. Disagreements almost always point at an ambiguous rubric level, not a model failure.
  • Watch for axis contamination. If design and execution scores move together on every paper, they are not being assessed separately.
  • Execution scores from abstracts are unreliable. Without full text, mark axis B as low-confidence rather than guessing.

Create a biomarker evidence report

Tiered grading applied to one biomarker question.

Check whether a finding has human evidence yet

Split preclinical from human evidence before grading it.

Best practices

Grounding rules and the five primitives.

All use cases

Browse the gallery by persona.