Method
How the scores work
A tool that asks you to trust a number owes you the formula. This is the one actually running — the weights below are read out of the code, not written by hand for this page.
Source score
A 0–100 measure of the evidence base, from five inputs:
- Freshness (35%) — how old the evidence is, by publication date where it is known. Under 14 days scores 95, under 60 days 78, under 180 days 58, older 32. A source with no publication date is capped at 70 and labelled undated, because re-fetching a six-year-old page does not make its evidence recent.
- Uniqueness (25%) — the share of sources that are not duplicates of each other, matched on a normalised URL.
- Extraction (20%) — the share with text actually pulled from them, rather than a bare link.
- Independence (15%) — how far each source is from the claim it supports. Independent reporting or a primary record counts full; a competitor comparison 0.6; a vendor writing about its own category 0.45; a recollection 0.2. What you type is checked against the address. A DOI, an arXiv identifier or a .gov domain is a primary record whatever the label says; a Reddit or Hacker News thread is a recollection whatever the label says. Where the two disagree the weaker one is used — you can downgrade a source, because you may know something the domain does not, but you cannot upgrade one.
- Domain diversity (5%) — how many distinct sites the evidence comes from. Normalised-URL matching cannot tell syndicated copy apart; three domains saying a thing beats one wire story on three URLs.
Penalties: −12 with fewer than three sources, −10 if nothing has been extracted from any of them.
Sources that disagree block publication outright. Figures in the sources are compared with each other — percentages, counts and amounts — and if two of them give materially different values for what looks like the same measurement, the venture is held until it is clear which is right. Two sources that contradict each other are not two sources supporting you. This is arithmetic, not judgement: it will not notice a disagreement stated only in prose.
One unsupported claim can block publication on its own
The audit score is an average, and an average will publish a venture whose entire premise is unevidenced as long as the peripheral claims are tidy. A claim marked load-bearing has to be carried on its own.
This product's own dossier is the example. It audits 74 — comfortably past the threshold — and is still held at build, because nothing on file supports the claim that founders want to be refused rather than warned. That is the premise the whole thing rests on.
Thresholds
Below the build threshold the venture is refused outright. Above build but below publish, it can be worked on but not published — the claims are not yet carried by the sources.
What this score does not measure
Worth stating plainly, because it is the honest limit of the method and it would be easy to leave unsaid.
- Whether a paraphrase is a copy. Duplicate detection now compares the text as well as the URL, so the same article republished on three domains counts once rather than three times. It matches republication, not rewriting: the threshold that would catch a reworded version also reaches genuinely unrelated articles, and wrongly telling you two independent sources are the same is worse than missing a paraphrase.
- Whether a source supports a claim — only whether it could. Two things are now checked mechanically rather than judged: a citation pointing at a source that does not exist, and a claim about something the cited source never mentions. Both block publication. Shared vocabulary is still not support, so what remains is the model's judgement, and it is wrong sometimes: on this product's own dossier it cited a paper about model sycophancy as the source for a claim about a competitor's features. The check above caught that; the model that wrote the audit did not.
- Whether a source is the right source. Independence is now weighted, but relevance is not. A first-rate independent study about a neighbouring question still counts.
- Whether the source supports the claim. That is the claim audit's job, and the audit is only as good as the model performing it.
- How good a study is — only what it says about itself. Sample size and study design are now read out of each source and shown beside it, and a source that describes itself as a recollection is scored as one whatever its label says. But the weighting is unchanged: four anecdotes still score well, because independence carries 15% and that is the whole swing. Weighting method properly needs a set of sources with known-good labels to calibrate against, and that does not exist yet — re-tuning it blind has twice broken another part of this score.
- Whether the idea is any good. A well-evidenced bad idea clears the gate. The gate checks your reasoning, not your judgement.
Independence used to be recorded and ignored. It is now weighted, which cost the demo venture on this site five points: three of its five sources are vendors writing about their own category, so it scores 84 rather than the 98 it had when quantity was all that counted.