How we rate AI tools

Last updated

Every tool in this directory carries a single score out of 10. That number is not a vibe, and it is not supplied by the vendor. It is a weighted sum of six criteria, applied the same way to all 264 tools, and every tool page shows the full breakdown so you can disagree with our weighting and re-rank in your head.

The six criteria and their weights

Criterion Weight What we look at
Capability 30% Does the core job well and completely: output quality, feature depth, how much of the workflow it covers without a second tool.
Value for money 25% Price against what you actually get, judged inside its own niche. A $20/mo tool competing with $8/mo rivals has to earn the gap. Free-tier caps count here.
Ease of use 15% Time from signup to first useful output. Onboarding, defaults, whether the interface hides the important settings.
Maturity and reliability 15% Years shipping, outage and data-loss history, whether pricing and limits have been changed abruptly on paying users.
Support and docs 10% Public documentation, API reference quality, response channels, whether support is gated behind the top plan.
Momentum 5% Shipping pace over the last 12 months. Deliberately the smallest weight, a busy changelog is not a working product.

The weighted total is rounded to one decimal. A tool scoring 7.5 means the weighted sum landed between 7.45 and 7.54, not that it is "a 7.5-class product".

What the score is not

It is not a popularity ranking. Traffic, funding, and social following carry zero weight. Several tools you have never heard of score above household names, because the score answers "is this good at the job for the money" and nothing else.

It is not comparable across niches. An 8.1 in AI meeting notes and an 8.1 in AI image generation were scored against different competitive sets. Compare within a category, not across the directory.

It is not a review-aggregate. We do not average G2 or Capterra stars. Those pools are heavily shaped by vendor review campaigns. We read user complaints as evidence for specific criteria (usually reliability and value), and when we cite a complaint we link where we found it.

It cannot be bought. There is no paid placement, no sponsored slot, and no pay-to-be-reviewed path. See our editorial policy.

How we collect pricing

Prices come from the vendor's own pricing page, read on the date recorded in the Verified stamp on each tool page, not from press coverage, not from a comparison site, not from our memory of what it used to cost.

For each tool we record the entry paid price, the billing basis (per seat or flat, monthly or annual), and the concrete free-tier caps, minutes, credits, seats, exports, whichever the vendor actually meters. Where a vendor shows only "Contact us", we write not disclosed and leave it at that rather than estimating.

A second pass re-checks the headline price and the free-tier limit of each tool against the source page independently, because a mis-stated price is the one error that makes a directory worthless.

Assessment is desk research, not lab testing

We are explicit about this: scores are assigned from public evidence, vendor documentation, pricing pages, changelogs, API references, trial-tier use, and sourced user reports. We do not run controlled benchmarks across 264 products, and any site claiming it does at that scale is describing something it did not do.

What that means in practice: Capability and Ease of use are the criteria where our judgement carries the most uncertainty, and where you should weigh your own trial above our number. Value for money, Maturity, and Support are grounded in verifiable public facts.

Re-checking and corrections

Pricing moves constantly in this category. Every tool carries a Verified month, listed pages are re-checked on a rolling schedule, and a dead product is marked deprecated with a note rather than quietly deleted, the page stays useful to anyone searching for what replaced it.

If a number here is wrong, tell us and we will fix it and re-stamp the date: request a correction.