How We Review AI Tools
Transparency matters. Here's exactly how we evaluate, score, and recommend the AI tools on our site.
Our Review Process
1. Research
Before testing, we research the tool's background, funding, user reviews, competitors, and market position. We read documentation, community forums, and existing reviews to understand the full picture.
2. Hands-On Testing
We sign up and use the tool on real tasks — not synthetic benchmarks. We test free tiers, paid plans, and edge cases. We evaluate the actual output quality, not just the feature list.
3. Comparison
Every tool is evaluated in the context of its competitors. We compare pricing, features, output quality, and ease of use against similar tools in the same category.
4. Scoring
We assign a score out of 5 based on the criteria below. Scores reflect overall value — a tool with a lower price and solid performance can score higher than an expensive tool with marginal improvements.
Scoring Criteria
Each tool is scored on a 0–5 scale across these factors:
| Factor | Weight | What We Evaluate |
|---|---|---|
| Output Quality | 30% | Accuracy, usefulness, and readiness of the AI output |
| Ease of Use | 20% | Interface design, learning curve, and onboarding experience |
| Value for Money | 25% | Pricing, free tiers, and ROI compared to alternatives |
| Features | 15% | Breadth and depth of features, integrations, and flexibility |
| Support & Docs | 10% | Documentation quality, customer support, and community |
Editorial Independence
Our reviews are editorially independent. While we earn affiliate commissions on some tools we recommend, this never influences our scores or opinions. Here's what that means in practice:
- We give low scores to tools with affiliate programs if they deserve it.
- We recommend free tools over paid ones when they're genuinely better.
- We clearly disclose affiliate relationships on every review page.
- No tool company has editorial input or approval over our content.
Keeping Reviews Updated
AI tools evolve quickly. We revisit and update our reviews when tools release major updates, change pricing, or add significant new features. Every review shows a “Last updated” date so you know how current the information is. If you spot something outdated, please let us know.
Scoring in Practice: Pictory Review Example
To make our methodology concrete, here's how we scored Pictory across each criterion during our 14-day hands-on test:
| Criterion | Weight | Score | Notes |
|---|---|---|---|
| Output Quality | 30% | 3.5/5 | Blog-to-video conversion works but AI voices need improvement |
| Ease of Use | 20% | 4.0/5 | No editing skills needed; intuitive interface |
| Value for Money | 25% | 3.5/5 | $19/mo starter plan is fair for the output quality |
| Features | 15% | 3.0/5 | Limited customization; no green screen or webcam |
| Support & Docs | 10% | 3.5/5 | Good documentation but slow email support |
| Final Score | 3.5/5 |
Every review on this site follows this same scoring methodology. You can see the breakdown in each review's verdict section.
See Our Reviews
Here are a few examples of our methodology in action:
Category-Specific Test Protocols
“We tested it” means nothing unless the test is the same for every tool in a category. These are the fixed protocols we run (or will run before scoring a tool in that category):
AI Voice
Identical scripts through each tool: a 60-second ad read, a 5-minute narration, and a technical passage with brand names. We compare pronunciation, pacing, emotional delivery, long-form consistency, the correction burden (retakes needed), and cost per usable minute.
AI Video & Avatars
The same script and prompt set per tool. We assess realism and lip-sync, consistency across renders, visible artifacts, rendering time, how much manual editing the output needs, and the usable-output rate — how many generations we would actually publish.
Automation Platforms
We build the identical multi-step workflow (trigger → transform → AI step → notify) on each platform and measure build time, number of steps/operations consumed, debugging experience, failure behavior, execution cost at 1,000 runs, and maintainability a month later.
AI Coding Tools
Identical tasks on the same repository: a scoped bug fix, a multi-file refactor, and a test-covered feature. We measure completion, files modified, errors introduced, human interventions required, wall-clock time, and plan usage consumed.
SEO & GEO Tools
The same website, the same article, the same query and prompt set, and the same competitor list across tools. For GEO monitoring we compare which brand mentions and citations each tool detects for an identical prompt set over the same period. Ranking outcomes have long feedback loops, and we say so rather than claim day-one causality.
Researched Profiles vs Hands-On Reviews
Not every tool page on ShelbyAI is a hands-on review, and we label the difference explicitly. A hands-on review carries a score, a tested date, and findings from the protocols above. A researched profile carries verified facts — pricing, plans, features from official sources with a fact-checked date — but no score and no testing claims, and a visible notice that our test is pending.
Two dates matter and they are not the same thing: last fact-checked (we re-verified prices and specs against the vendor) and last hands-on tested (we actually ran the tool through its protocol). Pricing pages additionally show a pricing verified date linked to the vendor source — see the AI Pricing Index for every tool's current verification status. We never publish invented scores, fabricated benchmarks, or placeholder test results.