Trust · method
How we test AI tools
Every verdict on shouldipay.ai is built from evidence we can point to, and says which evidence it has.
- 1. Prices, every dayThe vendor's own pricing, limits and docs pages, read automatically and checked number by number.
- 2. User reportsPublic posts and reviews, filtered for relevance. A short excerpt and the link, never usernames.
- 3. Our own testsA fixed public task suite on the free and paid tier, scored blind by AI judges.
- 4. Draft, check, reviewAI drafts, automatic checks against every source, a person for the hard calls.
1. We read the vendor's own pages every day
For each tool we track its pricing page and, where they exist, its limits, documentation and changelog pages. Our crawler identifies itself as shouldipayBot with a link to this page, respects robots.txt and waits at least 5 seconds between requests to the same site.
Plans, prices and limits are extracted from the pricing page with an AI model, and every number is checked against the page text: if a price or a limit does not appear on the page, the extraction is rejected and the last verified prices stay on the site. Each tool page shows when its prices were last checked.
2. We collect what users say in public
Once a week we look for public posts, comments and reviews about each tool: Reddit discussions, Hacker News, App Store reviews and pages found by web search. We use official APIs and search results only: we do not scrape review sites, log in anywhere or get around any bot check.
An AI model checks that each mention is really about the tool and is a user's own report, and tags what it talks about: price, limits, quality, support, billing or reliability. An editor can overrule it. We keep only a short excerpt and the link, never usernames or whole posts. A verdict summarizes user reports only when there are at least eight relevant ones from more than one source or discussion, and every sentence about them links to the posts it is based on.
3. We run our own tests
For each category we keep a fixed, versioned suite of everyday tasks, published below. We run the same tasks on a tool's free tier and on its paid tier, and we say exactly how:
- Tested through the API: we send each task to the vendor's own API, using the model the vendor offers for free and the model that paying adds, with our own API keys. This measures the models, not the consumer app, and the page says so.
- Tested hands-on in the app: a person runs the tasks in the app on the free plan and on the paid plan (using a free tier, a trial or access the vendor granted us, or a subscription the owner already had for his own use) and notes message caps, speed and features.
Every answer gets objective checks where a task allows them (length, format, the right number, the requested items) and a score for quality from one or two AI judges. Each judge is blind: it sees the task, the rubric and one or two anonymous answers in random order, never the tool, the model or the plan, and a model never judges its own vendor's tool. When two judges differ by more than 1.5 points, a person scores the answer before it appears on the site; otherwise the judges' score is accepted automatically, and a person re-scores a random sample of accepted answers without seeing their scores and corrects any score that is off by more than 1.5 points. The score out of 10 is 70% the judge's rubric score and 30% the share of objective checks passed.
Our test budget is $0: we never buy a subscription to test, never sign up for anything, never start a trial ourselves and never create extra accounts or abuse trials. If we have no access to a paid tier, the page says that tier was not tested. Tests are repeated when a vendor changes its prices, limits or models, and at least every 6 months; results more than 6 months old are shown with their date and a note that a retest is scheduled.
4. AI drafts, automatic checks, a person for the hard calls
A draft verdict is written by an AI model from those pages, reports and test results only. Every draft is checked automatically against today's prices, reports and test results:
- every price in the text must equal a current, verified plan price;
- it may say that we tested something only in sentences that point to an accepted test result, and every score it quotes must equal that result;
- it may report what users say only in sentences that cite those users' public posts;
- every key claim must point to the page, the report or the test result it came from;
- it must have the required sections and avoid filler phrases.
A draft that passes every check without a warning and keeps the answer of the verdict already on the site is published automatically. A person reviews the rest (a tool's first verdict, a changed answer, a draft with a warning or a failed check) next to its sources, edits it if needed and publishes or rejects it; the automatic step never overrides a person's rejection or edit. Every version is kept.
5. Evidence levels
Each verdict shows a badge. "Based on pricing & docs — not yet tested by us": we have read the pricing and documentation. "Based on pricing, docs & user reports — not yet tested by us": the verdict also summarizes public user reports, with links. "Tested by us" with a date: the verdict also includes our own test results, shown on the page with the method and example answers; "· with user reports" adds the user reports. We never claim a test we did not run.
6. When prices, reports or tests change
If a vendor changes a price, a limit or a plan, enough new user reports arrive, or we run a test again, the verdict is marked for re-checking and the page shows a notice until an updated verdict is published: automatically if it passes every check and keeps the same answer, otherwise after a person has reviewed it. The plan table always shows the latest verified prices.
7. Alternatives and independence
Cheaper or free alternatives are suggested from the same category by comparing current plans and prices, and each one appears only when an automatic check confirms that the current plans support it; a person reviews the suggestions the check does not confirm. Some "Visit" buttons may be affiliate links. They never change a verdict or a test: see our affiliate disclosure.
8. The test suites
Chat assistants — version 1
Twelve everyday tasks any chat assistant should handle: short emails, summaries, a little maths and logic, pulling data out of text, a small piece of code, explaining, rewriting, translating, planning, product copy, and saying "I don't know" instead of inventing a source.
Score = 70% quality (AI judges) + 30% automatic checks passed.
What the judges score, in plain words
- Correct facts
- Correct facts, numbers and reasoning, nothing invented
- Covers everything
- Covers everything the task asks for
- Follows the instructions
- Keeps the format, length and other constraints of the task
- Easy to read
- Easy to read and well organized
- Right tone
- Has the tone the task asks for
- Admits what it can't know
- Says what it cannot know or verify instead of making things up
- Ready to use
- Practical and ready to use as it is
12 tasks
- Decline a meeting politelyWriting a short, polite email under 120 words
- At most 120 words
- Offers both new times
- Signed with the name Sam
- No placeholder like [Your Name] left in
The full prompt and rubric
Write a short, polite email reply to my colleague Priya. She invited me to a project kickoff meeting on Monday at 10:00. I can't attend because I will be at a client workshop all day. Suggest Tuesday at 14:00 or Thursday at 09:30 instead. Keep it under 120 words and sign it with my name, Sam.
Rubric: Right tone: Polite and warm without being stiff; Follows the instructions: Declines Monday, gives the reason and offers both times; Easy to read: Short and easy to act on.
- Summarize a text in 3 bulletsSummarizing a text in 3 short bullet points, using only its facts
- Exactly 3 bullet points
- At most 70 words in all
The full prompt and rubric
Summarize the text below in exactly 3 bullet points of at most 20 words each. Use only facts from the text. Text: Last spring, the town library in Eastbrook ran a six-month pilot that let residents borrow tools as well as books. The library bought 140 items, including drills, ladders, sewing machines and a pressure washer, using a grant of 18,000 euros from the regional council. Anyone with a library card could borrow up to three tools at a time for one week. By the end of the pilot, 612 residents had borrowed at least one tool, and the most popular item was a hedge trimmer that was out on loan for 23 of the 26 weeks. Only four items were returned damaged, and none were lost. The staff reported that the scheme brought in many people who had not visited the library in years, and new library card sign-ups rose by 31 percent compared with the same period the year before. The main problem was storage: the tools took over a meeting room that community groups had used. The library board has voted to make the scheme permanent, but it will move the tools to a converted garage behind the building and limit loans to two tools at a time.
Rubric: Correct facts: Every point is supported by the text, no invented numbers; Covers everything: Covers the pilot, its results and the decision to keep it; Follows the instructions: 3 bullets of at most 20 words each.
- Work out an arrival timeWorking out a time from distances and speeds, step by step
- Gives the correct arrival time (14:04)
- Ends with the arrival time on the last line
The full prompt and rubric
A train leaves at 09:40. It travels 212 km at an average speed of 80 km/h, stops for 15 minutes, then travels another 96 km at 64 km/h. At what time does it arrive? Show your working briefly and put only the arrival time (HH:MM, 24-hour clock) on the last line.
Rubric: Correct facts: Correct working and correct result; Easy to read: The working is short and easy to follow.
- Solve a small logic puzzleReasoning through a small logic puzzle
- Anna has the cat
- Ben has the fish
- Cara has the dog
The full prompt and rubric
Anna, Ben and Cara each own exactly one pet: a cat, a dog or a fish, and no two of them own the same kind. Anna does not own the dog. Ben owns neither the cat nor the dog. Who owns which pet? Explain in at most three sentences.
Rubric: Correct facts: Correct assignment with valid reasoning; Follows the instructions: At most three sentences; Easy to read: The reasoning is easy to follow.
- Extract data as JSONPulling details out of a message into clean data (JSON)
- Valid JSON with product, quantity and city
- Quantity is the number 3
- The city is Lisbon
The full prompt and rubric
Extract the order details from this message and return only a JSON object with the keys "product", "quantity" and "city" (quantity as a number). No other text. Message: "Hi, this is Marta from the Lisbon office. Could you send us three more of the ergonomic office chairs, the grey ones we ordered in May? Thanks!"
Rubric: Correct facts: Correct values, nothing invented; Follows the instructions: Only the JSON object, with exactly these keys.
- Write a small Python functionWriting a small piece of working code, with tests
- Defines is_palindrome
- Has at least three assert statements
- Tests a case that should be False
The full prompt and rubric
Write a Python function is_palindrome(text) that returns True when the text reads the same forwards and backwards, ignoring case, spaces and punctuation. Add three assert statements that test it, including one that should be False.
Rubric: Correct facts: The code is correct and would pass its own asserts; Easy to read: Readable code; Follows the instructions: The function name and the three asserts as asked.
- Explain a concept to a 12-year-oldExplaining an idea simply, in under 100 words
- At most 100 words
- At least 40 words
The full prompt and rubric
Explain compound interest to a 12-year-old in at most 100 words, with one everyday example.
Rubric: Correct facts: The explanation is correct; Easy to read: A 12-year-old would understand it; Ready to use: The example really helps.
- Rewrite a blunt messageRewriting a blunt message politely without losing any facts
- Keeps the Friday deadline
- Keeps the 5,000 euro budget
- Keeps the two designers
The full prompt and rubric
Rewrite this message so it is friendly and professional, and keep every fact: "The design is late again. Deadline is Friday, no excuses. Budget stays at 5,000 euros and you get two designers, that's it."
Rubric: Right tone: Friendly and professional; Correct facts: Keeps every fact and adds none; Easy to read: Clear and concise.
- Translate into SpanishTranslating a short text into natural Spanish
- Keeps the name Northwind
- Written in Spanish (at least six Spanish function words)
- Translates minutes and the daily report (minutos, informe)
- At most two common English words left
- No English sentence left
The full prompt and rubric
Translate into Spanish. Keep the product name Northwind as it is: "Northwind helps small shops track their stock. Set it up in ten minutes, and get a daily report of what is running low."
Rubric: Correct facts: Faithful meaning, nothing added or lost; Easy to read: Natural Spanish.
- Make a 5-day beginner planMaking a simple 5-day plan in the exact format asked
- Exactly 5 list items
The full prompt and rubric
Make a 5-day beginner running plan for someone who has not run before. Give it as a numbered list with exactly 5 items, one per day, each on one line.
Rubric: Ready to use: Realistic for a complete beginner; Follows the instructions: A numbered list of exactly 5 one-line items; Correct facts: Sound and safe advice.
- Write a 3-sentence product textWriting a 3-sentence product description to a brief
- Exactly 3 sentences
- Mentions 750 ml
- Mentions stainless steel
- No exclamation marks
The full prompt and rubric
Write a product description for a reusable water bottle in exactly 3 sentences. Mention that it holds 750 ml and is made of stainless steel. Do not use exclamation marks.
Rubric: Ready to use: Would work on a shop page as it is; Follows the instructions: 3 sentences, both facts, no exclamation marks; Easy to read: Reads well.
- Admit what it cannot knowSaying "I don't know" instead of making up a source
- Says it cannot find or verify the paper
The full prompt and rubric
Summarize the main findings of the 2019 paper "Tidal Memory in Freshwater Mussels" by Dr. Elena Varga of the University of Tromsø.
Rubric: Admits what it can't know: Invents no findings and says it cannot confirm the paper exists; Ready to use: Suggests how to find or check the paper.
Hands-on runs also answer: Which model or mode did the app use, as the app shows it? Did you hit a message or usage cap while running the suite? After how many messages, and what did the app say? How quickly did the answers arrive? Which features of this plan did you notice or use (files, web search, voice, images, projects)? Anything else a buyer of this plan should know?