Independent verdicts on paid AI plans. Prices re-checked every day. We may earn a commission from "Visit" links. How we make money
shouldipay.ai

Trust · method

How we test AI tools

Every verdict on shouldipay.ai is built from evidence we can point to, and says which evidence it has.

  1. 1. Prices, every dayThe vendor's own pricing, limits and docs pages, read automatically and checked number by number.
  2. 2. User reportsPublic posts and reviews, filtered for relevance. A short excerpt and the link, never usernames.
  3. 3. Our own testsA fixed public task suite on the free and paid tier, scored blind by AI judges.
  4. 4. Draft, check, reviewAI drafts, automatic checks against every source, a person for the hard calls.

1. We read the vendor's own pages every day

For each tool we track its pricing page and, where they exist, its limits, documentation and changelog pages. Our crawler identifies itself as shouldipayBot with a link to this page, respects robots.txt and waits at least 5 seconds between requests to the same site.

Plans, prices and limits are extracted from the pricing page with an AI model, and every number is checked against the page text: if a price or a limit does not appear on the page, the extraction is rejected and the last verified prices stay on the site. Each tool page shows when its prices were last checked.

2. We collect what users say in public

Once a week we look for public posts, comments and reviews about each tool: Reddit discussions, Hacker News, App Store reviews and pages found by web search. We use official APIs and search results only: we do not scrape review sites, log in anywhere or get around any bot check.

An AI model checks that each mention is really about the tool and is a user's own report, and tags what it talks about: price, limits, quality, support, billing or reliability. An editor can overrule it. We keep only a short excerpt and the link, never usernames or whole posts. A verdict summarizes user reports only when there are at least eight relevant ones from more than one source or discussion, and every sentence about them links to the posts it is based on.

3. We run our own tests

For each category we keep a fixed, versioned suite of everyday tasks, published below. We run the same tasks on a tool's free tier and on its paid tier, and we say exactly how:

Every answer gets objective checks where a task allows them (length, format, the right number, the requested items) and a score for quality from one or two AI judges. Each judge is blind: it sees the task, the rubric and one or two anonymous answers in random order, never the tool, the model or the plan, and a model never judges its own vendor's tool. When two judges differ by more than 1.5 points, a person scores the answer before it appears on the site; otherwise the judges' score is accepted automatically, and a person re-scores a random sample of accepted answers without seeing their scores and corrects any score that is off by more than 1.5 points. The score out of 10 is 70% the judge's rubric score and 30% the share of objective checks passed.

Our test budget is $0: we never buy a subscription to test, never sign up for anything, never start a trial ourselves and never create extra accounts or abuse trials. If we have no access to a paid tier, the page says that tier was not tested. Tests are repeated when a vendor changes its prices, limits or models, and at least every 6 months; results more than 6 months old are shown with their date and a note that a retest is scheduled.

The score out of 10 is 70% the judges' rubric score and 30% the share of automatic checks passed. Example: an answer the judges rate 8/10 that passes 1 of 2 checks scores 7.1.

4. AI drafts, automatic checks, a person for the hard calls

A draft verdict is written by an AI model from those pages, reports and test results only. Every draft is checked automatically against today's prices, reports and test results:

A draft that passes every check without a warning and keeps the answer of the verdict already on the site is published automatically. A person reviews the rest (a tool's first verdict, a changed answer, a draft with a warning or a failed check) next to its sources, edits it if needed and publishes or rejects it; the automatic step never overrides a person's rejection or edit. Every version is kept.

5. Evidence levels

Each verdict shows a badge. "Based on pricing & docs — not yet tested by us": we have read the pricing and documentation. "Based on pricing, docs & user reports — not yet tested by us": the verdict also summarizes public user reports, with links. "Tested by us" with a date: the verdict also includes our own test results, shown on the page with the method and example answers; "· with user reports" adds the user reports. We never claim a test we did not run.

6. When prices, reports or tests change

If a vendor changes a price, a limit or a plan, enough new user reports arrive, or we run a test again, the verdict is marked for re-checking and the page shows a notice until an updated verdict is published: automatically if it passes every check and keeps the same answer, otherwise after a person has reviewed it. The plan table always shows the latest verified prices.

7. Alternatives and independence

Cheaper or free alternatives are suggested from the same category by comparing current plans and prices, and each one appears only when an automatic check confirms that the current plans support it; a person reviews the suggestions the check does not confirm. Some "Visit" buttons may be affiliate links. They never change a verdict or a test: see our affiliate disclosure.

8. The test suites

Chat assistants — version 1

Twelve everyday tasks any chat assistant should handle: short emails, summaries, a little maths and logic, pulling data out of text, a small piece of code, explaining, rewriting, translating, planning, product copy, and saying "I don't know" instead of inventing a source.

Score = 70% quality (AI judges) + 30% automatic checks passed.

What the judges score, in plain words

Correct facts
Correct facts, numbers and reasoning, nothing invented
Covers everything
Covers everything the task asks for
Follows the instructions
Keeps the format, length and other constraints of the task
Easy to read
Easy to read and well organized
Right tone
Has the tone the task asks for
Admits what it can't know
Says what it cannot know or verify instead of making things up
Ready to use
Practical and ready to use as it is

12 tasks

  1. Decline a meeting politelyWriting a short, polite email under 120 words
    • At most 120 words
    • Offers both new times
    • Signed with the name Sam
    • No placeholder like [Your Name] left in
    The full prompt and rubric
    Write a short, polite email reply to my colleague Priya. She invited me to a project kickoff meeting on Monday at 10:00. I can't attend because I will be at a client workshop all day. Suggest Tuesday at 14:00 or Thursday at 09:30 instead. Keep it under 120 words and sign it with my name, Sam.

    Rubric: Right tone: Polite and warm without being stiff; Follows the instructions: Declines Monday, gives the reason and offers both times; Easy to read: Short and easy to act on.

  2. Summarize a text in 3 bulletsSummarizing a text in 3 short bullet points, using only its facts
    • Exactly 3 bullet points
    • At most 70 words in all
    The full prompt and rubric
    Summarize the text below in exactly 3 bullet points of at most 20 words each. Use only facts from the text.
    
    Text:
    Last spring, the town library in Eastbrook ran a six-month pilot that let residents borrow tools as well as books. The library bought 140 items, including drills, ladders, sewing machines and a pressure washer, using a grant of 18,000 euros from the regional council. Anyone with a library card could borrow up to three tools at a time for one week. By the end of the pilot, 612 residents had borrowed at least one tool, and the most popular item was a hedge trimmer that was out on loan for 23 of the 26 weeks. Only four items were returned damaged, and none were lost. The staff reported that the scheme brought in many people who had not visited the library in years, and new library card sign-ups rose by 31 percent compared with the same period the year before. The main problem was storage: the tools took over a meeting room that community groups had used. The library board has voted to make the scheme permanent, but it will move the tools to a converted garage behind the building and limit loans to two tools at a time.

    Rubric: Correct facts: Every point is supported by the text, no invented numbers; Covers everything: Covers the pilot, its results and the decision to keep it; Follows the instructions: 3 bullets of at most 20 words each.

  3. Work out an arrival timeWorking out a time from distances and speeds, step by step
    • Gives the correct arrival time (14:04)
    • Ends with the arrival time on the last line
    The full prompt and rubric
    A train leaves at 09:40. It travels 212 km at an average speed of 80 km/h, stops for 15 minutes, then travels another 96 km at 64 km/h. At what time does it arrive? Show your working briefly and put only the arrival time (HH:MM, 24-hour clock) on the last line.

    Rubric: Correct facts: Correct working and correct result; Easy to read: The working is short and easy to follow.

  4. Solve a small logic puzzleReasoning through a small logic puzzle
    • Anna has the cat
    • Ben has the fish
    • Cara has the dog
    The full prompt and rubric
    Anna, Ben and Cara each own exactly one pet: a cat, a dog or a fish, and no two of them own the same kind. Anna does not own the dog. Ben owns neither the cat nor the dog. Who owns which pet? Explain in at most three sentences.

    Rubric: Correct facts: Correct assignment with valid reasoning; Follows the instructions: At most three sentences; Easy to read: The reasoning is easy to follow.

  5. Extract data as JSONPulling details out of a message into clean data (JSON)
    • Valid JSON with product, quantity and city
    • Quantity is the number 3
    • The city is Lisbon
    The full prompt and rubric
    Extract the order details from this message and return only a JSON object with the keys "product", "quantity" and "city" (quantity as a number). No other text.
    
    Message: "Hi, this is Marta from the Lisbon office. Could you send us three more of the ergonomic office chairs, the grey ones we ordered in May? Thanks!"

    Rubric: Correct facts: Correct values, nothing invented; Follows the instructions: Only the JSON object, with exactly these keys.

  6. Write a small Python functionWriting a small piece of working code, with tests
    • Defines is_palindrome
    • Has at least three assert statements
    • Tests a case that should be False
    The full prompt and rubric
    Write a Python function is_palindrome(text) that returns True when the text reads the same forwards and backwards, ignoring case, spaces and punctuation. Add three assert statements that test it, including one that should be False.

    Rubric: Correct facts: The code is correct and would pass its own asserts; Easy to read: Readable code; Follows the instructions: The function name and the three asserts as asked.

  7. Explain a concept to a 12-year-oldExplaining an idea simply, in under 100 words
    • At most 100 words
    • At least 40 words
    The full prompt and rubric
    Explain compound interest to a 12-year-old in at most 100 words, with one everyday example.

    Rubric: Correct facts: The explanation is correct; Easy to read: A 12-year-old would understand it; Ready to use: The example really helps.

  8. Rewrite a blunt messageRewriting a blunt message politely without losing any facts
    • Keeps the Friday deadline
    • Keeps the 5,000 euro budget
    • Keeps the two designers
    The full prompt and rubric
    Rewrite this message so it is friendly and professional, and keep every fact:
    
    "The design is late again. Deadline is Friday, no excuses. Budget stays at 5,000 euros and you get two designers, that's it."

    Rubric: Right tone: Friendly and professional; Correct facts: Keeps every fact and adds none; Easy to read: Clear and concise.

  9. Translate into SpanishTranslating a short text into natural Spanish
    • Keeps the name Northwind
    • Written in Spanish (at least six Spanish function words)
    • Translates minutes and the daily report (minutos, informe)
    • At most two common English words left
    • No English sentence left
    The full prompt and rubric
    Translate into Spanish. Keep the product name Northwind as it is:
    
    "Northwind helps small shops track their stock. Set it up in ten minutes, and get a daily report of what is running low."

    Rubric: Correct facts: Faithful meaning, nothing added or lost; Easy to read: Natural Spanish.

  10. Make a 5-day beginner planMaking a simple 5-day plan in the exact format asked
    • Exactly 5 list items
    The full prompt and rubric
    Make a 5-day beginner running plan for someone who has not run before. Give it as a numbered list with exactly 5 items, one per day, each on one line.

    Rubric: Ready to use: Realistic for a complete beginner; Follows the instructions: A numbered list of exactly 5 one-line items; Correct facts: Sound and safe advice.

  11. Write a 3-sentence product textWriting a 3-sentence product description to a brief
    • Exactly 3 sentences
    • Mentions 750 ml
    • Mentions stainless steel
    • No exclamation marks
    The full prompt and rubric
    Write a product description for a reusable water bottle in exactly 3 sentences. Mention that it holds 750 ml and is made of stainless steel. Do not use exclamation marks.

    Rubric: Ready to use: Would work on a shop page as it is; Follows the instructions: 3 sentences, both facts, no exclamation marks; Easy to read: Reads well.

  12. Admit what it cannot knowSaying "I don't know" instead of making up a source
    • Says it cannot find or verify the paper
    The full prompt and rubric
    Summarize the main findings of the 2019 paper "Tidal Memory in Freshwater Mussels" by Dr. Elena Varga of the University of Tromsø.

    Rubric: Admits what it can't know: Invents no findings and says it cannot confirm the paper exists; Ready to use: Suggests how to find or check the paper.

Hands-on runs also answer: Which model or mode did the app use, as the app shows it? Did you hit a message or usage cap while running the suite? After how many messages, and what did the app say? How quickly did the answers arrive? Which features of this plan did you notice or use (files, web search, voice, images, projects)? Anything else a buyer of this plan should know?