Which kind of AI suits the job? — questions and answers

The questions a sceptical reader asks of a tool like this, answered plainly, with the real limits stated.

You build and sell AI systems. Isn't this just a way of selling them?

That is the right question to ask of anyone publishing a tool like this, and the check is whether the thing can rule us out. It can, and it does: answer that you are tidying up writing, that nothing personal goes in, and that a person reads everything before it goes anywhere, and the tool says plainly that running your own hardware would be a large cost for a risk you do not have. That case is written into the rules and tested on every build — if it ever stopped ruling us out, the build fails. What the tool cannot claim is disinterest. We build this kind of infrastructure and the argument favours it in the cases where it favours it. Read the reasoning, which is shown, rather than the intention, which you cannot check.

Six questions cannot possibly capture my situation. Isn't this too crude to be useful?

It is crude, and that is deliberate rather than a limitation we are hiding. The tool is not trying to know your situation; it is trying to make a small number of consequences explicit so that some options can be eliminated on grounds you can inspect. Elimination survives crudeness in a way recommendation does not — "this cannot meet your requirement that information stays in the country" holds whether or not we understand your organisation, because it follows from the requirement, not from you. That is also why the output shows the answer behind every deduction. If a deduction is wrong for you, it is wrong visibly, and you can change the answer that caused it.

You tell me to put a prompt to a system and read its answers, then you tell me its answers are not evidence. Which is it?

Both, and the distinction is the point. What comes back is not evidence about the system; it is a set of claims by the system, which is a different object. Some of those claims are checkable — a policy with a resolvable link either says what was claimed or it does not. Some are usefully uncheckable: "I cannot establish that about myself" is the one kind of answer that cannot be an overclaim. And the comparison across two or three products tells you more than any single answer does, because the differences are visible even when the individual claims are not verifiable. The prompt carries all of this in the copied text, addressed to you rather than to the machine, because the guidance is useless if it stays on our website while you read the answers somewhere else.

Why won't you just tell me which product to use?

Because we would be wrong within months, and wrong in a way you could not detect. Products change, terms change, and a page that recommends becomes a page that misleads without anyone editing it. There is a second reason that matters more: the right answer depends on facts about your situation that we do not have and a questionnaire cannot obtain. What we can do is show which kinds of thing your own constraints exclude, which is a deduction rather than an opinion. You leave with a shortlist you can justify to somebody else, which is more use than a recommendation you would have to take on faith.

What if the system simply tells me what I want to hear?

It may, and worse, it may do so accurately-sounding. Systems are unreliable narrators about their own architecture and deployment, without any intention to mislead. There is a sharper version of this problem that we should name rather than wait to be caught by: once a set of questions circulates widely, systems get tuned to answer it well, and a polished answer stops being a signal. That is why the prompt asks for resolvable citations you can open yourself, why it asks what the system cannot establish, and why it tells you that a perfect score is itself worth holding at arm's length. None of that makes the answers reliable. It makes them checkable, which is the most a prompt can do.

Isn't "you may not need AI at all" just false modesty?

It is a real output, reachable from real answers, and it fires on the combination that should produce it: a decision about a person, made with nobody reading the result. In that case the tool says the question to settle first is not which system to use but whether this should be automated. We would rather that outcome existed and was rarely reached than not exist at all — a tool that can only ever answer "use one of these" is a catalogue with a questionnaire on the front.

I do not want to adopt AI. I want to make sure it is NOT used. Is there anything here for me?

Not yet, and it is the gap we take most seriously. It is a different question from the one this tool asks, and a harder one. Choosing a system means comparing what things do; excluding one means establishing that something did not happen, and absence is the thing evidence is worst at. Our own corpus has the worked example: a correction published on 31 July 2026 recorded that a process validating every citation against a primary source could not catch a claim that something did not exist, because a claim of absence has no citation to validate. That is exactly the position of an editor certifying no AI was used, or a procurement officer writing a warranty. The useful answer is not a shortlist but a specification — what you can require, what a supplier can warrant, what evidence would actually support it, and where the limit sits. That is being built next. In the meantime, if the answer to your situation is "use nothing", this tool will say so, and it says so more often than a tool published by an AI company comfortably should.

Why is "nothing at all" one of the options? Nobody publishing a tool like this includes that.

Because it is often the right answer, and a set of options that cannot contain it is a catalogue. Every mainstream browser translates pages for free. A saved outline beats a prompt retyped weekly. If you are drafting something with nothing personal in it and the worst case is embarrassment, an assistant is genuinely useful; if you are doing the same thing but a template would do, adopting anything at all is a cost with no return. The tool surfaces "nothing" wherever the stakes are low enough to deserve it — measured across every combination of answers rather than asserted, because a claim like that is exactly the kind we would otherwise be making about ourselves.

Some of the reasons for ruling something out feel like assumptions rather than facts.

They sometimes are, and where that is so the reason now says which assumption it is making. Telling you that running your own hardware is a large cost assumes you would be building it from scratch — if you already run infrastructure, that cost is paid and the exclusion does not apply. Telling you a specialist service translates better assumes you need publishable translation rather than the gist. Those premises are stated inside the reason so you can reject one without discarding the assessment. An exclusion you cannot contest is not a deduction; it is an assertion in a deduction's clothing, and the difference matters more here than almost anywhere, because the whole tool is an argument that you should check reasoning rather than accept conclusions.

What has actually been built, as against argued?

Built and running: the six situation questions; the elimination rules with the answer behind each shown; the shortlist; the prompt, with its reading guidance travelling inside the copied text; and a self-test that fails the build if the tool stops being able to rule out its author's own product. Argued but not built: the claim that this changes how anyone actually buys. Half-built: the glossary is linked but not yet written to the level a newcomer deserves at every entry. Not built: any way of checking the answers you get back — you do that, and the tool tells you how rather than doing it for you.

What are the limits?

The elimination rules are ours, and they encode judgements you may not share; they are visible so you can disagree with a specific one rather than the whole tool. Nothing here has been tested with people who are not us. The six questions were chosen because they map to consequences we can reason about, not because they are the six that matter most in every setting. And the tool stops at eliciting — it will not tell you whether a given answer to "can it forget?" is a good one, because that judgement depends on obligations only you know you are under.

What would falsify this?

A setting where the eliminations are systematically wrong — where the category the tool rules out turns out to be the one that suited, repeatedly and for reasons the rules could have captured. Or, closer to home, evidence that no real user ever reaches an outcome that excludes what we sell, in which case the self-test is passing on fixtures we wrote while the instrument behaves as a funnel in the world.


← Back to the tool · Understanding AI: a reading guide · The Least Performant AI · The Marks It Leaves · We ran the test on ourselves