I paid 52 strangers to poke holes in Stet
Stet is a small, early product. Before I talked myself all the way into my own pitch, I wanted to hear from people who had never met me and had no reason to be nice.
So I ran a study. I recruited people who write with AI every week, asked them about a real problem they had hit, then showed them the Stet landing page and asked the blunt questions. Does this fix your problem? What would stop you from using it? 351 people started, and 52 made it to the end. That 52 is the denominator for every number below.
Small sample, paid panel, all of it self-reported. Read what follows as a set of signals worth chasing, nothing sturdier than that.
The problem is real, for about half of them
26 of the 52 rated the fit "Very" or "Perfectly."
Four problem statements were on offer and people picked whichever was theirs. Two carried almost the whole thing: the AI that quietly drops a constraint while rewriting, at 15 picks, and the document that loses the trail of who changed what after passing through several hands and tools, at 13. People did not just rate those highly. They told their own versions, unprompted and in detail. Claude pulling a deadline out of a literary pitch. A production memo that lost a delivery date and cost someone a material delivery. A PRD where "under 2 seconds" became "fast" with no record of who decided that.
The framings I had more invested in did worse. Spec attribution got 9 picks. The one about markdown losing its review layer got 4, dead last, which is uncomfortable given how much of the product rests on exactly that idea.
And 11 people, better than a fifth of the sample, chose "none of these describe my work." Their fit ratings were correspondingly grim. The screener should have caught them and did not.
The wall was somewhere else
I expected the objection to be "meh." It was not.
Price was the most common thing people said would stop them, at 19 mentions. Integration came second at 10, mostly Google Docs and a general weariness about installing another app. Platform was 7, every one of them some version of Mac-only.
I am discounting the price number, and I want to be honest that this happens to be the conclusion most convenient to me. On a paid panel, "it might be expensive" is close to a reflex, and it came back as the top blocker in every single role, which is what a reflex looks like. Am I right to throw out my largest number because it is inconvenient? I genuinely do not know. A pricing-sensitivity probe is on the follow-up list to separate real resistance from panel noise, and until that runs, this paragraph is a hypothesis I have a stake in.
What I trust more is where the credible blockers sit. The people who picked the two winning problems and rated the fit Very or Perfectly were stopped, over and over, by reach rather than by doubt. "It's only on Mac; I use Windows." "If it does not work with Google Docs, that is where 80% of my collaboration happens." The best-fit respondent in the entire set, a security engineer who lost about a fifth of a 30 to 40 hour threat assessment to an AI rewrite, is blocked by precisely one thing, and that thing is Windows.
So I built a review tool for people who write with AI, and the loudest credible complaint is that they cannot open it. Hard to argue with that one.
What I am changing
Four things.
Stop bouncing Windows visitors into a dead end. If you cannot download it yet, I still want to know you exist, so I can tell you when you can.
Lead with the two problems that actually landed, instead of the one I happened to write the headline around. Both have had most of my build time since. The version history, so you can see and recover what an edit took out. Stet sync, so the trail survives when a document passes between people and machines.
Move the markdown framing out of the headline. This is the one I had to think hardest about, because it came last and I am the most attached to it. The four people who picked it also rated the fit higher than any other group did. Four is not a base you can build on. It is not nothing, either. My read is that I had the right idea in the wrong slot. "Markdown lost its review layer" is a bad wound to open with, because almost nobody wakes up feeling it. It is a good explanation of why the wound exists, for someone who already feels the AI dropping things on them. So it moves from the first thing I say to the answer I give second. I have written that argument out on its own, and it will go up separately.
And take the friction worry seriously even though it arrived from fans. Two of the highest-fit people in the study volunteered it without being asked, one of them putting it at "90% of permission requests will be busy work." They are right to ask. Stet now grades every change with a local classifier and holds only the ones that look significant, so routine edits land marked instead of waiting in a queue. Clearing a batch shows you the riskiest thing in it before you clear it. That is a first answer, not a finished one. Review only earns its place if it stays cheap, and that is a design problem, not an objection to argue away.
This is the part of building in public I actually like. I did not have to defend anything. I just had to shut up and read.
Download for Mac and tell me where I am wrong.