If it didn’t close, it didn’t ship.
A look at exactly how every prompt in Promptifi gets graded, run, and re-run before it lands in your library — and what happens when it falls below the line.
Most prompt libraries are generated. Ours is tested.
The internet is drowning in prompts no one ever ran in a real deal. We built the testing protocol because we hated trying them ourselves and finding out, mid-call, that nothing came out the other end.
Three things we don’t do: ship from theory, score with an AI judge, or publish anything we wouldn’t run on our own pipeline.
Tested in real deals, not test data
Every prompt is run against an actual prospect, an actual call transcript, an actual CRM record. Synthetic test cases are too forgiving.
Graded by the rep, not the model
Output gets a five-point score from the rep who ran it. AI-on-AI evaluation is where prompt quality goes to die.
Three runs minimum, five preferred
One opinion isn’t a signal. Three across different verticals starts to be. We average and weight by deal-stage match.
Re-tested every quarter
Models drift, sales tactics evolve, what worked in Q1 may flop in Q3. Every published prompt cycles back through the protocol.
From draft to library.
Same five gates for every prompt — discovery question, cold email, MEDDPICC builder, renewal play. No shortcuts, no special cases.
Draft
An in-house operator or beta tester writes V1 against a specific deal scenario. Stage-tagged from day one. Time-boxed to 20 minutes.
∼ 20 minDry run
Output is reviewed against three reference deals. Obvious failures get killed here — hallucinated company facts, generic copy, wrong stage tone.
internalField test
Three to five testers run the prompt in live pipeline. Output goes into actual emails, calls, CRM notes. They log time saved and outcome.
5 to 14 daysScore & revise
Five-criterion rubric averaged across testers. Below 4.0 goes back to revision. Above and it earns the “Tested” flag.
≥ 4.0 to shipPublish & re-run
Goes live in the library with tester names and scores. Auto-flagged for re-test 90 days later. Score below 3.5 on retest — pulled.
quarterlyThe protocol, by the numbers.
Cumulative numbers since we started recording the protocol.
The prompts we didn’t ship.
“Aggressive renewal nudge”
Output sounded like a debt collector. Two of three CSMs said they’d never send it. We tried three rewrites — none cleared 3.0 on tone fit.
“LinkedIn post generator”
Hallucinated stats, fabricated quotes, sounded like every other AI thread. Wasn’t a sales workflow — bad fit for the library.
“Cold email opener — generic”
Shipped at 4.1 last quarter. Re-test scored 3.2 — reply rates collapsed once the pattern got common across the wider market.
Want to see the tested prompts?
Or apply to be one of the reps who tests them. Both take less than three minutes.