I ran 1,000 prompts through my little side project and 340 of them broke the exact same way
I built a small tool last month that cleans up messy customer emails, and I thought I had it dialed in around 300 test runs. Then I pushed it to a thousand, and 340 of them failed on the same dumb thing: people writing dates like 3/4 without a year. My model kept guessing the current year, which was wrong about half the time. That number stopped me cold because I had been telling my buddy in Tacoma my setup was basically finished. It was not. What surprised me more was that the fix took twenty minutes once I stopped guessing and just added a regex step before the model ever saw the text. We spend so much time tuning the smart part and forget the boring part in front. So my real question: what percent of your test runs actually fail before you ship, and is it ever the model's fault or always your own sloppy inputs?