Kept hand-tuning my prompts for weeks until a coworker asked why I wasn't just testing ten at once
I was so sure that sitting there tweaking one prompt word by word was the only way to get good output from the model we use for customer tickets. My coworker watched me do it for an hour at our Austin office and pointed out I could run batches of variations overnight and compare them, which cut my work from 15 hours a week to about 4. Has anyone else fallen into the trap of treating prompt work like a craft when it's really just trial and error?
The batch thing works but it's not exactly just trial and error. You still need a way to judge which output is better, and most people skip that part. If you run ten prompts overnight and pick the one that "looks good," you're kind of back where you started, just faster. The real fix was deciding upfront what a good ticket reply even means (tone, length, does it actually solve the problem) so you can score them. Otherwise you're just guessing in bulk.