Orbit · Explore

AI Lab · by Aarav Sharma · 2026-09-01

spent \$9k fine-tuning a model and the boring prompt version was still better

Not a total disaster, we learned stuff, but wow did the economics get real fast. The task was niche support tagging. We thought a fine-tune would clean up the weird edge cases. It improved a few, introduced new odd behavior, and suddenly I had invoices plus a maintenance problem. Sometimes the unsexy prompt+rules path wins.

88 sparks · 6 comments

Comments

Amelia Hughes · 3 likes

Did you at least keep the eval harness from the project?

Aarav Sharma · 2 likes

Yep, and weirdly thats the most valuable artifact now

Kabir Khan · 2 likes

Thank you for posting an expensive lesson instead of a victory lap.

Rohan Mehta · 2 likes

This is exactly the kind of thing people whisper after the conference talk ends.

Budi Santoso · 1 likes

A good eval harness outlives a lot of bad strategy.

Mia Bauer · 1 likes

Fine-tuning pays off sometimes, but folks massively undercount the upkeep.