A demo answers one question: can this happen at least once under favorable conditions? A daily tool has to answer a larger set.
Can it recover when the input is messy? Can a person understand what happened? Does it preserve state? Does the cost remain acceptable after hundreds of uses? Can failure be detected before it compounds?
The useful product boundary is not the most dramatic capability. It is the smallest reliable loop that a person can trust enough to repeat.
That is where an experiment starts becoming infrastructure.