I'm building iPulse AI, an Open Agentic Investment Research Platform, and there's a time-saving claim I'd like to test properly. Getting an answer quickly is useful. But if someone then spends ages checking the numbers and fixing the summary, have we saved time or just moved the work?
I'd measure the whole task, not just the wait for the answer. Start with a research question. Stop when the person has checked the sources and has a note they can actually use. Count corrections too. A fast answer with the wrong numbers doesn't get a medal.
For a small test, I'd use two similar research tasks. Each person does one with AI and one with their usual tools. Swap which task gets AI help across users, so an easier question doesn't make the product look better. Then have someone check the finished notes without knowing which method was used.
I haven't run this test yet. My guess is that making answers easier to check could save more time than making them longer. How are you measuring the checking and fixing that happens after your AI responds?