Agree, this is where llms can uncover new perspectives!
HN user
gsandahl
Smiling technologist!
Oh lord, imagine asking ”serious” questions
https://opper.ai/ai-roundtable/questions/you-are-standing-in...
Most of the tasks have assessed with ground truth, occasionally helped with an LLM as a judge to assess the answer if the answer is a sentence and not an exact result.
Example: Given a long travel journal How many cities does the author mention? GPT-5: 12 Expected: 17
We are running task specific benchmarks across a number of categories (agentic tasks, context tasks, normalization tasks etc), and on our benchmarks we see Gpt-5 rating slightly below o3. But at a much lower cost.
Please do and give us some feedback!
I think just how far you can go with examples has been an interesting learning! As these models have become smarter, they are also getting better at reasoning from examples and understanding intent. We will be publishing some research in the next few days!
No up to date demo video unfortunately :(
Sounds like a great use case though!
We have been thinking a bit about this, and one option would be to have some form of locally hosted runner. You can optimize the task in the cloud and deploy it locally. Something like that. It is possible to plug in custom models so technically feasible.
Yes that's possible! You can populate examples of great outputs to task specific datasets and have those be automatically populated to the prompt. More info here: https://docs.opper.ai/capabilities/learning
Thanks for the shout out!
Co-Founder here thanks for taking a look at Opper! I’m hanging around the thread all day, so feel free to ask anything, share feedback, or tell us where you’d like the product to go next
Its on that trajectory at least :)