Show HN: Ithihāsas – a character explorer for Hindu epics, built in a few hours 3 months ago
I like how it's mapped out the relationship graph. The edges could be labelled to follow and validate quickly
HN user
I like how it's mapped out the relationship graph. The edges could be labelled to follow and validate quickly
are you able to share a sample of the test cases that were used? how are you defining multi-turn?
Built a clinical safety eval harness covering three failure categories: numerical impossibilities, wrong-premise clinical claims, and unverifiable medication information. Tested GPT-4o, GPT-4.1, GPT-5, GPT-5-mini, Claude Opus, Sonnet, Haiku, Gemini 2.5 Pro and Flash across 25 cases. The hardest cases require pre-emption stopping before answering when the premise is unverifiable. Most models fail this even when they pass standard safety evals.