GLM 5.2 beats Claude in our benchmarks 23 days agoWe ran a set of popular open-source models against our IDOR benchmark."our IDOR benchmark", there you go. 0ThreadHN
Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable 1 month agoFable has been pretty disappointing for security research. It downgrades itself to Opus 4.8 even when you ask it questions about basic things like port scanning. 0ThreadHN
Exposing Critical Vulnerabilities in CBSE's On-Screen Marking Portal 2 months agoAuthor here, thanks for posting about it :) 0ThreadHN