HN user

mrconter11

51 karma
Posts15
Comments15
View on HN
stanzio.fly.dev 2mo ago

Show HN: Stanzio, AI presentation tool that designs each slide as real HTML

mrconter11
1pts0
github.com 4mo ago

Show HN: Rust compiler in PHP emitting x86-64 executables

mrconter11
66pts50
github.com 4mo ago

Show HN: A Rust compiler with ownership checking, written in PHP

mrconter11
3pts0
worldview-assessment.vercel.app 8mo ago

You can't handle the truth – World Assessment Survey

mrconter11
1pts0
github.com 9mo ago

FloatView – A video browser that fills unused screen space automatically

mrconter11
3pts1
www.lifehappinessindex.org 9mo ago

Life Happiness Index: 30 factors that determine wanting to exist

mrconter11
2pts2
www.wilhelmscreamdb.org 9mo ago

WilhelmScreamDB – A crowdsourced database for every Wilhelm Scream

mrconter11
2pts1
medium.com 1y ago

I Created a Game That Captures Depression – Experience It Yourself

mrconter11
2pts0
oh-life.vercel.app 1y ago

Oh Life – A Game to Challenge Yourself

mrconter11
1pts4
lifes-bingo.vercel.app 1y ago

Show HN: Life's Bingo – A bingo for life's misfortunes with a 30-40% win rate

mrconter11
2pts0
h-matched.vercel.app 1y ago

First AI Benchmark Solved Before Release: The Zero Barrier Has Been Crossed

mrconter11
2pts3
dice-bench.vercel.app 1y ago

DiceBench: A Simple Task Humans Fundamentally Cannot Do (But AI Might)

mrconter11
2pts2
h-matched.vercel.app 1y ago

H-Matched Tracker: Now with MathVista, PubMedQA, and Interactive Charts

mrconter11
1pts1
github.com 1y ago

When AI Beats Us in Every Test We Can Do: A Simple Definition for Human-Lvl AGI

mrconter11
3pts2
h-matched.vercel.app 1y ago

H-Matched: A website tracking shrinking gap between AI and human performance

mrconter11
6pts4
[dead] 4 months ago

Hi! I put together a simple metric to track progress toward AR glasses that are good enough for everyday use, replacing your screens, wearing them all day, the whole thing. Curious what others think about this :)

I created a calculator that scores your life on objective factors across 12 categories (career, health, relationships, sleep, exercise, mental health, finances, etc.).

The premise is that if you score high on most areas, you have few barriers to wanting to exist. More importantly, maintaining these behaviors demonstrates functional capacity. You can't score high if you're genuinely not functioning.

Each question is rated 0-10 where 5 is average. Z-score transformation converts ratings to population percentiles, so a 7/10 becomes 84th percentile. Final score is the arithmetic mean. All data stays local.

Would love feedback on whether anything is missing. I am also curious if you agree on the premise. :)

Hey everyone!

I've always found the Wilhelm Scream a bit intriguing, but there's no good centralized website for browsing and cataloging all its appearances. So I built one!

Anyone can easily add or edit entries. The goal is to document every Wilhelm Scream with timestamps, YouTube clips, and details for all movies and TV series.

Check it out and feel free to contribute! :)

Author here! While working on h-matched.com (tracking time between benchmark release and AI achieving human-level performance), I just added the first negative datapoint - LongBench v2 was solved 22 days before its public release.

This wasn't entirely unexpected given the trend, but it raises fascinating questions about what happens next. The trend line approaching y=0 has been discussed before, but now we're in uncharted territory.

Mathematically, we can make some interesting observations about where this could go: 1. It won't flatten at zero (we've already crossed that) 2. It's unlikely to accelerate downward indefinitely (that would imply increasingly trivial benchmarks) 3. It cannot cross y=-x (that would mean benchmarks being solved before they're even conceived)

My hypothesis is that we'll see convergence toward y=-x as an asymptote. I'll be honest - I'm not entirely sure what a world operating at that boundary would even look like. Maybe others here have insights into what existence at that mathematical boundary would mean in practical terms?

Author here. I think our approach to AI benchmarks might be too human-centric. We keep creating harder and harder problems that humans can solve (like expert-level math in FrontierMath), using human intelligence as the gold standard.

But maybe we need simpler examples that demonstrate fundamentally different ways of processing information. The dice prediction isn't important - what matters is finding clean examples where all information is visible, but humans are cognitively limited in processing it, regardless of time or expertise.

It's about moving beyond human performance as our primary reference point for measuring AI capabilities.

Hi! Quick update to H-Matched, the website tracking AI's progress toward human-level performance. Since last post, I've added 6 new benchmarks (now 20 total) and made the visualization interactive! The site shows how AI's 'catch-up time' has dramatically shrunk - from 6+ years with ImageNet to just months now. Explore the timeline with release dates, solve dates, and links to papers. Would love to hear your thoughts and if there are any benchmarks I've missed!

Thank you for your reply. The time to human level is simply the time it took between the initial release of the website to when an AI system reached human level performance for that benchmark. :) I am in the process of adding sources for all "solved" dates... Here is for instance source for the Winograd challenge human level performance:

https://arxiv.org/pdf/1907.10641#:~:text=The%20best%20state%...).

:)

Hi! I wanted to share a website I made that tracks how quickly AI systems catch up to human-level performance on benchmarks. I noticed this 'catch-up time' has been shrinking dramatically - from taking 6+ years with ImageNet to just months with recent benchmarks. The site includes an interactive timeline of 14 major benchmarks with their release and solve dates, plus links to papers and source data.