HN user

mlmonkey

938 karma

redocpot@yahoo.com

Posts0
Comments257
View on HN
No posts found.

You are evaluated by launches. If you have nothing of substance to launch, then launch a rebrand.

It is a sign of an organization that has nothing of substance to show off. Sadly, that's where Google is now: while its competitors are busy launching new features and improved models, Sir Demis' merry band of pranksters is busy renaming and rebranding things.

We're in crisis. As of 2025, 40% of fourth graders are reading below basic levels

How much of this crisis is due to the social engineering being attempted in school districts across America? Case in point: San Francisco schools decided a couple of years ago that they would no longer teach Algebra in 8th grade. Why? Because too many kids of a certain demographic were failing it. So let's just not teach it! No class, so nobody fails it, right?

It took a proposition on a ballot (i.e., an election) [1] to force the SFUSD to put Algebra back in 8th grade!

I have kids in SFUSD. It often feels like the SFUSD does not care about the average and above average kids; all they focus on is the bottom layer. And even there, they do a terrible job. There was a student who got straight F's in each and every class, and still managed to be a senior in High School! [2]

[1] https://ballotpedia.org/San_Francisco,_California,_Propositi...

[2] https://www.sfchronicle.com/bayarea/article/A-child-left-beh...

GPT-5.6 13 days ago

Where is Gemini in all this? Lately it's not even been in the running. Sir Demis asleep at the wheel? Or Google too scared to release a SOTA model?

Or ... maybe Gemini 4 is too good and the NSA is using it to break into systems worldwide ...?

GPT-5.6 13 days ago

Avoid generic brevity instructions: GPT-5.6 is more sensitive than GPT-5.5 to instructions such as “Be concise,” “Keep it short,” or “Use minimal text.”

What about my favorite, "no yapping"?

GPT‑Live 14 days ago

I speak from experience. The last few weekends I've been taking a friend to see her grandma at a Senior Facility in Sonoma County, and it is really depressing to watch them either lying in bed, listlessly, or in a wheelchair, gazing away in the distance.

They need interaction. And a suitably prompted LLM _can_ provide such interaction. I'm not saying hook them up with ChatGPT and let them loose, but that with the right harnesses and guardrails, they could have a more interactive life.

GPT‑Live 14 days ago

Try being a 90-year old with minimal social contact and nobody to talk to, wasting away in front of a TV blasting NewsMax or Fox News ... does that sound like Heaven or Hell?

GPT‑Live 14 days ago

Is it possible to create a "companion" of sorts with this model, using, say, an RPi and a speaker + microphone? Not for advanced scientific brainstorming, but for seniors who are often alone in their homes.

The problem with such a culture of fear is that the Good Ones(tm) take it as a hint and jump ship; as they're good, they have no trouble finding a job and leaving. But the Mediocre Ones(tm) know that they won't be able to find a comparable job, so they bring the knives out and it becomes a Lord of The Flies situation.

Thus the company gets hurt 2 ways: good ones leave, and bad ones stay, making the lives of everyone else miserable.

Here are the numbers from their bar chart:

    1. SWE-bench Pro
    Model Score (%)
    GLM-5.2 62.1
    GLM-5.1 58.4
    Claude Opus 4.8 69.2
    GPT-5.5 58.6
    Gemini 3.1 Pro 54.2

    2. Terminal-Bench 2.1
    Model Score (%)
    GLM-5.2 81.0
    GLM-5.1 63.5
    Claude Opus 4.8 85.0
    GPT-5.5 84.0
    Gemini 3.1 Pro 74.0
    
    3. NL2Repo
    Model Score (%)
    GLM-5.2 48.9
    GLM-5.1 42.7
    Claude Opus 4.8 69.7
    GPT-5.5 50.7
    Gemini 3.1 Pro 33.4
    
    4. DeepSWE
    Model Score (%)
    GLM-5.2 46.2
    GLM-5.1 18.0
    Claude Opus 4.8 58.0
    GPT-5.5 70.0
    Gemini 3.1 Pro 10.0
    
    5. ProgramBench
    Model Score (%)
    GLM-5.2 63.7
    GLM-5.1 50.9
    Claude Opus 4.8 71.9
    GPT-5.5 70.8
    Gemini 3.1 Pro 39.5
    
    6. MCP-Atlas
    Model Score (%)
    GLM-5.2 77.0
    GLM-5.1 71.8
    Claude Opus 4.8 77.8
    GPT-5.5 75.3
    Gemini 3.1 Pro 69.2
    
    7. Tool-Decathlon
    Model Score (%)
    GLM-5.2 48.2
    GLM-5.1 40.7
    Claude Opus 4.8 59.9
    GPT-5.5 55.6
    Gemini 3.1 Pro 48.8
    
    8. Humanity's Last Exam
    Model Base Score (%) Score w/ Tools (%)
    GLM-5.2 40.5 54.7
    GLM-5.1 31.0 52.3
    Claude Opus 4.8 49.8 57.9
    GPT-5.5 41.4 52.2
    Gemini 3.1 Pro 45.0 51.4
Seems to be handily beating Gemini 3.1 Pro. What _is_ Google DeepMind doing (other than bleeding talent to A\ ) ?

I have seen enough examples of guys prioritizing mental health over money after a certain point.

A friend made mid-8 figures after exit and became a highschool math teacher.

In general: if money was everything, wouldn't the top faculty in every top school be quitting and joining FAANGs ? Who would want a professor's job, making a mid-level SWE amount?