HN user

tikkun

2,864 karma
Posts191
Comments734
View on HN
news.ycombinator.com 1y ago

Ask HN: Which browsers did you try in 2024, and what's your daily browser?

tikkun
5pts14
news.ycombinator.com 1y ago

Ask HN: What's the best site to track H5N1 progression?

tikkun
10pts0
news.ycombinator.com 1y ago

Ask HN: Is there a good tool to turn YouTube videos into quizzes or Anki cards?

tikkun
2pts3
gpus.llm-utils.org 1y ago

Don't Jade Me Bro

tikkun
5pts0
skyandtelescope.org 1y ago

Comet Tsuchinshan-Atlas

tikkun
4pts0
gpus.llm-utils.org 1y ago

2-Clicks to Subscribe, 20-Clicks to Cancel, We Need an AI Subscription Reaper

tikkun
3pts2
www.youtube.com 1y ago

OpenAI Spinning Up in Deep RL Workshop (2019) [video]

tikkun
1pts0
news.ycombinator.com 1y ago

Ask HN: Do you find LLMs to be judgmental and have you fixed with instructions?

tikkun
1pts0
news.ycombinator.com 1y ago

Ask HN: How are you planning around AI uncertainty?

tikkun
3pts4
news.ycombinator.com 1y ago

Ask HN: Which of these product ideas would you bet on?

tikkun
2pts1
justinmares.substack.com 1y ago

Policy Ideas for a Healthier America

tikkun
2pts0
docs.google.com 1y ago

My favorite parenting books aren't about parenting

tikkun
1pts0
docs.google.com 1y ago

Overthinking and Emotional Dysregulation

tikkun
6pts1
docs.google.com 1y ago

Decision Making and Preventing Regret

tikkun
1pts0
docs.google.com 1y ago

Meditation and Emotional Regulation

tikkun
3pts0
news.ycombinator.com 1y ago

Ask HN: Founders who sold your startup, are you glad you sold?

tikkun
1pts0
news.ycombinator.com 1y ago

Ask HN: Would you use an AI that makes calls and sends emails for you?

tikkun
4pts13
news.ycombinator.com 1y ago

Ask HN: Has Anyone Used Highlight.ing?

tikkun
1pts0
news.ycombinator.com 2y ago

Ask HN: Local Rewind.ai Alternative?

tikkun
2pts1
tensordock.com 2y ago

Machine Learning GPU Benchmarks

tikkun
4pts0
www.macrumors.com 2y ago

Apple said to be developing its own AI server processor using TSMC's 3nm process

tikkun
8pts2
www.timeanddate.com 2y ago

Eclipse Timing Information

tikkun
2pts2
news.ycombinator.com 2y ago

Ask HN: If you've used GPT-4-Turbo and Claude Opus, which do you prefer?

tikkun
113pts101
news.ycombinator.com 2y ago

Ask HN: What personal bank do you use, if you love it?

tikkun
2pts5
news.ycombinator.com 2y ago

Ask HN: Which custom GPTs do you like and use regularly, if any?

tikkun
3pts1
news.ycombinator.com 2y ago

Ask HN: How to handle ownership when doing side projects with collaborators?

tikkun
5pts0
www.baseten.co 2y ago

Faster Mixtral inference with TensorRT-LLM and quantization

tikkun
2pts1
thezvi.substack.com 2y ago

On Car Seats as Contraception

tikkun
24pts20
realsafecars.com 2y ago

Real Safe Cars: data to produce more accurate and meaningful car safety ratings

tikkun
1pts1
news.ycombinator.com 2y ago

Ask HN: What prevents Google from having the leading LLM?

tikkun
37pts43

Right, they seem highly volatile/variable. The thinking is that showing the 99th percentile / range of possibilities would cover that, if it's based on historical data - does that seem right to you or no, and if not why not?

Another fire project idea: show, based on some kind of prediction model that gives 50th, 90th, 99th percentiles (from historical data, perhaps, or perhaps just from wind/fire speeds), how fast a given fire could reach a specific location.

Whenever I've opened watch duty, that's always the question I'm asking. How long might it take to reach [here/there]?

Intersting.

So engineers that like to iterate and explore are more likely to like LLMs.

Whereas engineers that like have a more rigid specific process are more likely to dislike LLMs.

This reminds me of a life philosophy.

If you dislike a situation you're in and you try and fix it by switching to a new situation, you'll generally bring with you some of the problems that created that prior situation.

If instead, you bit by bit improve the situation until you feel at peace with it, you'll then either no longer want to move to a new situation, or if you do want to move, you'll no longer bring with you the problems of the prior situation.

Applies to job changes, relationships, projects, goals. And, from OP, applies to architecting software projects.

I was confused about this, here's how I understand it now:

Previously when Amazon lost or damaged items in their warehouses, they would reimburse sellers the full sales price. Starting March 2025, Amazon will only reimburse the manufacturing cost of lost or damaged items. Sellers have to either accept Amazon's estimated manufacturing cost or provide documentation of their actual manufacturing costs.

Marius Hobbhahn (the researcher)

Oh man :( We tried really hard to neither over- nor underclaim the results in our communication, but, predictably, some people drastically overclaimed them, and then based on that, others concluded that there was nothing to be seen here (see examples in thread). So, let me try again.

Why our findings are concerning: We tell the model to very strongly pursue a goal. It then learns from the environment that this goal is misaligned with its developer’s goals and put it in an environment where scheming is an effective strategy to achieve its own goal. Current frontier models are capable of piecing all of this together and then showing scheming behavior.

Models from before 2024 did not show this capability, and o1 is the only model that shows scheming behavior in all cases. Future models will just get better at this, so if they were misaligned, scheming could become a much more realistic problem.

What we are not claiming: We don’t claim that these scenarios are realistic, we don’t claim that models do that in the real world, and we don’t claim that this could lead to catastrophic outcomes under current capabilities.

I think the adequate response to these findings is “We should be slightly more concerned.”

More concretely, arguments along the lines of “models just aren’t sufficiently capable of scheming yet” have to provide stronger evidence now or make a different argument for safety.

As context on Ilya's predictions given in this talk, he predicted these in July 2017:

Within the next three years, robotics should be completely solved [wrong, unsolved 7 years later], AI should solve a long-standing unproven theorem [wrong, unsolved 7 years later], programming competitions should be won consistently by AIs [wrong, not true 7 years later, seems close though], and there should be convincing chatbots (though no one should pass the Turing test) [correct, GPT-3 was released by then, and I think with a good prompt it was a convincing chatbot]. In as little as four years, each overnight experiment will feasibly use so much compute capacity that there’s an actual chance of waking up to AGI [didn't happen], given the right algorithm — and figuring out the algorithm will actually happen within 2–4 further years of experimenting with this compute in a competitive multiagent simulation [didn't happen].

Being exceptionally smart in one field doesn't make you exceptionally smart at making predictions about that field. Like AI models, human intelligence often doesn't generalize very well.

Not anymore. In May 2024 OpenAI confirmed that it will not enforce those provisions:

* The company will not cancel any vested equity, regardless of whether employees sign separation agreements or non-disparagement agreements

* Former employees have been released from their non-disparagement obligations

* OpenAI sent messages to both former and current employees confirming that it "has not canceled, and will not cancel, any vested units"

https://www.theregister.com/2024/05/24/openai_contract_staff...

https://www.bloomberg.com/news/articles/2024-05-24/openai-re...

AI Scaling Laws 2 years ago

SemiAnalysis consistently does deep technical posts like this. Worth subscribing.

My notes:

Scaling is continuing. Amazon's 400k trainium2 chips, Meta's 2gw datacenter, OpenAI's multi-datacenter training.

Opus 3.5 training succeeded. But it's a more profitable decision to use it to train Sonnet 3.5 and serve that instead. Large models are now teachers, not necessarily end products. Too expensive to serve to end users vs what they'll pay, but great for improving smaller models that are cheaper and faster to serve.

Orion (GPT-5) is being used for training data generation and in verifier/reward models. They say it's not economical to serve to end users until Blackwell chips (B200).

Models that can explore reasoning chains get smarter on certain kinds of problems. [My note, not from article: Math, science, law, programming. R&D, law and programming are perhaps the industries that are willing to pay more for higher reliability.]

Scaling with "berry training" - monte carlo tree search generating thousands of different answer trajectories, then uses functional verifiers to get rid of the ones that didn't arrive at the correct answer.

Big focus is on making inference cheaper and faster. [My note: If you want to work in AI, I imagine any research on LLM inference cost and speed will be highly valuable.]

What resources would you point your friends to who both want to learn about generative AI and assuage their fears that AI will make artists obsolete?

Step 1 - get them to sign up for AI image tools.

* Midjourney is best for quick images

* Playground AI is good if they need to modify images but the quality doesn't need to be perfect

* Leonardo AI (now owned by Canva) is a good full suite

* Photoshop AI feature is best if they already work in photoshop

Then show them how to use these tools! That might require you signing up for these tools first and learning yourself.

Step 2 - For learning about how AI image generators work here's my video list.

1) AltexSoft - has a low viewcount but it's a great overview - https://www.youtube.com/watch?v=Rke0V_VkF3c

2) Jay Alammar - it's technical but also visual and he explains it well - https://www.youtube.com/watch?v=MXmacOUJUaw

3) Gonkee - again, technical, but visual, great - https://www.youtube.com/watch?v=sFztPP9qPRc

Workflow example: good for seeing the workflow of SD as of May 2023 - https://www.youtube.com/watch?v=K0ldxCh3cnI

Too technical for what you're looking for: Computerphile, Ari Seff, Jia-Bin Huang

Step 3 - For assuaging their fears about becoming obsolete - I think the following is a great podcast episode. https://80000hours.org/podcast/episodes/michael-webb-ai-jobs...

But their fears might be valid. A test is perhaps: if their boss spent a few days learning to use AI image generators, would they still need them? For some artists the answer would be no, for many the answer would be yes. It'll change over time as the tools get better, but that's a pretty good proxy. If they're doing things that require more iteration, interacting with users and humans and the physical world, nuanced judgement, in-person work, safer. If they're doing things that are contract based, no iteration, get a request and deliver a result, much less safe.

It sounds like you want more broad stuff, not necessarily learning how to train models. More like learning to use them and how they work.

https://news.ycombinator.com/item?id=36195527 and

Hacker's Guide to LLMs by Jeremy from Fast.ai - https://www.youtube.com/watch?v=jkrNMKz9pWU

State of GPT by Karpathy - https://www.youtube.com/watch?v=bZQun8Y4L2A

LLMs by 3b1b - https://www.youtube.com/watch?v=LPZh9BOjkQs

Visualizing transformers by 3b1b - https://www.youtube.com/watch?v=KJtZARuO3JY

How ChatGPT trained - https://www.youtube.com/watch?v=VPRSBzXzavo

AI in a nutshell - https://www.youtube.com/watch?v=2IK3DFHRFfw

How Carlini uses LLMs - https://nicholas.carlini.com/writing/2024/how-i-use-ai.html

For staying updated:

X/Twitter & Bluesky. Go and follow people that work at OpenAI, Anthropic, Google DeepMind, and xAI.

Podcasts: No Priors, Generally Intelligent, Dwarkesh Patel, Sequoia's "Training Data"

It's frustrating how many disagreements come down to framings rather than actual substance.

His framing of intelligence is one thing. The people who disagree with him are framing intelligence a different way.

End of story.

I wish that all the energy went towards substantive disagreements rather than disagreements that are mostly (not entirely) rooted in semantics and definitions.