HN user

tartakovsky

112 karma
Posts14
Comments76
View on HN

Well, task == Resolving real GitHub Issues

Languages == Python only

Libraries (um looks like other LLM generated libraries -- I mean definitely not pure human: like Ragas, FastMCP, etc)

So seems like a highly skewed sample and who knows what can / can't be generalized. Does make for a compelling research paper though!

2-ish questions:

Is this level of fear typical or reasonable? If so, why doesn’t Anthropic / AI code gen providers offer this type of service? Hard to believe Anthropic is not secure in some sense — like what if Claude Code is already inside some container-like thing?

Is it actually true that Claude cannot bust out of the container?

Hyprland Premium 1 year ago

I clicked 3 different links failing at answering that question for myself and then I stopped caring, though not enough to pass up one-upping your comment.

Bop Spotter 2 years ago

Huh? “Total Shazams ever detected: 240. That's an average of 240 songs per day.”

What is your goal? if d1, d2, d3, etc is the dataset over which you're trying to optimize, then the goal is to find some best performing d_i. In this case, you're not evaluating. You're optimizing. Your acquisition function even says so: https://rentruewang.github.io/bocoel/research/

And in general if you have an LLM that performs really well on one d_i then who cares. The goal in LLM evaluation is to find a good performing LLM overall.

Finally, it feels that your Abstract and other snippets sound like an LLM wrote them.

Good luck.

The webpage discusses a Bayesian approach to experimentation, focusing on interpreting and extrapolating experimental results, mainly in a tech environment aiming to maximize user retention. It addresses challenges like the inference problem, the extrapolation problem, the explore-exploit problem, and a culture problem within tech companies around misuse of experiments. The author suggests providing decision-makers with benchmark statistics to help them estimate true effects of different policies, and discusses a model of experimentation dealing with observed and true effects along with the noise in experiments 1 .

What’s better, train a model with 10X parameters once on some default hyperparameter setting or to search for a good hyperparameter configuration by training on X parameters 10 times? While I’m at it, how many LLMs of the size of GPT3 were trained until they landed on the capability of GPT3? How much of this is dependent on the data, or do good settings transcend the type of text that a model is trying to train on?

Just having a simple conversation with ChatGPT (the upgraded GPT-4 version) this morning. In the end searching GitHub myself was faster. Any tips on this or research on this topic? Is this what is referred to as In Context learning? Or In Context reminders?

Wow! I don't know how accuracy translates, I do see charts that look strong but unless I'm missing something, this is incredible. Would be curious and also terrified to see an endpoint so I can play around with it. I thought we were stopping this kind of research?