HN user

dongobread

670 karma
Posts0
Comments42
View on HN
No posts found.

What a strangely hostile statement on an open weight model. Running like 20 benchmark evaluations isn't trivial by itself, and even updating visuals and press statements can take a few days at a tech company. It's literally been 5 days since this "new generation" of models released. GPT-5.3(-codex) can't even be called via API, so it's impossible to test for some benchmarks.

I notice the people who endlessly praise closed-source models never actually USE open weight models, or assume their drop-in prompting methods and workflow will just work for other model families. Especially true for SWEs who used Claude Code first and now think every other model is horrible because they're ONLY used to prompting Claude. It's quite scary to see how people develop this level of worship for a proprietary product that is openly distrusting of users. I am not saying this is true or not of the parent poster, but something I notice in general.

As someone who uses GLM-4.7 a good bit, it's easily at Sonnet 4.5 tier - have not tried GLM-5 but it would be surprising if it wasn't at Opus 4.5 level given the massive parameter increase.

Open models by OpenAI 12 months ago

It is absolutely awful at writing and general knowledge. IMO coding is its greatest strength by far.

Open models by OpenAI 12 months ago

How up to date are you on current open weights models? After playing around with it for a few hours I find it to be nowhere near as good as Qwen3-30B-A3B. The world knowledge is severely lacking in particular.

This is a little misleading. The data they quote is based on their previous article[1], which just uses this analysis[2] provided by a VC company. Funnily enough the same VC company put a seperate clickbaitish article just a year before that one, claiming the exact opposite findings (about startups ditching SV).

I would guess a lot of these annual trends are just random fluctuations in their dataset, though to be honest I wonder how they're even trying to estimate this kind of information.

[1] https://www.wsj.com/articles/austins-reign-as-a-tech-hub-mig...

[2] https://www.signalfire.com/blog/signalfire-state-of-talent-r...

[3] https://www.signalfire.com/blog/state-of-talent-tech-trends

The corporate politics at Meta is the result of Zuck's own decisions. Even in big tech, Meta is (along with Amazon) rather famous for its highly political and backstabby culture.

This is because these two companies have extremely performance-review oriented cultures where results need to be proven every quarter or you're grounds for laying off.

Labs known for being innovative all share the same trait of allowing researchers to go YEARS without high impact results. But both Meta and Scale are known for being grind shops.

I'm very skeptical on this, the paper they linked is not convincing. It says that GPT-4 is correct at predicting the experiment outcome direction 69% of the time versus 66% of the time for human forecasters. But this is a silly benchmark because people are not trusting human forecasters in the first place, that's the whole purpose for why the experiment is run. Knowing that GPT-4 is slightly better at predicting experiments than some human guessing doesn't make it a useful substitute for the actual experiment.

I think what you say is true when comparing transformers to CNNs/RNNs, but not to MLPs.

Transformers, RNNs, and CNNs are all techniques to reduce parameter count compared to a pure-MLP model. If you took a transformer model and replaced each self-attention layer with a linear layer+activation function, you'd have a pure MLP model that can model every relationship the transformer does, but can model more possible relationships as well (but at the cost of tons more parameters). MLPs are more powerful/scalable but transformers are more efficient.

Compared to MLPs, transformers save on parameter count by skimping on the number of parameters devoted to modeling the relationship between tokens. This works in language modeling, where relationships between tokens isn't that important - you can jumble up the words in this sentence and it still mostly makes sense. This doesn't work in time series, where relationships between tokens (timesteps) is the most important thing of all. The LTSF paper linked in the OP paper also mentions this same problem: https://arxiv.org/pdf/2205.13504 (see section 1)

From experience in payments/spending forecasting, I've found that deep learning generally underperform gradient-boosted tree models. Deep learning models tend to be good at learning seasonality but do not handle complex trends or shocks very well. Economic/financial data tends to have straightforward seasonality with complex trends, so deep learning tends to do quite poorly.

I do agree with this paper - all of the good deep learning time series architectures I've tried are simple extensions of MLPs or RNNs (e.g. DeepAR or N-BEATS). The transformer-based architectures I've used have been absolutely awful, especially the endless stream of transformer-based "foundational models" that are coming out these days.

I get what this piece is trying to say, but it's ignoring the fact that schools are trying to maximize learning with pupils who often don't want or care about learning (unlike with athletes or musicians who are generally learning their craft by choice).

A significant part of teaching disinterested students (not just in a grade school but in general) is about making the subject interesting enough that students will want to spend time on learning and continue to delve further in their free time.

If you're trying to teach someone web development, would you have them churn through a stack of predetermined bootcamp-style projects, or would let them try to build something they have personal interest in? I bet the latter method would turn out much better for the student in the long run.

I'm not sure what would lead to you believe this. I've worked in the data science/ML space for over a decade now and I see the majority of pure analytics projects started in R, including at big tech companies I've worked at recently.

Of course, ML projects and other things that need to result in production-grade models are almost always done in Python. This is currently the most visible form of "data project" due to all the ML/AI hype, but it is far from the only data work going on.

Assuming you already know some basic linear algebra and calculus, know Python (or R), and have a decent-but-not-advanced grasp of statistics, I'd recommend working through these books. They are very readable and focus on intuitive understanding/practical applications, but give enough technical foundation for you to jump into more specific subfields if needed.

Stats & ML - https://www.statlearning.com/

Deep Learning - https://udlbook.github.io/udlbook/

Reinforcement Learning - https://web.stanford.edu/class/psych209/Readings/SuttonBarto...

As with anything else, people usually fail to learn ML not because of content quality but because of lack of effort/time/consistency. Take handwritten notes, solve exercises, etc., and expect to spend at least a hundred hours on each book.

We tried using a multi-agent system for a complex NLP-type task and we found:

- Too many errors that just propogate on top of each other, if a single agent in the chain generates something even a little bit off then the whole system goes off the rails.

- You often end up having to pass a massive amount of shared context to every agent which just increases the cost dramatically.

Curiously enough we had an architect from OpenAI tell us the same thing about agent systems a few days ago (our company is a big spender so they serve a consulting function), so I don't think anybody is really finding success with multi-agent systems currently. IMO the core tech is nowhere near good enough yet.

Legality aside, I think the "payment" people get from posting free knowledge on the Internet is the human connection, and the satisfaction of knowing that other people are reading and appreciating it directly.

Injecting an LLM middleman between your post and the end user changes this dynamic quite a bit - without the human component, the feeling is that you're just doing unpaid labor for a profit-oriented company (OpenAI).

I don't think either of those theories is right. (1) doesn't explain the rise in corporate profits, and (2) is of course silly.

Here is my theory:

Consumers generally have an "acceptable" price range in their head for each product. When most retailers have prices within that range, it is really hard for a single company to raise prices above that range, as they'll lose a lot of business. But, COVID forced input prices to rise and fluctuate quite a bit. Once companies raised prices to account for input costs, consumers lost their sense of "normal". Then, companies were able to get away with charging prices even more, and could get away with raising prices above their competitors without losing any business.

Gemini AI 3 years ago

This isn't apples to apples - they're taking the optimal prompting technique for their own model, then using that technique for both models. They should be comparing it against the optimal prompting technique for GPT-4.

I don't think the target market for this is people looking for extremely knowledgeable LLMs that can handle deep technical tasks, given that you can't even finetune these models.

I'd guess this is more of an attempt to poach the market of companies like character.ai. The market for models with a distinct character/personality is absolutely massive right now (see: app store rankings) and users are willing to spend insane amounts of money on it (in part because of the "digital girlfriend" appeal).

Clickbait unfortunately, this is median not mean. If you check the source data[1], 1m net worth is at roughly 90th percentile of US households.

Also interesting in the original survey - for the median household, over >80% of their wealth is in their house (and 2/3s of households have houses). Whereas stocks are mostly the domain of the wealthy.

[1] https://www.federalreserve.gov/publications/files/scf23.pdf

The idea of splitting off payment plans into its own abstraction is super interesting. Personally I would really like if the abstraction layer also was independent of the payments processor, currently it seems heavily dependent on Stripe. Even though I currently use Stripe I find them to be untrustworthy and don't want anything that ties me to them even further. And more generally, using a platform built on top of a SaaS is a recipe for future headaches.

TimeGPT-1 3 years ago

As someone who's worked in time series forecasting for a while, I haven't yet found a use case for these "time series" focused deep learning models.

On extremely high dimensional data (I worked at a credit card processor company doing fraud modeling), deep learning dominates, but there's simply no advantage in using a designated "time series" model that treats time differently than any other feature. We've tried most time series deep learning models that claim to be SoTA - N-BEATS, N-HiTS, every RNN variant that was popular pre-transformers, and they don't beat an MLP that just uses lagged values as features. I've talked to several others in the forecasting space and they've found the same result.

On mid-dimensional data, LightGBM/Xgboost is by far the best and generally performs at or better than any deep learning model, while requiring much less finetuning and a tiny fraction of the computation time.

And on low-dimensional data, (V)ARIMA/ETS/Factor models are still king, since without adequate data, the model needs to be structured with human intuition.

As a result I'm extremely skeptical of any of these claims about a generally high performing "time series" model. Training on time series gives a model very limited understanding of the fundamental structure of how the world works, unlike a language model, so the amount of generalization ability a model will gain is very limited.