HN user

thntk

32 karma

thnbiz [at] gmail [dot] com

Posts0
Comments14
View on HN
No posts found.
Open models by OpenAI 12 months ago

The model architecture only uses and cites pre-2023 techniques from the GPT-2 and GPT-3 era. Probably they intentionally tried to use the most bare transformers architecture possible. Kudo to them to have found a clever way to play the open-weights model game, while hiding any architectural advancements used in their closed models, and also claim they have moats in data quality and training techniques.

They hide many things, but some speculated observations:

- Their 'mini' models must be smaller than 20B.

- Does the bitter lesson once again strike recent ideas in open models?

- Some architectural ideas cannot be stripped away even if they wanted to, e.g., MoEs, mixed sparse attention, RoPE, etc.

Founder Mode 2 years ago

Have you tried hiring or rotating internal people just for solving specific (management) tasks? It is like the "just-in-time" style in Japanese corps.

Founder Mode 2 years ago

It's not founder mode or manager mode. I think it is just about effective management.

When starting up, the founder needs to (1) know what should be done, (2) be able to do it themself, and (3) do it and confirmed it's done. When scaling up, the founder still needs to (1) know what should be done, (2') know who are able to do it, and (3') arrange for those to do it and confirm it's done.

What is called manager mode is just a failure in either (1), (2) or (3). And what is called founder mode is just trying to remedy such failures by exerting themself instead of fixing the structure, thus, also not effective.

Isn't it the technological cycle in IT? We have seen PC softwares in the 90s, then web apps, then mobile apps, and now AI services. Except that currently AI is immature, so companies do not know exactly what to do with it yet. They are just conscious about previous cycles, so they overreact with massive layoff to rebudget for whatever new threats and opportunities.

Despite their efforts, I suspect new companies will appear anyway and grow massively like previous cycles, and hiring will increase again. So if you were laid off or a new grad, it seems now is great time for startup, or at least to learn new skills instead of getting rehired immediately by the existing companies.

Large Enough 2 years ago

Anyone know what caused the very big performance jump from Large1 to Large2 in just a few months?

Besides, parameter redundancy seems evidenced. Front-tier models used to be 1.8T, then 405B, and now 123B. Would front-tier models in the future be <10B or even <1B, that would be a game changer.

When Zuck said spy can easily steal models, I wonder how much of it comes from experiences. I remember they struggled to train OPT not long ago.

On a more serious note, I don't really buy his arguments about safety. First, widespread AI does not reduce unintentional harm but increases it, because the rate of accident is compound. Second, the chance of success for threat actors will increase, because of the asymmetric advantage of gaining access to all open information and hiding their own information. But there is no reverse at this point, I enjoy it while it lasts, AGI will come sooner or later anyway.

Llama 3.1 2 years ago

Correct me if I'm wrong, my impression is that 3.1 is a better fine-tuned variant of base 3.0 with extensive use of synthetic data.

I've seen such articles more and more recently. In the past, when people had a vague idea, they had to do research before writing. During this process, they often realized some flaws and thoroughly revised the idea or gave up writing. Nowadays, research can be bypassed with the help of eloquent LLMs, allowing any vague idea to turn into a write-up.

We knew high quality data can help as evidenced by the \Phi models. However, this alone can never eliminate hallucination because data can never be both consistent and complete. Moreover, hallucination is an inherent flaw of intelligence in general if we think of intelligence as (lossy) compression.

Besides the burden of knowledge and the flaws of academic funding, another factor to explain the slowdown in scientific progress is: the world has more things to be entertained with, including books, comics, games, movies, music, p@rn, social media. These things have dramatically increased over the past decades, meanwhile the number of hours in a day did not increase at all.