HN user

mgrund

126 karma
Posts2
Comments54
View on HN

The big question is if they manage to keep on their enterprise consumers. Vowing changes nothing, they will be looking at the actual numbers.

Dissatisfaction is kind of expected (my cost goes up 2 orders of magnitude and I already cancelled since there are better options at market rate). Complaining without change won’t matter.

swe-bench is a standardized evaluation suite so that's why I'm asking - hopefully there are well-defined criteria on whether this is an open/closed book benchmark.

As I understand it, it is designed to evaluate the LM itself and not agentic systems with online access (very high likelihood of unintentional cheating/solution leaking). The paper and docs are not super clear on the concrete requirements (although reproducibility is emphasized which goes against online access). So I was hoping for someone with more familiarity to chip in.

Obviously not a problem for internal evaluations, but for fair scoreboard submissions it matters. It's not a matter of whether internet searches are useful, but rather what the benchmark is intended to benchmark.

I was under the impression that swe-bench (and I guess most other benchmarks) were supposed to be run offline?

I get that you may accidentally include something in local git history, but it feels off to me to run these kinds of benchmarks online.

I really really want to like local AI, but I highly doubt it will see wide adoption for a long time.

The additional up-front cost for hardware designed to run an LLM in addition to normal workload is unlikely to be accepted by most consumers.

The scale will be very constrained (like Apples on-device models which are small, heavily quantized, and have a small 4K token context window). It’s also terrible for battery life.

AI as it is implemented today is simply just computationally expensive and unless you put in dedicated hardware (like the ANE) for only this purpose - a large cost driver - I don’t really see it getting large scale adoption.

Companies will probably need a server-backed solution as fallback if they want reasonable user experience, so why even invest in diverse hardware support.

There is but I don’t think this is it.

I’ve worked most of my career in US tech satellite offices and I have not experienced EU team members to be less productive than US team members, nor spend less time on work (if anything, more really since they also need to be available for US time zone overlap).

It’s true there are chill jobs here, as there are in the US.

But ambitious people tend to work as much as ambitious US people (and it’s really more like 40 hours work weeks - 39,5 where I live since lunch is not work time). But again, many are not really counting, it’s just a full time job.

Vacations (typically 3 weeks summer holiday and additional weeks to distribute over the year) does create longer time on skeleton crew. Skilled tech labour is also cheaper so you can just hire more to make up for it.

More likely crash looping of so many VMs overloading some system with insufficient back pressure, possibly combined with unfortunate cluster management scheduler behavior at this scale of crash looping (e.g. too eager to retry scheduling instances, maybe even on new hosts which causes more infrastructure load).

I find it very difficult to find meaning in a large portion of the jobs available today. Most workers are just another cog in the wheel. The system is so large that one cannot directly appreciate the effects of their work, or even know whether they are positive at all.

My experience is that the same job also has a large range for meaningfulness depending on how well leadership manages to facilitate it.

Also required in public sector jobs here. When you are on a contract with compensation time off for overtime, and it's sometimes even computed with >1:1 compensation on average hours over some time window, you kind of need to track it.

Personally this is my biggest resistance to in-office work; if the commute was recognized as part of the working hours, I don’t care if I spend some of my workday on the road (although I might feel that time could be better spent, but my employer gets to make that tradeoff).

Great walk-through on the production challenges, but this thread is very precise in its description of the superconducting mechanisms in LK-99. I haven’t followed recent developments but is there any kind of consensus to support that this is the underlying mechanism?

The thread comes off as a pitch for a startup, hence my skepticism.

That’s not what the number means. From the manual:

“Each setting corresponds to a time-based distance that represents how long it takes for Model 3, from its current location, to reach the location of the rear bumper of the vehicle ahead of you”

I think it might have been in the past, at least I was told the same thing years ago.

That said, it is actually an issue that it now keeps too much distance under some conditions on the lowest setting.

Update on Sharing 3 years ago

Would feel like less of a money grab if they rolled this out with price reductions, given the increase in subscriptions I assume they expect as a consequence. They might even be able to sell it as a way to bring justice to those that do not share accounts and who are currently covering the cost of other people doing it.

Absolutely, this is a central subject of study in Tactile Internet from what I’ve heard; next-gen backhaul/WAN technology. It starts with first hop to support local Tactile Internet networks. Private, ultra low latency 5G LANs are already being used in e.g. some production facilities.