If it was the easy part, then why did they pay us hundreds of thousands, sometimes millions, sometimes more - to do it? The fact of the matter is that it wasn't easy, not for a brain that's architected the way a human's is. The fact that computers can now do it much more quickly and arguably - in many cases - better doesn't diminish the act itself - it just shows how far AI has come, and how easily human intelligence will be dwarfed as it continues to make progress.
HN user
Aqueous
sorry you can’t keep up
Anyone who thinks LLMs have not come a long way in approximating human linguistic capabilities (and associated thinking) are in fact, engaging in (delusional) wishful thinking regarding human exceptionalism.
With respect to consciousness, you are doing nothing more than asserting a special domain inside the brain that, unlike the rest of the mechanisms of the brain, has special "magic" that creates qualia where classical mechanisms cannot. You are saying that there is possibly a different explanation for intelligence as consciousness, when it would be much simpler to say the same mechanisms explain both. Furthermore, you have no explanation for why this quantum "magic", even if it was there, would solve the hard problem of consciousness - you are just saying that it does. Why should quanta lend themselves anymore to the possibility of subjective experience/qualia than classical systems? Finally, a brain operates at 98.6° and we can't even create verifiable quantum computing effects at near absolute zero, the only place where theory and experiment both agree is the place quantum effects start to dominate. The burden of proof is on you and Penrose as what you are both saying is wildly at odds with both physics, experimental and theoretical, and recent advancements in computing. Penrose is a very smart guy but I fear on these questions he's gone pretty rogue scientifically.
In just 2-3 years we've gone from primitive LLMs to LLMs reaching Graduate PhD-level knowledge and intelligence in multiple domains. LLMs can complete almost any code I write with high accuracy given sufficient context. I can have a naturalistic dialog with an LLM that goes on for hours in multiple languages. Frankly (and humblingly, and frighteningly) they have already surpassed my own knowledge and intelligence in many, probably most, domains. Obviously they aren't perfect and make a lot of errors - but so do most humans.
What's odd about the current moment is that in the very same era in which it seems there is conclusive evidence (LLMs) that quantum explanations are not necessary to explain at the very least linguistic intelligence as advanced linguistic intelligence is possible in a purely classical computing domain, there is at the same time an insistence elsewhere that consciousness must be a quantum phenemonon. Frankly I am increasingly skeptical that this is the case. LLMs show that intelligence is at least mostly algorithmic, and the brain is far too warm and wet for quantum effects to dominate. Why should intelligence be purely classical but consciousness (another brain phenemenon) be quantum? It lacks parsimony.
The fact that it takes decades to master such a mundane task may mean the entire approach is wrong. The article hand-waves a lot of the complexity of "automating as much as possible."
In my opinion, the solution lies in append-only software as dependencies. Append-only means you never break an existing contract in a new version. If you need to do a traditional "breaking change" you instead add a new API, but ship all old APIs with the software. In other words - enable teams to upgrade to the latest of anything without risking breaking anything and then updating their API contracts as necessary. This creates the least friction. Of course, it's a long way for every dependency and every transitive dependency to adopt such a model.
"Doing them swiftly, efficiently, and -- most of all -- completely is one of the most critical skills you can develop as a team."
That all sounds great. However, I'd like to understand what teams are actually able to do this, because it seems like a complete fantasy. Nobody I've seen is doing migrations swiftly and efficiently. They are giant time-sucks for every company I've ever worked for and any company anyone I know has ever worked for.
Not talking about just monitoring outputs though. I'm talking about monitoring the internals of the model as it reaches its output. The entire issue around interpretability / observability inside the LLM's model is the hard problem, one for which considerable resources are being dedicated to solve - not simply hooking the public-facing APIs up to observability tools like any other service API. This is just conventional telemetry. Calling this LLM observability implies there is something special about it and unique to LLMs in particular that enhances introspection into the AI model itself, which is not true. The title is highly misleading, classic startup-bro fake-it-til-you-make-it hustling crap, and deserves to be called out.
Correct- the summary is misleading marketing. This is just normal system / service observability. What people mean by observability in the LLM context is specific.
I thought Observability in this context means the ability to introspectively make sense of why the LLM output what it did, which is a difficult problem because the model parameters are effectively an unintelligible morass of numbers. Does this help with that and if so how?
I guess I'm not sure that there's a practical difference between "It's marketing" and "This is what the market rate is for a CEO." In other words, firms need to pay this much because that's what other firms are paying (at least.) That's not marketing - that's just the unfettered market at work. Which is why it needs to be fettered. I agree with the salary cap idea, because the market will naturally keep raising the price without bound unless there's something to counteract that, and it is terrible for the labor market as a whole (and even the company as a whole) to be paying so much for a CEO when the same amount could buy dozens or even hundreds of workers.
Not buying it. Tesla has a huge share of a growing pie. They will continue to grow as the EV market grows. While I agree that the past few months were not good, the long-term outlook for EVs is positive and therefore Tesla's outlook is positive.
Just because we have all decided we hate Elon Musk doesn't mean everything he has done is bad and destined to fail.
Oh give me a break. That’s so condescending to African Americans. If African Americans bought a certain kind of shoe, it’s because they liked the shoes. It’s fun to try shoes. People make their own choices.
To be sure, this kind of research, whether ‘craft’ or ‘natural’ is the correct word, is simply too risky to continue. The juice is not worth the squeeze.
Those companies largely made the decision for monorepos before there was tooling to support a proper multi-repo existence.
Literally just flew on a United Airlines 737 MAX 9 one week ago. It seems like the craft I flew on has probably been grounded in the week since. I noticed that we were flying on a MAX before boarding and nearly asked to switch flights, but consoled myself that I was being irrational and that the planes were almost certainly fine now. Guess my confidence was misplaced.
It's not applicable at all once you get beyond a few teams.
A mono-repo makes decisions for your teams.
However, it would also be folly to assume good will.
I actually don’t want the developer experience of co-location. There are millions of things that are totally irrelevant happening in my company’s (thousands of engineers) monolith. The noise in the commit log is considerable. Isolated repos are smaller, and reduce useless coupling.
So "the fewer people are buying electric cars" canard is still living in a caption under the banner: "Fewer people are buying electric cars — the slowdown hints at a problem at the heart of America's EV push."
Sorry, but in what world do sales of electric cars going up and market share of electric vehicles increasing lead to the thesis that "fewer people are buying electric cars?"
Oh - I get it. It's the world where this publication wants clicks for ad revenue.
"Sure, electric vehicles are becoming more and more widely adopted, but wouldn't it be better for this article if they weren't?"
While I don’t doubt that Threads addition of ActivityPub is good for both Threads and the fediverse, I wonder if it will become federated in the sense that email is still federated, which is to say email is dominated by a small number of players who control the vast majority of email addresses (i.e gmail) and their control represents a de-facto centralization of a once decentralized protocol. Is there a reason to think that ActivityPub is fundamentally different from SMTP, and Threads from Gmail, to suggest that won’t happen here?
It’s still preferable to have an open protocol, but only slightly if the related market is monopolized.
The board he drew was totally unintelligible but also immediately told me that Netflix didn’t or doesn’t know what its actual domains are. Adding functionality like this, even in a microservices world, should have been trivial and obvious, but nobody thought ahead and now they’ve ended up with a spiderweb of tangled services that have unclear responsibilities.
An advertised price that is racing to the bottom by offloading more and more of the cost to tips from customers.
Meanwhile the employees blame the customers, the customers blame the corporation, and the corporation blames the employees.
The corporation is the only one that is laughing all the way to the bank.
Just to clarify, I'm talking about integration testing the service itself - posting a payload, saving to the database, producing messages to a mock queue, etc. Test the entire service in isolation from other services and validate that it is behaving correctly. Not end to end across services. You should be mocking out all service dependencies and testing against the contracts for those systems.
Our API tests run flawlessly every time because they write against an isolated database with well-defined endpoint and messaging contracts. They also execute all remote operations against a mock API that conforms to those contracts. This is perfectly achievable.
I can think of few better investments than to have a reliable suite of end-to-end tests running as part of your merge pipeline even if they're difficult to set up. Sleep quality improves once you have this in place. Your tests don't have to be brittle or flaky. If your tests are brittle it is definitely a good investment to fix whatever's causing the brittleness rather than accept it as a fact of life. Having the tests run against an isolated data and infrastructure environment without additional noise from shared activity is a good first step.
High level architects should focus primarily on creating a system architecture that is designed to incrementally and composably add functionality.
This often means adopting some sort of overarching architecture that nudges people towards composable implementations. For write-heavy, highly stateful systems with complex business logic, this means using something like a workflow engine where you can simply declaratively add tasks and conditions to pre-existing DAGs while being confident the existing workflow will not fundamentally change. To create new functionality, it's often enough to use existing functionality as a drop-in template. Duplicate and add / remove - no need to worry about what's there because you're not touching it.
For read heavy systems thinking about composition at the very highest levels is also very important. This allows engineers to easily add functionality without being concerned about breaking what's already there.
Always favor additive models over models where updating existing functionality is the norm. This means junior engineers can "color within the lines" so to speak, and the risk to the rest of the system is low.
Would also emphasize that while unit testing is often useful for the individual developer working on a piece of code, but as far as ROI, end-to-end integration testing has the most bang for the buck. If you have the full endpoint tested, for instance, from end-to-end, including database writes, your confidence level goes up by an order of magnitude when you have to modify that functionality. If you have to choose between investing in extensive unit testing and extensive end-to-end API or contract tests, always choose the latter.
Then they can make their own sandbox that doesn’t treat their readers badly. But I see no signs of that.
Yes, I'm an Apple News subscriber, and it is the closest thing to what I'm talking about, no doubt because Apple is the only company with enough sway to convince publishers this is in their own interests. But even Apple can't push them nearly far enough. Having to look up the same article in Apple News when I'm paywalled through the web site even though I can legitimately access it is extremely annoying. These publications are definitely not going out of their way to take me to my legitimate, paid-for Apple News copy of their article - and I'm sure that's by design. Secondly, only a subset of publications are available. I want there to be equal access to all articles. But media does not want to compete on a level playing field like that.
I can only speak for myself, but here is what I would happily pay for: I would happily pay up to 20 cents to read a single article, ad free (or even, I suppose, not ad free), and for perpetual access to that article. And I would happily pay up to 50 bucks a month for all of the articles I read from all publications combined. What I will not do is purchase 20 subscriptions with recurring charges to 20 different media publications, none of which let me read articles from any other publication, and then remember to cancel them and reactivate them on a continuous basis so I can fly through paywalls. What I will also not do is hand over a bunch of personal contact information for the privilege of reading a single article, subjecting myself to all manner of unsolicited marketing emails. I have a feeling I'm not alone.
The sooner we can get to where this is all going, which I believe is tiny, contactless, trust-less, anonymous micropayments on a per-page basis, not a coercive subscription model built on dark patterns like recurring charges and call-to-cancel, the better it will be for everyone . Paywalls and other gating have totally ruined the user experience of the Internet and it is totally orthogonal to how the Internet and its protocols were designed to function. The sooner media realizes that they have to band together to adopt some sort of per-page or per-view micropayments solution in order remain aligned with the content model baked into the web's architecture, the sooner they will figure out how to support their profession with a continuous revenue stream and also, yes, even profit.
And for those who say that will warp the incentives of news reporting: they are already warped by page views and ad revenue. It's self-deluding to think that subscriptions somehow prevent that.