HN user

mumblemumble

17,246 karma
Posts11
Comments3,020
View on HN

Perhaps only if you can also be very certain that the output is correct whenever the logprobs don't trigger the filter.

If that's not the case then it might just trigger bad risk compensation behavior in the model's human operators.

I'm not an expert, either, but I've poked at this a little. From what I've seen, token logprobs are correlated enough with correctness of the answer to serve as a useful signal at scale, but it's a weak enough correlation that it probably isn't great for evaluating any single output.

My best guess is that somewhere close to the root of the problem is that language models still don't really distinguish syntagmatic and paradigmatic relationships. The examples in this article are a little bit forced in that respect because the alternatives it shows in the illustrations are all paradigmatic alternatives but roughly equivalent from a syntax perspective.

This might relate to why, within a given GPT model generation, the earlier versions with more parameters tend to be more prone to hallucination than the newer, smaller, more distilled ones. At least for the old non-context-aware language models (the last time I really spent any serious time digging deep into language models), it was definitely the case that models with more parameters would tend to latch onto syntagmatic information so firmly that it could kind of "overwhelm" the fidelity of representation of semantics. Kind of like a special case of overfitting just for language models.

I'm not sure it's easy to understand what a big change there has been in the perceived pace of computer technology development if you weren't there. I'm typing this on a laptop that I purchased 11 years ago, in 2013. It's still my one and only home computer, and it hasn't given me any trouble.

In 1994, though, an 11 year old computer would already be considered vintage. In 1983 the hot new computer was the Commodore 64. In 1994 everyone was upgrading their computers with CD-ROM drives so they could play Myst.

Testable hypotheses are at the core of the scientific method, yes. But that's not just limited to the actual testing of hypotheses. All the work that goes into formulating hypotheses is also explicitly part of the scientific method.

Worth noting, too, that the paper outlines several possible experiments. It also specifically mentions some relative shortcomings of the model, and lists existing observations that they haven't tried to reconcile with it yet.

I think that similar arguments were made about Ma Bell and Bell Labs back in the day. And it's true, a lot of great things did come out of Bell Labs.

In fact, it almost seems like the only people able to produce great things in the 1970s were massive entrenched corporations like Ma Bell.

Funny, that.

Come to think of it, wasn't there a much more vibrant browser ecosystem in the late 90s and early 2000s, before Google used its dominant position in the ad market to undercut the competition? There used to be a lot more mobile operating systems out there, too.

I wonder what happened to all that competition? It's almost like some sort of massive anti-competitive influence came into force in the tech scene somewhere in the 2000s. . .

The central premise of the article seems to me to be a likely misunderstanding of the problem. I would bet that a Scrum Product Owner who uses a "command and control" leadership style and doesn't know how to properly delegate authority would be doing the same under any other development framework, too.

In general I'm a big fan of the "single wringable neck" principle. I've seen it put to great effect in the hands of a skilled leader, in both Scrum and non-Scrum teams. Better yet, when the leader isn't managing things well, it also leaves no question that they're the one who needs to figure out how to set things straight. Same goes for their delegates.

And for the ICs it makes collaboration easier - and therefore, ironically, enables them to work more autonomously. When everyone unambiguously knows what they're in charge of and what the team's big-picture objectives are, they have everything they need to independently figure out how best to make it happen. And it's also a lot easier to figure out who to talk to when they need to call attention to a problem.

I've seen a lot less luck with shared authority and informal delegation. On the best of days, it turns decisionmaking into an unnecessarily political process. More likely, the team will settle into an informal consensus process that typically operates as "rule by the obdurate" in practice. And when things get tough, the leaders will tend to slide into unproductive bickering that all but precludes actually fixing the problem.

Favorite readings that touch on this kind of thing: The Tyranny of Structurelessness by Jo Freeman, and Turn the Ship Around! by David Marquet.

The old chestnut about AI just being a term for things we haven't quite figured out yet might apply here. "Products that are well-known... and are used on a daily basis by a significant amount of people" are almost by definition not AI.

But here are some examples of things that used to fall under the AI umbrella but don't really anymore:

  - Fulltext search with decent semantic hit ranking (Google)
  - Fulltext search with word sense disambiguation (Google)
  - Fulltext search with decent synonym hits (Google)
  - Machine translation
  - Text to speech
  - Speech to text
  - Automated biometric identification (Like for unlocking your phone)
If you're more specifically asking for everyday applications of GPT-style generative large language models, I don't think that's going to happen for cost reasons. These things are still far too expensive for use in everyday consumer products. There's ChatGPT, but it's kind of an open secret that OpenAI is hemorrhaging money on ChatGPT.

But do I even want an actually smart Siri?

Microsoft's been trying to ram essentially that down my throat for the better part of a year now, and it's mostly convinced me that the answer is "no". I don't want to have arbitrary conversations with my computer.

I still just want the same thing I've been wanting from my digital assistant for 30 years now: fewer "eat up Martha" moments, and handling more intents so that I can ask "When does the next east-bound bus come?" and it stops answering questions like "Will it rain today?" as if I had asked "Is it raining right now?". None of those are particularly appropriate problems for a GPT-style model.

Clarke's Third Law is, has always been, and always will be the best explanation for how futurists think about these kinds of things.

Moreover, this is exactly the frustration I've experienced when working with outsourced developers.

Which tells me the problem may be fundamental, not a technical one. It's not just a matter of needing "more intelligence". I don't question the intelligence or skill of the people on the outsourced team I was working with. The problem was simple communication. They didn't really know or understand our business and its goals well enough to anticipate all sorts of little things, and the lack of constant social interaction of the type you typically get when everybody's a direct coworker meant we couldn't build that mind-meld over time, either. So we had to pick up the slack with massive over-specification.

It's important to accept that you will screw it up. Repeatedly. Interfaces have to be designed before you can start using them, which means that you will never have less information about how a module will be used than you do when you design its interface.

The best defense against this that I've found is to ensure, as much as possible, that interfaces can be replaced. The single responsibility and interface segregation principles can help here. Using small, focused interfaces and letting modules implement more than one of them makes it easier to use the strangler pattern to replace interfaces that no longer work well with new and improved ones.

Also avoid temporal coupling as much as is feasible. Unnecessary statefulness is the easiest way to make this sort of thing harder than it needs to be.

Similar experience here. I went to four high schools. The best ones, academically, were the bog-standard schools that people love to use as an easy punching bag.

The ones where classes were a breeze and barely challenging at all were the private college prep school and the not-quite-as-rich-but-still-pretty-wealthy high school in the suburban white flight community. My read on the situation was that teachers had figured out that every single parent believed their kid deserved a constant stream of gold stars, and had the leisure time and resources to make their lives miserable until it started happening. Also the white flight school cared a lot about its football team so you better make for damn sure that the classes aren't so challenging that the quarterback has trouble balancing their time spent on homework, practice, and partying.

My favorite thing I've ever heard said by the principal of our kids' elementary school: "Our test scores are down, which is great, maybe that will keep some of the school shoppers away this year."

Our city has a school choice program that includes a portal where you can look up these kinds of quantitative measures, and I think I agree with him. Tiger parents slosh from school to school as they chase after rankings, and, much like ill-contained liquid cargo in ships, all that motion tends to destabilize and capsize schools.

Sadly, I don't think smaller higher education institutions can afford to take such a relaxed attitude about it. They don't get to have an enrollment backstop in the form of a semi-captive audience of parents who live nearby and aren't hyperactive enough to commit to spending upwards of an hour every weekday trucking their kids back and forth across town.

Why Haskell? 2 years ago

It is. But I think that, for that purpose, I like F# even better. Even beyond getting access to the .NET ecosystem, you also get some language design decisions that were specifically meant to make it easier to maintain large codebases that are shared among developers with varying skill levels.

Lack of typeclasses is a good example. Interface inheritance isn't my favorite, but after years working as the technical lead on a Scala project I've been forced to concede that haranguing people who just want to do their job and go home to their family about how to use them properly isn't a good use of anyone's time. Everyone comes out of school already knowing how to use interfaces and parametric polymorphism, and that is fine.

The ruling's section on transformativeness explains the distinction. Note that "derivative works" under US copyright law works differently from how it gets defined in typical open source licenses.

My understanding is that, for the purposes of determining fair use, a derivative work is substantially the same thing but in a different format. Transformative work must involve significant additional creative contribution "Changing the medium of a work is a derivative use rather than a transformative one." They cite previous case law that holds repackaging a print book as an e-book as a "paradigmatic example of a derivative work." The law also offers some paradigmatic examples of transformative work, such as criticism, commentary and scholarship.

Based on all of that, I would guess that, for the purposes of copyright law, a JPEG of a painting is absolutely a derivative work and not a transformative one.

We've also got to think about the actual value of preserving all of these works in a completely indiscriminate manner. Curation is important. Even assuming, for the sake of argument, that we could keep everything forever, actually doing so would ultimately harm the value of the archive, due to Sturgeon's Law. The truth is that the vast majority of cultural output is of only ephemeral value. It's relevant to a place and a time, but not necessarily great enough to also be interesting to people from a different place and a future time.

And I've only got a little bit of time in this life; I'd much rather read a trashy romance novel that was written this year and meant to entertain me than the trashy romance with politics that make me cringe that my mom was reading 50 years ago.

This is why, for example, the Library of Congress doesn't just keep a copy of everything. It's not just a space constraints or storage costs issue; it's a signal-to-noise ratio issue. As Mark Crislip is fond of saying, when you mix apple pie and cow pie it doesn't make the cow pie better, it just makes the apple pie worse.

The ruling discusses this starting on page 33. The gist is that they set up a non-transformative service that is substantially equivalent to competing ebook services and CDLs, but unlike those it is not paying the customary price to publishers.

It also discusses that there is a very good reason why digital libraries don't typically get to have perpetual rights to a work at the retail (or used) price for a print book. Basically, physical books wear out with use, ebooks don't, so there's a built-in mechanism for revenue recurrence that happens with print books but not ebooks. The ruling points out that publishers originally sold ebooks to libraries at the same pricing as print books, but abandoned the practice because they discovered that it was not financially sustainable.

And that's ultimately where the harm comes in. The IA is trying to create a loophole that subverts the income stream of all the people who work on a book by offering derivative works - which are never fair use; fair use is for transformative works - without paying the market's customary price for acquiring rights to create and distribute derivative works.

(As an aside, when I see authors speaking for themselves on these sorts of issues they will typically point out that editors and typesetters and cover artists and all the other folks who work on a book also deserve to get paid. It seems to only be people who are tokenizing authors for rhetorical purposes who want fixate on authors specifically and erase the value-adding contributions of "the publishers".)

I'm in both groups and I'm getting tired of how Python's syntax additions increasingly turn it into a mechanism for mid-level engineers to assert their dominance over colleagues by writing code that only people who have been deeply immersed in Python for years can understand.

Our Python users' Slack channel at work is already overcrowded with messages to the effect of, "halp what's this syntax how does this code work."

Its predecessor launch system had 10 successes and 2 failures for the crewed flights. One, they got the crew home safely, but it was close. So that's a 17% failure rate and a 8% rate of failures killing the crew.

Not saying the Shuttle's success rate was awesome; I'm glad we demand more nowadays. But it still represented a pretty decent crew safety improvement for the USA's human spaceflight program.

I wish code like this still felt normal to me, but over the past ~10 years it seems that many people have come to value brevity over explicitness.

I strongly prefer the explicitness, at least for important code like this. More than once in my career I've encountered situations where I couldn't figure out if the current behavior of a piece of code was intentional or accidental because it involved logic that did things like consolidating different conditions and omitting comments explaining their business context and meaning.

That's a somewhat dangerous practice IMO because it creates code that's resistant to change. Or at least resistant to being changed by anyone who isn't its author, though for most practical purposes that's a distinction without a difference. Unnecessarily creating Chesterton's fences is anti-maintainability.

I have a family member who used to be the director of an urban area's primary emergency department. (Now retired.)

At some point he and his administrative staff saw a similar trend, and also figured out that their top 10 most expensive patients were all people who had chronic health conditions and no health insurance or medicaid. Not having health insurance meant they couldn't effectively manage their conditions, which meant they were repeatedly getting to a crisis point and having a family member call an ambulance. Which is just incredibly expensive - all emergency services are - and of course if they couldn't pay out of pocket for a family physician then they couldn't pay for emergency services, either.

The solution was to start just paying out of the ED's own budget for these folks to see regular care providers, buying them their medications, arranging cab rides, etc. It saved the hospital millions per year.

I don't think you need that specific interpretation to decide that this work reveals something important. We have a phenotypic difference that leads to a big in vitro difference, and it's not a huge jump to infer that that would also lead to an in vivo difference, even if the in vivo presentation is unknown, or the exact details of the presentation aren't the same.

Here's the conclusions section from the research paper this article is summarizing:

By embryogenesis, the biological bases of two subtypes of ASD social and brain development - profound autism and mild autism — are already present and measurable and involve dysregulated cell proliferation and accelerated neurogenesis and growth. The larger the embryonic BCO size in ASD, the more severe the toddler’s social symptoms and the more reduced the social attention, language ability, and IQ, and the more atypical the growth of social and language brain regions.

This is not making any huge logical jumps that I noticed. All it's saying is they found a strong correlation. And the researchers seem to be well aware that there are still dots to connect. In the limitations section, it explicitly points out more-or-less the very thing that the researchers are being accused of not thinking about in this HN thread:

The genetic causes and cellular consequences of decreased Ndel1 activity and expression correlated with ASD BCOs enlargement remain to be specified. A limitation of most previous ASD patient-derived iPSC-based models is lack of within-subject statistical linkage of ASD molecular and cellular findings with variation in ASD social phenotypes. Without this, future ASD iPSC reports will continue to have limited impact on our understanding of the genetic, molecular and cellular mechanisms that cause the development and variation in the central feature of ASD: social affect and communication.

Which brings us to an important thing about interpreting popular science literature: it's unwise to assume that what's in the popularization of the research accurately reflects everything the scientists who published the work think or know. Attempting to eliminate these kinds of details is one of the primary goals of science journalism. For better or for worse.

https://molecularautism.biomedcentral.com/articles/10.1186/s...