HN user

benjismith

760 karma
Posts18
Comments96
View on HN
duiduidui.app 9d ago

Show HN: Duiduidui is a new Chinese dictionary and flashcard app

benjismith
2pts0
machinecreativity.substack.com 4mo ago

Your Brain Is a Ball of Electric Spaghetti

benjismith
2pts0
github.com 4mo ago

I asked Claude for 37,500 random names, and it can't stop saying Marcus

benjismith
91pts72
machinecreativity.substack.com 4mo ago

Benji's Guide to Machine Creativity

benjismith
1pts0
chatgpt.com 1y ago

Show HN: Hypnotizing ChatGPT

benjismith
2pts0
www.prosecraft.io 8y ago

Show HN: Five Thousand Novels, Ranked by Vividness

benjismith
54pts27
prosecraft.io 8y ago

Five Thousand Novels, Ranked by Vividness

benjismith
1pts0
blog.shaxpir.com 8y ago

My Philosophy of Subscription Software

benjismith
1pts0
blog.shaxpir.com 8y ago

Prosecraft: Linguistics for Literature

benjismith
1pts0
solo-founder.com 9y ago

Why I Decided to Go Solo

benjismith
1pts0
www.solo-founder.com 9y ago

Solo Founder Magazine

benjismith
1pts0
blog.shaxpir.com 9y ago

Founder and Chief Everything

benjismith
1pts0
news.ycombinator.com 9y ago

Show HN: Shaxpir 4: Everyone

benjismith
1pts0
vimeo.com 9y ago

Reflection VOID: Time Lapse Milky Way with Motion-Control Mirrors

benjismith
1pts0
orchestrate.io 11y ago

Geospatial Search Adds Location to Your Applications

benjismith
14pts0
orchestrate.io 11y ago

Bring Order to Search Results with Field Sorting

benjismith
4pts0
orchestrate.io 11y ago

Schemalessness Gone Wrong: How We Improved Elasticsearch Indexing

benjismith
10pts0
orchestrate.io 12y ago

Server Generated Keys: Unique IDs for Distributed Databases

benjismith
10pts0

I used to work for a company (~50 people) whose entire business was based on the SBIR pipeline. We did a lot of super-interesting work!

Here's my advice:

1) Your proposal needs to be completely solid and well-structured:

- Describe the problem. Put it into a Defense/Intel context. Talk about the needs of the warfighter.

- Do a literature review of the field, and explain what the state-of-the-art looks like today.

- Explain what previous approaches to the problem have been attempted in the past.

- Demonstrate why those approaches are flawed.

- Describe your novel approach.

- Explain why your approach will succeed where others have failed.

- Talk about what you'll deliver in your Phase I deliverable, so that you can demonstrate proof-of-concept.

- Talk about how your eventual Phase II will put your proof-of-concept into a real-world scenario, and offer at least a glimpse of how your Phase III+ will commercialize.

- Talk about your team. Why are you uniquely capable of solving this problem?

- Talk about your budget. How will you spend the money toward satisfaction of the deliverables (salaries, subcontractors, equipment and supplies, etc)

2) As soon as you know you're interested in a topic, send an email to the Principal Investigator, telling them you'd like to meet with them to talk about the topic. Before the phone call, research the PI's history with this topic. Also, lookup the archive of SBIR topics, to see if this person has been a PI on similar topics in the past.

When you meet with them, ask clarifying questions that demonstrate you know the domain. Try to get as much specificity as you can... Ask them what their success criteria look like. See if you can get them excited!

Most importantly, by the time you submit, the PI should already know your name and to expect your submission.

3) If you're not already a recognized expert, with published academic papers on the topic, that's okay! But you'll improve your chances of winning a grant if you hire a known researcher as an advisor. For example, I've hired a Computer Science professor to supervise one of their own grad students, while doing paid work on a SBIR project. So the professor's credentials and the grad student's previous publications also became part of the SBIR proposal.

Anyhow, good luck!

It makes perfect sense.

Apple's philosophy is that new APIs need some time to stabilize before they can be baked-in as a commitment to third-party developers.

So new APIs are almost always first-party only. Apple designs the API and becomes the first consumer of it. This experience of dogfooding their own APIs lets them iterate and learn without breaking compatibility with third-party developers consuming the API.

Only after an API has been hardened in this way does it become eligible for third-party consumption, where Apple can promise to document and support those APIs publicly.

It makes sense then, that if the DMA mandates equal access to new APIs for third-parties, then Apple will just disable new first-party APIs in the region until they've gotten their bake-in period elsewhere in the world. Sorry, EU!

I thought this pull-quote was interesting:

"Interoperability only works when it is built into the platform from the start"

-- Lucas Lasota, FSFE Legal Programme Manager

To my mind, this is almost exactly opposite of true. Most new capabilities need to be incubated in private first, so that the APIs can get real-world usage and have a chance to evolve into a stable state before they become public interoperability promises.

This is true, but I also think the input context isn't the only function of those tokens...

As those tokens flow through the QKV transforms, on 96 consecutive layers, they become the canvas where all the activations happen. Even in cases where it's possible to communicate some detail in the absolute minimum number of tokens, I think excess brevity can still limit the intelligence of the agent, because it starves their cognitive budget for solving the problem.

I always talk to my agents in highly precise language, but I let A LOT of my personality come through at the same time. I talk them like a really good teammate, who has a deep intuition for the problem and knows me personally well enough to talk with me in rich abstractions and metaphors, while still having an absolutely rock-solid command of the technical details.

But I do think this kind of caveman talk might be very handy in a lot of situations where the agent is doing simple obvious things and you just want to save tokens. Very cool!

I think the biggest injury to the hiring of junior devs happened after COVID made remote-work ubiquitous. It's a lot harder for a junior dev to get real mentorship, including the ambient kind of mentorship-by-osmosis, when everyone works alone in a sad dark room in their basement, rather than in an office with their peers and mentors.

The advent of agentic coding is probably punch #2 in the one-two punch against juniors, but it's an extension of a pattern that's been unfolding for probably 5+ years now.

I read it. I also searched the page for the word "Opus" and it didn't appear anywhere. The word "Sonnet" appears, but only once.

There's also "GPT-4.1 or GPT-5", but that's not what my question implied, which was that it's weird to offer Sonnet but not Opus.

A similar kind of question about "understanding" is asking whether a house cat understands the physics of leaping up onto a countertop. When you see the cat preparing to jump, it take a moment and gazes upward to its target. Then it wiggles its rump, shifts its tail, and springs up into the air.

Do you think there are components of the cat's brain that calculate forces and trajectories, incorporating the gravitational constant and the cat's static mass?

Probably not.

So, does a cat "understand" the physics of jumping?

The cat's knowledge about jumping comes from trial and error, and their brain builds a neural network that encodes the important details about successful and unsuccessful jumping parameters. Even if the cat has no direct cognitive access to those parameters.

So the cat can "understand" jumping without having a "meta-understanding" about their understanding. When a cat "thinks" about jumping, and prepares to leap, they aren't rehearsing their understanding of the physics, but repeating the ritual that has historically lead them to perform successful jumps in the past.

I think the theory of mind of an LLM is like that. In my interactions with LLMs, I think "thinking" is a reasonable word to describe what they're doing. And I don't think it will be very long before I'd also use the word "consciousness" to describe the architecture of their thought processes.

Is there way to get "speech marks" alongside the generated audio?

FYI, Speech marks provide millisecond timestamp for each word in a generated audio file/stream (and a start/end index into your original source string), as a stream of JSONL objects, like this:

{"time":6,"type":"word","start":0,"end":5,"value":"Hello"}

{"time":732,"type":"word","start":7,"end":11,"value":"it's"}

{"time":932,"type":"word","start":12,"end":16,"value":"nice"}

{"time":1193,"type":"word","start":17,"end":19,"value":"to"}

{"time":1280,"type":"word","start":20,"end":23,"value":"see"}

{"time":1473,"type":"word","start":24,"end":27,"value":"you"}

{"time":1577,"type":"word","start":28,"end":33,"value":"today"}

AWS uses these speech marks (with variants for "sentence", "word", "viseme", or "ssml") in their Polly TTS service...

The sentence or word marks are useful for highlighting text as the TTS reads aloud, while the "viseme" marks are useful for doing lip-sync on a facial model.

https://docs.aws.amazon.com/polly/latest/dg/output.html

If I'm reading the pricing correctly, these models are SIGNIFICANTLY cheaper than ElevenLabs.

https://platform.openai.com/docs/pricing

If these are the "gpt-4o-mini-tts" models, and if the pricing estimate of "$0.015 per minute" of audio is correct, then these prices 85% cheaper than those of ElevenLabs.

https://elevenlabs.io/pricing

With ElevenLabs, if I choose their most cost-effectuve "Business" plan for $1100 per month (with annual billing of $13,200, a savings of 17% over monthly billing), then I get 11,000 minutes TTS, and each minute is billed at 10 cents.

With OpenAI, I could get 11,000 minutes of TTS for $165.

Somebody check my math... Is this right?

Nope. They needed the maximum amount of thrust from those boosters in order to propel the spacecraft toward Jupiter, so they couldn't save enough fuel for the boosters to land themselves. This was the 6th flight of these boosters, so we thank them for their service!

Awesome, I'm been following Seph's work for many years! Always thoughtful and well-executed. Probably the most prolific and insightful engineer in the "collaborative text editing" universe.

I use ShareDB every day, which originated from Seph's excellent work on OT algorithms. Good stuff!

It doesn't sound like they "failed" any actual safety test, but rather that they rushed their safety tests, thereby "failing" (in the eyes of many people) to conduct sufficiently rigorous tests.

Now that the 4o model have been out in the wild for 2 months, have there been any claims of serious safety failures? The article doesn't seem to imply any such thing.

I find myself wondering if Apple applied some kind of back-channel pressure to oust Riccitiello.

With Unity at such a privileged position in the developer ecosystem of the upcoming Apple Vision Pro, I can imagine that Apple execs were pissed off that Unity would do something so stupid and shortsighted to jeopardize their developer ecosystem.

I haven't heard anyone float that idea yet, and the term "Apple" doesn't appear anywhere (yet!) in the comments of this post, so it doesn't seem to be on most people's minds. But still, I wonder...

I'm a software engineer, and I think the fully remote-work culture can often be less personally fulfilling.

I enjoy going to an office, having a change of scenery, interacting with people (both close friends and casual acquaintances), brainstorming ideas with a small group around a whiteboard, getting lunch with colleagues, etc, etc...

Sure, the "writing code" part of the job is easier when there are fewer distractions. And a crowded open office can be an annoying source of distractions.

And I definitely enjoy having lunch more often with my family, or hanging out with the dogs, or sitting out on my deck under the trees while I read my morning email.

But working in an office was nice too, and I miss it.

I'm still trying to get a handle on that part myself... But my ever-evolving understanding goes something like this:

The "Query" matrix is like a mask that is capable of selecting certain kinds of features from the context, while the "Key" matrix focuses the "Query" on specific locations in the context.

Using the Query + Key combination, we select and extract those features from the context matrix. And then we apply the "Value" matrix to those features in order to prepare them for feed-forward into the next layer.

There are multiple "Attention Heads" per layer (GPT-3 had 96 heads per layer), and each Head performs its own separate QKV operation. After applying those 96 Q+K->V attention operations per layer, the results are merged back into a single matrix so that they can be fed-forward into the next layer.

Or something like that...

I'm still trying to grok it myself, and if anyone here shed more light on the details, I'd be very grateful!

I'm still trying to understand, for example, how many QKV matrices are actually stored in a model with a particular number of parameters. For example, in a GPT-NeoX-20B model (with 20 billion params) how many distinct Q, K, and V matrices are there, and what is their dimensionality?

EDIT:

I just read Imnimo's comment below, and it provides a much better explanation about QKV vectors. I learned a lot!

Okay, here's my attempt!

First, we take a sequence of words and represent it as a grid of numbers: each column of the grid is a separate word, and each row of the grid is a measurement of some property of that word. Words with similar meanings are likely to have similar numerical values on a row-by-row basis.

(During the training process, we create a dictionary of all possible words, with a column of numbers for each of those words. More on this later!)

This grid is called the "context". Typical systems will have a context that spans several thousand columns and several thousand rows. Right now, context length (column count) is rapidly expanding (1k to 2k to 8k to 32k to 100k+!!) while the dimensionality of each word in the dictionary (row count) is pretty static at around 4k to 8k...

Anyhow, the Transformer architecture takes that grid and passes it through a multi-layer transformation algorithm. The functionality of each layer is identical: receive the grid of numbers as input, then perform a mathematical transformation on the grid of numbers, and pass it along to the next layer.

Most systems these days have around 64 or 96 layers.

After the grid of numbers has passed through all the layers, we can use it to generate a new column of numbers that predicts the properties of some word that would maximize the coherence of the sequence if we add it to the end of the grid. We take that new column of numbers and comb through our dictionary to find the actual word that most-closely matches the properties we're looking for.

That word is the winner! We add it to the sequence as a new column, remove the first-column, and run the whole process again! That's how we generate long text-completions on word at a time :D

So the interesting bits are located within that stack of layers. This is why it's called "deep learning".

The mathematical transformation in each layer is called "self-attention", and it involves a lot of matrix multiplications and dot-product calculations with a learned set of "Query, Key and Value" matrixes.

It can be hard to understand what these layers are doing linguistically, but we can use image-processing and computer-vision as a good metaphor, since images are also grids of numbers, and we've all seen how photo-filters can transform that entire grid in lots of useful ways...

You can think of each layer in the transformer as being like a "mask" or "filter" that selects various interesting features from the grid, and then tweaks the image with respect to those masks and filters.

In image processing, you might apply a color-channel mask (chroma key) to select all the green pixels in the background, so that you can erase the background and replace it with other footage. Or you might apply a "gaussian blur" that mixes each pixel with its nearest neighbors, to create a blurring effect. Or you might do the inverse of a gaussian blur, to create a "sharpening" operation that helps you find edges...

But the basic idea is that you have a library of operations that you can apply to a grid of pixels, in order to transform the image (or part of the image) for a desired effect. And you can stack these transforms to create arbitrarily-complex effects.

The same thing is true in a linguistic transformer, where a text sequence is modeled as a matrix.

The language-model has a library of "Query, Key and Value" matrixes (which were learned during training) that are roughly analogous to the "Masks and Filters" we use on images.

Each layer in the Transformer architecture attempts to identify some features of the incoming linguistic data, an then having identified those features, it can subtract those features from the matrix, so that the next layer sees only the transformation, rather than the original.

We don't know exactly what each of these layers is doing in a linguistic model, but we can imagine it's probably doing things like: performing part-of-speech identification (in this context, is the word "ring" a noun or a verb?), reference resolution (who does the word "he" refer to in this sentence?), etc, etc.

And the "dot-product" calculations in each attention layer are there to make each word "entangled" with its neighbors, so that we can discover all the ways that each word is connected to all the other words in its context.

So... that's how we generate word-predictions (aka "inference") at runtime!

By why does it work?

To understand why it's so effective, you have to understand a bit about the training process.

The flow of data during inference always flows in the same direction. It's called a "feed-forward" network.

But during training, there's another step called "back-propagation".

For each document in our training corpus, we go through all the steps I described above, passing each word into our feed-forward neural network and making word-predictions. We start out with a completely randomized set of QKV matrixes, so the results are often really bad!

During training, when we make a prediction, we KNOW what word is supposed to come next. And we have a numerical representation of each word (4096 numbers in a column!) so we can measure the error between our predictions and the actual next word. Those "error" measurements are also represented as columns of 4096 numbers (because we measure the error in every dimension).

So we take that error vector and pass it backward through the whole system! Each layer needs to take the back-propagated error matrix and perform tiny adjustments to its Query, Key, and Value matrixes. Having compensated for those errors, it reverses its calculations based on the new QKV, and passes the resultant matrix backward to the previous layer. So we make tiny corrections on all 96 layers, and eventually to the word-vectors in the dictionary itself!

Like I said earlier, we don't know exactly what those layers are doing. But we know that they're performing a hierarchical decomposition of concepts.

Hope that helps!

No, I don't pay for it myself. My employer setup the WeWork and pays the monthly bill, on the premise that there would be a handful of people in our city (Portland, OR) who also want to work from an office. There have a been a handful of other people who used to come in occasionally, but now it's dwindled down to mostly just me.

I also ride my bike ~10 miles each way (even in Portland winter!) because it's a great way for me to make sure I get daily exercise and breathe some fresh air.

Without the commute, I get cabin-fever in the dreary Portland winter.

I have a pretty robust social life with my wife and other non-work friends, but if I work from home every day, I start to get some major cabin fever...

And I specifically miss the PROFESSIONAL SOCIAL LIFE that comes from having a close face-to-face relationship with my collaborators.

I must be one of the very very few software engineers that would prefer to RTO. (I'm not at Amazon, btw).

I do enjoy the flexibility of being able to occasionally work from home. But I miss the bustling office culture, going out to lunch with coworkers, forming friendships with people from the office who don't necessarily work in the same department/team.

Most days, I go into town and work from a WeWork (because it's nice to have a daily change of scenery), but 95% of the time, I'm the only person there (in an office with nine desks). Before the pandemic there were 80~100 people in our office.

Sigh...

Okay, somebody posted a thread on Twitter explaining how this works...

The language model is capable of generating python scripts to solve certain text-processing tasks, and then it re-prompts itself by reading the python outputs back into the language model. Very clever!

https://twitter.com/goodside/status/1598253337400717313

Other tricks include... prompting itself to lookup wikipedia entries, and then re-prompt itself with snippets from the resulting wikipedia page. Each user prompt is inserted into a template prompt with instructions to the model about the limitations of its capabilities.

Lol, good point!

I just meant "this isn't related to Thinking Fast and Slow. It's just the tokenizer".

But yeah, the inner workings of the language model are so complicated as to be almost completely incomprehensible, even after years of study. Touche!

It's really not so complicated. This is just an issue with text tokenization, and the fact that the learning model never actually sees the raw input bytes.

All modern LLMs use a tokenizer to convert a sequence of bytes into a sequence of tokens. Short, common words like "the" and "why" are represented as single tokens, while longer and less-common words are represented by multiple tokens. For example, the word "fantastic" is three tokens ("f", "ant", "astic").

Each of these tokens is assigned an arbitrary integer value ("fantastic" becomes [69, 415, 3477]) and then those integer values are used to lookup embedding vectors for each word.

Each embedding vector represents the MEANING of the tokens, by plotting them into a 4096-dimensional vector-space. At runtime, the model looks up each token ID in a dictionary and finds its embedding vector.

For the word "fantastic", those embedding vectors might look something like this:

  "f"        (69) = [  0.123,  0.456, ...etc...  0.789, -0.890 ]
  "ant"     (415) = [  0.111, -0.222, ...etc...  0.333, -0.444 ]
  "astic"  (3477) = [ -0.101,  0.202, ...etc... -0.303,  0.404 ]
All of these vectors are assembled into a matrix, and then passed into the layers of neural network, where the actual training/inference occurs.

So the language-model has NO IDEA how any of the words are spelled, because the tokenization (and embedding vector lookup) happens as a pre-processing step, outside the bounds of the learning algorithm.

If you want a LLM to understand spelling, you have to include exhaustive spelling information in its training data. For example:

  "The word 'fantastic' is spelled f-a-n-t-a-s-t-i-c."
  "The word 'FANTASTIC' is spelled F-A-N-T-A-S-T-I-C."
  ...etc...
And even then, even with 100k+ English words all spelled out in your training data, you'd be hard-pressed to infer any ROT-13 tokens in your output data, because the learning model has probably never seen a token like "qvq" or "pebff".

You can play with the GPT tokenizer directly here:

https://beta.openai.com/tokenizer

It will show you the tokenization of any block of text, and the token IDs of the resultant tokens. It's very handy if you spend much time working with GPT-3 (or any other modern language-model!)