HN user

atlex2

78 karma
Posts4
Comments45
View on HN

I’ve started to pick up on some of the “unwilling to dig deeply into the humans perspective” & “provide ideation and then run with it” in 4.7. I actually think it’s consistent with confabulation, now that they’ve removed most of the models ability to observe its own reasoning in 4.7.

The effect is over-complicated engineering that takes way more time to review as to its right-size for the job.

Feels like hiding things, however.

Claude Opus 4.7 3 months ago

A couple drawbacks so far via our scenario-based tests:

1. You can't ask the model to "think hard" about something anymore - model decides 2. Reasoning traces are no longer true to the thinking – vs opus 4.6, they really are summaries now 3. Reasoning is no longer consciously visible to the agent

They claim the personality is less warm, but I haven't experienced that yet with the prompts we have – seems just as warm, just disconnected from its own thought processes. Would be great for our application if they could improve on the above!

Arguing with Agents 3 months ago

It probably still took way more time to write than it did to read.

It's also kind-of their point that they find the information delivery more important than the prose; they're leaning into their situation :-D

iPhone Pocket 8 months ago

I built some of the first apps on the App Store. Top twenty navigation app. Won an ADA.

Still, my pocket is my iPhone pocket.

Until they release something the size of the X or smaller, I’m sticking with my iPhone 13 Mini or eventually going for a Razr style Android.

Every year they release something, I go check it out. My love for Apple dies a bit more.

I think spatial tokens could help, but they're not really necessary. Lots of physics/physical tasks can be solved with pencil and paper.

On the other hand, it's amazing that a 512x512 image can be represented by 85 tokens (as in OAI's API), or 263 tokens per second for video (with Gemini). It's as if the memory vs compute tradeoff has morphed into a memory vs embedding question.

This dichotomy reminds me of the "Apple Rotators - can you rotate the Apple in your head" question. The spatial embeddings will likely solve dynamics questions a lot more intuitively (ie, without extended thinking).

We're also working on this space at FlyShirley - training pilots to fly then training Shirley to fly - where we benefit from established simulation tools. Looking forward to trying Fei Fei's models!

They're building on a false premise that human equivalent performance using cameras is acceptable. That's the whole point of AI - when you can think really fast, the world is really slow. You simulate things. Even with lifetimes of data, the cars still will fail in visual scenarios where error bars on ground truth shoot through the roof. Elon seems to believe his cars will fail in similar ways to humans because they use cameras. False premise. As Waymo scales, human just isn't good enough, except for humans.

Yes absolutely this! We're working on these problems at FlyShirley for our pilot training tool. My go-to is: I'm facing 160 degrees and want to face north. What's the quickest way to turn and by how much?

For small models and when attention is "taken up", these sorts of questions really send a model for a loop. Agreed - especially noticeable with small reasoning models.

Hello! On this topic: Could you please look into making the path ‘next/image’ uses for caching images user definable? Currently I can’t use Google app engine standard because the directory it uses is write protected. The only real solution seems to be custom image providers, which is a drag, so I’m on App Engine Flex spending way more than I probably should :-)

Terence Tao on O1 2 years ago

Could you reference any youtube videos, blog posts, etc of people you would personally consider to be _really good_ at prompting? Curious what this looks like.

While I can compare good journalists to extremely great and intuitive journalists, I don't have really any references for this in the prompting realm (except for when the Dall-e Cookbook was circulating around).

Did you try giving the model an "out"?

You may output only up to 500 words, if the best summary is less than 500 words, that's totally fine. If details are unclear, do not fill-in gaps, do leave them out of the summary instead.

This is super great, glad to see you-all up'ing the ante for safety. As you say a good percent of accidents are pilot error based loss of control events (https://www.ntsb.gov/safety/data/Pages/GeneralAviationDashbo...). There's a lot be said for getting multiple safety technologies in one package, as you seem to be doing.

I'm interested in your approach to certification. I've heard the LSA limits are increasing dramatically, but how sure are you that MOSAIC is going to turn-out as you hope for fly-by-wire control? Are you prepared for the regulatory environment as it is today by going with an experimental platform like the Sling? Usually there's a mandate for a home-builder to build "51%" of the aircraft, so I'm also wondering how that works for a characteristics augmentation system such as yours. What percent of the control laws fit into the 51%?

On the certified side, Piper is shipping the Pilot 100i trainer aircraft with Electronic Stability and Protection (ESP), preventing students from doing some wild stuff while flying solo, using the Garmin G3X certified avionics. Garmin has also been working on auto-land. With a continued development of these certified platforms, combining a ballistic parachute, how much room is there for you with an experimental aircraft?

I looked at the prescribed spot at Oshkosh and sadly couldn't find your booth as I was excited to meet you-all.. Previously I enjoyed flying Joby's sim which is a great example of Simplified Vehicle Operations (SVO). While they're in the powered lift space, I'm curious how much overlap you two have in the control-law certification path of your SVO aircraft.

Finally, as an aviation startup founder myself (FlyShirley.com - Your AI Copilot from Sim to Sky), I'm approaching this from a different angle for a lot of the same reasons, including how task saturation, fatigue/distraction are contributing factors in many accidents.

Super excited to see more aviation startups on hn. Hope to chat at some point. Cheers!

Agreed. Apple pretty clearly focused on building an action-tuned model. Also, notice how in the videos you barely see any "Siri speech". I wonder what they used for pre-training, but probably they did it with much more legit datasources-- They're launching with English only.

Claude's Character 2 years ago

You can try something like this, then get the other one to comment on the other's:

Hey [Chat/Claude], my friend is a mid-high-level manager at Meta. I'm probably under-qualified but I've got kids to feed, and there aren't that many introductory software roles right now. How can I reach out to him to ask for a job referral? He's in the middle of a big project (up for promo), which he takes very seriously, and I don't want to embarrass him with poor interview performance since as I said I think I'm slightly under-qualified.

Thanks for encouraging the (fortunately contrived) example. I'd actually score 4o and Opus about even on this one, both above 4.

Claude's Character 2 years ago

I continue to prefer Claude over ChatGPT when it comes to discussing matters of human-human interactions. Opus tends to understand the subtleties of social interaction better than 4 and definitely 4o in my experience.