HN user

Me1000

2,093 karma

Working on something new!

Previously worked on RunKit, Google Search App on iOS, and core team member of Cappuccino.

Posts13
Comments395
View on HN

Stopped using it after about a week or two of usage. The only interesting use case was screen mirroring from my Mac, but that wasn't compelling enough to endure the weight of it on my face. I expected watching a movie would be a good use case, but in reality the brightness of the screen would reflect (I guess?) off my face and create a glare... so it ultimately wasn't a good movie device. Gave it to a friend who was excited about it, and he also stopped using it after about a week or two.

A native application that further locks users into some single platform? Or accept all the maintenance and development costs and burdens that keep the application one step behind Photoshop if they wanted to support multiple platforms?

Wouldn’t it be better to use a grammar in the token sampler? Tuning is fine, but doesn’t guarantee a syntactical correct structured output. But if the sampler is grammar aware it could.

Why Objective-C 5 months ago

You can send messages to null, sendings messages to a deallocated pointer is going to be a bad time.

Orion 1.0 8 months ago

Hi! Congratulations on the launch. Is your intention to ship using WebKit on Window and Linux too?

Not OP, but yes believe it or not it's impossible to find certain movies anywhere other than pirating them. One example is "Pirates of Silicon Valley", I watched it when I was young and recently wanted to watch it again. I pay for basically all the streaming services, I'm would have been happy to rent it from any service at all. I spent several hours trying to find a way to pay to watch it and never could.

There’s an important distinction between the open weight model itself and the deepseek app. The hosted model has a filter, the open weight does not.

Yet it's interesting how we put the blame and punishment on the people being taken advantage of, and not the employers who are exploiting them. If both parties are breaking the law shouldn't we at the very least ensure that the business owner who is exploiting any number of workers is held to the same standard as an undocumented person whose only crime was not having the proper paperwork?

This is a technology demo, not a model you'd want to use. Because Bitnet models are only average 1.58 bits per weight you'd expect to need the model to be much larger than your fp8/fp16 counterparts in terms of parameter count. Plus this is only a 2 billion parameter model in the first place, even fp16 2B parameter models generally perform pretty poorly.

I'm also confused by that, but it could just be the model being agreeable. I've seen multiple examples posted online though where it's fairly clear that the COT output is not included in subsequent turns. I don't believe Anthropic is public about it (could be wrong), but I know that the Qwen team specifically recommend against including COT tokensfrom previous inferences.

This administration under through Elon is pushing to cut 50% of NASA's science funding. Mapping galaxies we'll never visit is a purely scientific endeavor. Trump seems to care more about military expansion or for lack of a better term more "masculine" expansion of space. The science stuff is not interesting to him, and I'm honestly not sure I think Musk cares about it that much anymore either.

I'm not a hater (or OP), but I think it's because GLP-1s are often talked about as if they're a kind of miracle drug. And it's very rare that drugs don't have some kind of side effects, especially after long term use. They might not even be universal, and we might not know what they are for years.

Yes, I understand there's an obesity epidemic, I also understand that GLP-1 drugs can have benefits outside of overeating. But with any drug, it's worth being thoughtful about its use.

DeepSeek-R1 2 years ago

Although I haven’t used these new models. The censorship you describe hasn’t historically been baked into the models as far as I’ve seen. It exists solely as a filter on the hosted version. IOW it’s doing exactly what Gemini does when you ask it an election related question: it just refuses to send it to the model and gives you back a canned response.

DeepSeek-R1 2 years ago

You’re just seeing a short summary of it, not the actual monologue.

The 32B parameter model size seems like the sweet spot right now, imho. It's large enough to be very useful (Qwen 2.5 32B and the Coder variant our outstanding models), and they run on consumer hardware much more easily than the 70B models.

I hope Llama 4 reintroduces that mid sized model size.

ChatGPT Search 2 years ago

It's not like ChatGPT (or ChatGTP as half of people call it) is much better.