HN user

chad1n

343 karma
Posts0
Comments76
View on HN
No posts found.

Anyone who has some experience with native apis knows that a standard library should never rely on unstable apis. Ntdll is not "stable" as in Microsoft can change it at any time since they expect anyone to use kernel32. It's questionable that they referenced a random book on this top claiming that ntdll is more performant than kernel32 which is doubtful. There are some specific cases where this is true (the ntfs stuff), but, in general, it's not, at least not in a significant matter. A standard library should never do this, it might break binaries for no reason, other than making a cool blog post. I, as a developer, can choose to use ntfs, but a standard library should never.

https://news.ycombinator.com/item?id=25997506 https://github.com/golang/go/issues/68678

So the author is in a clear conflict of interest with the contents of the blog because he's an employee of Anthropic. But regarding this "blog", showing the graph where OpenAI compares "frontier" models and shows gpt-4o vs o3-high is just disingenuous, o1 vs o3 would have been a closer fight between "frontier" models. Also today I learned that there are people paid to benchmark AI models in terms of how close they are to "human" level, apparently even "expert" level whatever that means. I'm not a LLM hater by any means, but I can confidently say that they aren't experts in any fields.

OpenAI o3-pro 1 year ago

The guys in the other thread who said that OpenAI might have quantized o3 and that's how they reduced the price might be right. This o3-pro might be the actual o3-preview from the beginning and the o3 might be just a quantized version. I wish someone benchmarks all of these models to check for drops in quality.

To be honest, checking if there is a path between two nodes is a better example of NP-Hard, because it's obvious why you can't verify a solution in polynomial time. Sure the problem isn't decidable, but it's hard to give problems are decidable and explain why the proof can't be verified in P time. Only problems that involve playing optimally a game (with more than one player) that can have cycles come to mind. These are the "easiest" to grasp.

The idea is correct, a lot of people (including myself sometimes) just let an "agent" run and do some stuff and then check later if it finished. This is obviously more dangerous than just the LLM hallucinating functions, since at least you can catch the latter, but the first one depends on the tests of the project or your reviewer skills.

The real problem with hallucination is that we started using LLMs as search engines, so when it invents a function, you have to go and actually search the API on a real search engine.

These "OCR" tools who are actually multimodals are interesting because they can do more than just text abstraction, but their biggest flaw is hallucinations and overall the nondeterministic nature. Lately, I've been using Gemini to turn my notebooks into Latex documents, so I can see a pretty nice usecase for this project, but it's not for "important" papers or papers that need 100% accuracy.

It's not really a hot take, considering the price, they probably released it to scam some people when they to `benchmark` it or to buy the `pro` version. You must be completely in denial to think that gpt4.5 had a successful launch, considering that 3 days before, a real and useful model was released by their competitor.

I don't think that's the case, when a model is reasoning, it sometimes starts gaslighting itself and "solving" other problems completely than the one you've shown. Reasoning can help "in general", but very frequently, reasoning also makes it more "nondetermistic". Without reasoning, usually it ends up just writing some code from its training data, but with reasoning, it can end up hallucinating hard. Yesterday, I asked Claude thinking to solve me a problem in c++ and it showed the result in python.

OpenAI O3-Mini 1 year ago

I think that OpenAI should reduce the prices even further to be competitive with Qwen or Deepseek. There are a lot of vendors offering Deepseek R1 for $2-2.5 per 1 million tokens output.

Sounds like a good thing overall, my biggest annoyance when I was writing a flutter app was the codegen for annotations (which sure it's better iteratively, but the first one was taking minutes), but if you move these seconds that happen once in a while to seconds during "hot" reload, you're just losing. Honestly, I think they should try to come with a faster codegen, maybe write it in c++ or rust and fix these problems, because macros aren't a silver bullet. They introduce complexity, a new "thing to learn" and sometimes lead to Turing complete machines.

This is not exactly right, they said they spent $6M on training V3, there aren't numbers out there related to the training of R1, I can feel it will be cheaper than o1, but it's hard to tell how much cheaper. I can guess that overall deepseek spent way less than openai to release the model, because I have the feeling that the R&D part was cheaper too, but we don't have the numbers yet. Anyway, we can assume that deepseek and Alibaba will try to get the most out of their current GPUs however.

Who's this "we"? Is there anything that runs on the Bluesky protocol outside of the Bluesky itself which has its own extensions which can't be federated. Also, when I opened this site, all the posts were from a certain political ideology. The algorithm is probably more or less the same as Twitter in pushing contents loved by their creators.

So when I buy a Dell laptop now (won't happen), I need to ask for Dell Pro Premium package and make sure that the seller doesn't mistake it for Plus version or Base. Why would you go for "easier" names and go for 3 subcategories with the same name within 3 other categories. Just sell Dell/Dell Pro/Dell Pro Max (even the names are copied from Apple) with different specs, why give them subcategories

It truly protected web3 from the "normies" that could have learned about crypto from this video. AI moderation is such a joke, every reupload (or a completely different video on the same subject) can take a video down because they look "similar" enough for the AI and no person would bother checking it. I expected a different treatment of their bigger creators, but that's what it is.

A minecraft server isn't exactly a small side project. There are some in works for 3-5 years and they are not yet complete, some have very specific features (like https://github.com/MCHPR/MCHPRS which is meant for redstone showcases). This COBOL server doesn't yet implement lighting and that's one of the hardest parts since mob generation also depends on it. It also didn't fully implement some blocks. You need years to finish a minecraft server so getting something done fast isn't the best path along the way.

Personally, I use it for everything right now. It's faster to do `uv init` and then add your dependencies with `uv add` and than just `uv run <whatever>`. You can argue that poetry does the same, but `uv` also has a pipx alternative, which I find myself using more than the package manager that my distro offers, since I never had compatibility issues with packages.

I've built 3 iterations of captcha solvers for that crappy website based on https://github.com/drunohazarb/4chan-captcha-solver/issues/1 . The only thing I've learned along the way is that it's mostly pointless outside of a "learning" exercise, since they'll change the captcha (in terms of letter count or the entropy background). Initially, it was 4 characters with pretty obvious background, then it turned to 5, then it was both 4 and 5 and the current iteration which is also either 4 or 5, but with a lot of entropy surrounding the characters.

So instead of blaming the websites that do this and require you to disable the cookies for their 2000 tracking partners, let's blame the EU. The only thing I can blame EU in this manner is that they don't enforce that the websites should respect DNT or something similar. Also, there are plugins to remove the popups, it's not really that big of a deal.

I know chat rooms that have been nuked for Pornography etc. I reported some chats where I've seen inappropriate content and I received notifications that they were deleted. A lot of users are muted/banned too for illegal activities. It isn't exactly unmoderated, but the staff can't exactly search every single server under the sun for illegal material or activities. You probably don't know how bad Matrix is, out of 200k servers, 70k were banned for CSAM and there are still a lot of them around.

EU has been complaining about Telegram's end-to-end encryption for a long time and they want to implement some regulations to basically add backdoors into all messaging apps. I don't really see how this case will go on since at least private chats are encrypted so Telegram (theoretically at least) can't see the contents.

From Telegram sources: >Pavel Durov faces up to 20 years in prison in France. The trial will take place very soon – sources close to the investigation.

In addition to drug trafficking, he is accused of collaborating with an organized crime group, covering up for pedophiles, fraud and money laundering.

I don't know how reliable this is, but I've seen in 3-4 sources that he's arrested for terrorism, child abuse, drug trafficking (not providing data to prosecutors).