HN user

nickalaso

87 karma
Posts0
Comments44
View on HN
No posts found.

So, I have done a bit of research on this with the writeup+skills+scripts on my personal git: https://github.com/NickalasLight/codex-reasoning-bug-512-tok...

I think it is very interesting that: A. Removing the section in question seems to greatly fix performance on the benchmark candy question. B. Removing the section does not appear to change at all mean reasoning token use or the 512 reasoning token hit problem

It truly is impressive to me how good Nvidia has gotten at understanding 'Bubblenomics'.

Why don't you buy my overpriced hardware, then let me use the overpriced hardware you just paid for in exchange for extremely overvalued shares that border on funny money, so I can then take the same overpriced hardware you just bought and sell it at an extremely overpriced rate to some inference provider.

When the inevitable happens I really do think it going to be pretty bad this time.

Fable 5 is Back 20 days ago

Yeah, I tried using it before the 'safety' blocks got too much for me. Its pretty good as an 'orchestrator' calling subagents and being the PR reviewer/tast tracker. Seems really not that much better for coding though, yeah it is a bit better, but not worth it.

But when it did work it was making steady progress on a vague git issue backlog and actually following instructions to carefully break into atomic sub issues and carefully assign.

Fable 5 is Back 20 days ago

OpenRouter would like a word. Also my in progress SaaS would like a word. Models are incredibly interchangable, just find best performance per cost and optimize harness and prompts, easy to have multiple configs for each model.

Fable 5 is Back 20 days ago

Just wait patiently for the chinese labs to inevitably distill Fable without the 'policy layer/neurons' and offer it for pennies.

That one will be fun to see.

Fable 5 is Back 20 days ago

Yeah, and I also don't buy the semi-conspiratorial beliefs of some of my friends who think there is super intelligence somewhere being built.

If fable is the best we got everybody still has some more time till the apocalypse I think.

Fable 5 is Back 20 days ago

Yes, I distinctly remember the recent past when the roles were reversed, and Anthropic was the kind good hearted startup trying to compete with big evil OpenAI corporation.

Wonder how many times we get to see it flip back and forth before this bubble explodes.

Fable 5 is Back 20 days ago

It is surprisingly half baked for something the company seems to be proudly charging my first born child and both of my kidneys for.

Especially considering its all for the privilege of getting half way through a standard dev task before it just bricks out of nowhere on the 'safety' guardrails because of a keyword that often appears to be created by the model itself?

I've been running Fable on a few different projects and just by itself on same codebase it'll brick, so looks like it can just straight up create its own demise.

I do not think any of us yet have to worry about being 'replaced with ai' if the 'premiere ai company' still can't put out their flagship product with minimal qa. This is embarassing.

And the short 7 day window for subscribers is just taking the piss. So happy I'm paying hundreds per month to be free qa for anthropic, god knows they need the help.

This feels like an article written at least 2+ years ago when 'AI' was still a bit more 'mysterious'.

Mona barely thought about profit because you didn't update the many configurable levers to increase the code loops generation related to profit.

You can update the sys prompt, perform basic harness changes, such as sys prompt optimization, toolset optimization, tool description optimization, 'multi-agent' and 'composed-agent' flows, etc. etc. Get it to do whatever you want.

Its like if someone wrote an article about how their windows 95 pc they named paul didn't think about finances too much because they never installed quickbooks on it.

I'm not really sure who this article is meant to appeal to.

Yeah, but that has literally been their business model from the start yes?

Even now, after all the development, if you can somehow get Claude Fable to run without bricking due to the safety 'features', whatever it imitates is still a poor replacement compared to the slew of available open source code repositories that it stole its training from.

Problem is humans are lazy and its really easy to just yell at a chatbot until you get some slop that mostly does it.

So I went ahead and quickly vibecoded a working harness with a barebones tool interface and some constraints on output (credit to noperator for the idea). github: https://github.com/NickalasLight/VibeHarness.git

Its meant for a Windows machine using ollama but I'm sure anyone who wants to mess with it can point claude code at it to convert it for your own operating system and requirements. After install you can ask it to do something with "vibe 'create me a poem about cheese in cheese.txt'" its workspace is by default the directory the cli was located in when you called it.

Sublight speed Von Neuman probes (self replicating) are the kind of tech we are likely going to be able to produce ourselves sometime in the future.

Which is tech that would allow complete saturation/exploration of the entire galaxy on the order of millions of years. As essentially by the time one probe is able to reach one side of the galaxy from the host system the entire galaxy will have already been filled with probe copies.

It wouldn't be particularly difficult for a more advanced civ with potentially millions or billions of years evolutionary lead time on ourselves to have explored the galaxy and by extension have probes in our solar system.

These claims in the hearing are definitely extraordinary, but the tech required to "make it here" doesn't particularly need to be.

For me, I am also having an issue where the text is too large for the cards, causing large portions of text to not fit the screen and be unreadable. Interesting design concept, but appears to be buggy/not-tested.

The social isolation of this whole period has definitely had an effect on me personally.

I'm assuming most people already have a strong social circle that they've been able to rely on during this, but as someone who moved to a new city to work as a young professional in his late 20's and no friends in the area, it's been rough.

I am in the same camp, though, drinking doesn't exactly expand the options that much, just so you know. The real problem is our generation doesn't really seem to have many places to actually meet and make friends with others our own age. It is either some club (not exactly conducive to long term friendship building) or online in some way. If you don't make lasting friendships during college, many of us 20-somethings are kind of screwed unless we bust our asses trying to join communities outside of work. I've been thinking about volunteer work for instance.

Any comment like this is immediately hit with a flood of "well I love working from home". I'm under 30 and absolutely miserable working from home. The lack of any social interaction with my coworkers combined with a personal lack of a social network outside of work has left me lonely and isolated. On a positive note the lockdown has pushed me to begin exercising, start eating healthier, stop drinking, and start learning a second language in prep for a long sabbatical abroad.

I think the real way to make politics less crazy is to create barriers to voting, actually.

The real problem with our politics is that the more informed members of the populace make up their minds early on, leaving the lowest common denominator voters as the only votes left to be swayed. Which causes our political cycle to be driven by abject stupidity rather than useful information and policy discussion.

Imagine how much better things could become if we could find some fair, reasonable way to create a truly representative voting base that were required to be actually quantitatively informed on policy and policy consequences.

Various polls throughout the pandemic have shown that both zoomers and millenials have been terrible on average estimating their personal risk as well as the risk of the other age groups. Believing they have at least a 2% fatality risk, and a 7.5% chance of hospitalization. Which is wrong by several orders of magnitude. Source: https://www.economist.com/graphic-detail/2020/07/21/young-pe...

Social media has magnified the natural fear and anxiety that many of us naturally feel surrounding disease. The generations most exposed to it are the least able to accurately estimate risk associated with Covid as a result.

I personally believe that historians will look back on this period as one driven by panic and extreme emotion. Social media is to blame for this, and will continue to magnify these kinds of events.

If by "best" we mean, most educated, most capable of solving engineering problems, "highest IQ" etc. (Which is how we generally define best in this context) Then yes.

There used to be a big incentive for the smart scientific and engineering minded types to go into academia, but now the majority go into either financial services or tech work because that is where the high salaries and actual problem solving are.

As a conservative, its not that educated conservatives think society is "fair", in that their aren't people with more or less advantages. Rather, we accept that life is "fair enough", and true equality, in the sense that everyone has the exact same capital/life outcomes is worse for society than it is good.

To be more fair, is better, in the sense that rewarding meritocracy is the goal, but even here there is a fundamental unavoidable unfairness born out of the natural talents and genetic predispositions of individuals making them uniquely advantaged for a given sport, field, career, etc.

I believe it has something to do with someones level of "institutionalization", or perhaps, how comfortable they are with institutions and authorities themselves.

A PhD holder is a PhD holder, at least in part, because they trusted that the many 1000's of hours of study and hard work they put into to receive a certificate from an institution was worth it. And this is because they, at some level, trust the institution itself.

Presumably this trust can also transfer to other "authorities", like US cable new media companies.

The top level comments are arguing that the cause for this is due to the effectiveness of various health measures. The problem with this argument is we would expect for flu cases to continue to be reported in countries and US states where lockdown measures were never properly enacted.

This doesn't appear to be happening. Instead, "the drop-off in flu numbers was both swift and universal." (article quote)

This universality suggests a data problem to me rather than an environmental change.

Could you explain why this is "literally so stupid" when this appears to be what is happening?

If lockdown measures, social distancing, mask wearing, etc. were the true cause of the massive reduction in flu cases, we would expect to also see that countries and US states that failed to enact measures would be reporting high flu numbers to the WHO.

This doesn't appear to be happening. Instead, "the drop-off in flu numbers was both swift and universal." (article quote)

This universality suggests a data problem to me rather than an environmental change.

They don't need to claim ownership of the drama to write the paper, in fact, my first thought was that they would specifically try to avoid taking ownership and instead write a paper "discovering" the vulnerability(ies).