Fair enough. I really like the tarpit analogy, wasn't familiar with it. You can keep pulling your feet out faster than the tar rises, as long as you're willing to keep spending the energy, possibly with diminishing returns over time.
HN user
gmaster1440
https://www.markfayngersh.com/
I think we're basically agreeing here. Your point (if I'm reading it right) is that taste and discernment do scale, but the gains come through pretraining/parameter scaling, which is slow and expensive compared to the fast, cheap wins in math/coding from smaller models. So taste is more of a lagging indicator of scale. it improves, but it's the last thing people notice because the benchmarkable stuff races ahead. Which also means taste isn't really a moat, just late to get commoditized.
If you're properly bitter-lesson-pilled then why wouldn't better models continue to develop and improve taste and discernment when it comes to design, development, and just better thinking overall?
No Pinball :(
What if the slowdown isn't a bug but a feature? What if AI tools are forcing developers to think more carefully about their code, making them slower but potentially producing better results? AFAIK the study measured speed, not quality, maintainability, or correctness.
The developers might feel more productive because they're engaging with their code at a higher level of abstraction, even if it takes longer. This would be consistent with why they maintained positive perceptions despite the slowdown.
This is also an uncomfortable direction. All investors have been betting on the application layer. In the next stage of AI evolution, the application layer is likely to be the first to be automated and disrupted.
Highly agree with the sentiments expressed in this post, I wrote about something similar in my blog post on "Artificial General Software": https://www.markfayngersh.com/posts/artificial-general-softw...
the entire premise of this economic index is that they're showing you actual usage and insights from millions of anonymized claude conversations.
i think it says, amongst other things, that there is a salient difference between competitive programming like codeforce and real-world programming. u can train a model to hillclimb elo ratings on codeforce, but that won't necessarily directly translate to working on a prod javascript codebase.
anthropic figured out something about real world coding that openai is still trying to catch up to, o3-mini-high notwithstanding.
The "second year university student" analogy is interesting, but might not fully capture what's unique about LLMs in strategic analysis. Unlike students, LLMs can simultaneously process and synthesize insights from thousands of historical conflicts, military doctrines, and real-time data points without human cognitive limitations or biases.
The paper actually makes a stronger case for using LLMs to enhance rather than replace human strategists - imagine a military commander with instant access to an aide that has deeply analyzed every military campaign in history and can spot relevant patterns. The question isn't about putting LLMs "in charge," but whether we're fully leveraging their unique capabilities for strategic innovation while maintaining human oversight.
The paper argues against using LLMs for military strategy, claiming "no textbook contains the right answers" and strategy can't be learned from text alone (the "Virtual Clausewitz" Problem). But this seems to underestimate LLMs' demonstrated ability to reason through novel situations. Rather than just pattern-matching historical examples, modern LLMs can synthesize insights across domains, identify non-obvious patterns, and generate novel strategic approaches. The real question isn't whether perfect answers exist in training data, but whether LLMs can engage in effective strategic reasoning—which increasingly appears to be the case, especially with reasoning models like o1.
How generalizable are these findings given the rapid pace of AI advancement? The paper studies a snapshot in time with current AI capabilities, but the relationship between human expertise and AI could look very different with more advanced models. I would love to have seen the paper:
- Examine how the human-AI relationship evolved as the AI system improved during the study period
- Theorize more explicitly about which aspects of human judgment might be more vs less persistent
- Consider how their findings might change with more capable AI systems
AI appears to have automated aspects of the job scientists found most intellectually satisfying.
- Reduced creativity and ideation work (dropping from 39% to 16% of time)
- Increased focus on evaluating AI suggestions (rising to 40% of time)
- Feelings of skill underutilization
Does Codebuff / the tree sitter implementation support Svelte?
I imagine it's difficult to be a good teacher and find effective ways to encourage students to rigorously think about things they care about in spite of the discomfort it might cause.
I also believe increasingly capable and sophisticated AI systems will play a formative role in transforming education, not as the current chatbots that are disrupting education as mentioned in the article, but as active participants in the reimagined classrooms of the future. The transition will probably be rough, but it has the potential to bring about a better future and more fruitful learning and writing.
With more capable models, reliable test time compute, and more sophisticated RAG (https://openai.com/index/openai-acquires-rockset/) I genuinely struggle to see meaningful use cases for traditional user interfaces.
The biggest hurdle I'm personally struggling with now as a software engineer trying to find the motivation to continue writing code and developing products isn't simply that AI is getting better and better at writing more complex code and doing more of my job for me (though it certainly doesn't help), but more importantly that it will soon do away with entire classes and categories of desktop, web, and mobile applications as the human interface evolves towards conversational, intent-driven interactions with AI.
Vast majority of all apps are just tables, forms, and JSON over the wire—I don't see that continuing to be the case for much longer.
I've enjoyed working with Reflex (https://reflex.dev) as a pure Python wrapper over React.
Relevant to this discussion is the fact that if an LLM can’t come up with that, it wouldn’t be due to the inability to mix and match to form novel ideas, but something else, and that something else hasn’t been clearly articulated yet.
I'm happy to leave the conversation here. I don't necessarily disagree with what you're saying, but we appear to be making different points, or at least at different levels of description, and it's not really productive anymore.
AlphaFold is really impressive and made scientific advancements and discoveries in the field of protein folding, and is now even expanding into more molecules and biology, but it was explicitly trained to do just that. You're not going to see AlphaFold write compelling science fiction.
We can build models that are specifically trained and fine tuned on scientific fields to make advancements in them, but that's different from what I'm talking about, which is building a model that forms its own hypothesis, designs its own experiments, and contributes to the wide and deep wealth of knowledge that, crucially, goes well beyond the scope of its training data.
I would say the opposite: AI are very good at searching through information spaces, much better than we are.
We are likely talking past each other here. By "searching" I don't mean how inference is currently carried out by efficiently analyzing the context window using weights trained on large data sets fine tuned on specific goals.
I mean the process by which novel information is discovered, which is why many proponents of AI will concede that it's not currently capable of "doing science" or making novel discoveries.
we don't know what that is, it's just what we do.
Not sure I understand, we have a pretty good understanding of what qualia actually is, even if it can be difficult or awkward to talk about conceptually. The gap between having a subjective experience and not having one is a large one, just ask anyone who's alive but under general anaesthesia that induces loss of consciousness. Qualia is simply what arises from the quality and character of having a subjective experience.
Human beings are currently capable of productively searching through the space of possible knowledge and experience in ways no current AI systems are. This is not to say AI will never do this, but I think it's fair to say there are things human beings are capable of doing today that AI is not, and it remains very much unclear whether AI will ever be able to achieve important milestones like being conscious in the sense of having a subjective experience and therefore forming special knowledge that can only come from that, like Qualia.
What are the use cases for going with this framework over something like ASP.NET Core Minimal APIs? https://learn.microsoft.com/en-us/aspnet/core/fundamentals/m...
It's refreshing to see a science fiction writer underplay the capabilities of AI, but if anyone can speak to the nuances and implications of generative AI on art and writing it's probably someone like Ted Chiang.
We can debate his generalized definition of art as making creative choices that carry subjective, intentional, and performative value for human beings (and therefore LLMs fall short of this), but I think he makes a couple strong points nonetheless:
1. The argument others like François Chollet have also made, that we have yet to see any AI systems capable of exhibiting intelligence beyond stylistic mimicry or forming generalized knowledge about concepts from large data sets.
2. The subjective experience of human interaction is valuable and desirable, and will remain so in the face of increasingly capable models, not because they won't be able to compete in producing inspiring art or enjoyable fiction, but because of the inherent primacy of human intentionality and experience.
How about Svelte 5 with runes mode?
One of the biggest challenges I've encountered so far while working on my SvelteKit with Svelte 5 codebase is that all frontier models struggle to understand differences between major versions of languages and frameworks, leading to a lot of annoying hallucinations. They're fantastic at writing code, but the last mile of figuring out how to fix issues for incorrect APIs or syntax becomes very tedious.
Scale beats all else. The best performance improvements come from increasing scale, rather than incremental insights in novel architectures.
...until the next novel architecture is discovered, which won't happen without said AI research.
If I'm in charge of forging a presidential election, how difficult is it for me to use realistic, "ugly" numbers to sell it more effectively?
Apple's New Foundation Language Models (AFMs)
1. Two Main Models: - AFM-on-device: ~3 billion parameters, for efficient on-device use - AFM-server: Larger model for Private Cloud Compute
2. Architecture and Training: - Based on Transformer with optimizations - Three-stage training: core, continued, and context-lengthening - LoRA adapters for task-specific fine-tuning - Innovative quantization: 3.5-3.7 bits per weight
3. Performance and Benchmarks: - AFM-on-device outperforms larger models (e.g., Gemma-7B, Mistral-7B) - AFM-server competitive with GPT-3.5 - HELM MMLU (5-shot): AFM-on-device 61.4%, AFM-server 75.4% - GSM8K (8-shot CoT): AFM-server 83.3% - Strong in instruction-following (IFEval) - Best overall on Berkeley Function Calling Leaderboard
4. Capabilities: - Excels in instruction following, tool use, writing, math - Long context support up to 32k tokens - Specialized for tasks like summarization
5. Responsible AI: - Focus on user privacy and responsible AI principles - Extensive safety measures (red teaming, human evaluations) - Lower violation rates on safety prompts vs. other models
6. Unique Aspects: - "Accuracy-recovery adapters" post-quantization - Novel RLHF framework: "Iterative Teaching Committee" (iTeC) - New RL algorithm: MDLOO
Semantic Clarity: Converts web content to a format more easily "understandable" for LLMs, enhancing their processing and reasoning capabilities.
Are there any data or benchmarks available that show what kind of text content LLMs understand best? Is it generally understood at this point that they "understand" markdown better than html?