For capabilities reference:
I made a lower effort but similar scaffold for LLMs to do iterative drawing in Nov 2024, with Sonnet 3.5 as the artist: https://paritybits.me/llm-drawing-with-eyes-open/
Quite a difference.
HN user
https://letterspractice.com - https://patched.network - https://github.com/patched-network/vue-skuilder
Working on FOSS and user-friendly alternatives to things like khanacademy, anki, MathAcademy, Alpha School, etc.
Modern, open edtech tooling.
Also http://paritybits.me
For capabilities reference:
I made a lower effort but similar scaffold for LLMs to do iterative drawing in Nov 2024, with Sonnet 3.5 as the artist: https://paritybits.me/llm-drawing-with-eyes-open/
Quite a difference.
You can have these things without owning a house.
Yes this is feasible, but respecting it as a design problem, renting is more transient than ownership and tilts the floor away from deep communal relationships.
Comparing the neighborhood I grew up in with the one I now live in is night and day. My mom has had the same next door neighbors for 44 years. Up and down the street there are many similarly familiar persons.
By comparison, from my own front door, I can only physically see two houses that are owner occupied. There are good neighbors (and friends!) in the rental houses as well, but investing in those relationships pays off with lower certainty because circumstance is very likely to uproot them at any given moment.
Thanks much. Source code is for customers only I guess :)
I had the same nit, but I imagine deforming text / inline content generally would be a much larger effort.
I like this a lot and am going to experiment w/ incorporating in my early literacy app.
Heads up: the "See it live in the showcase → " links in the API documentation do not go back to the showcase - they just reload the same current API section.
Question: maybe I've missed it, but what exists here wrt distribution / packaging / bundling / source availability? I see MIT listed, but no repo. I see src="https://jelly-ui.com/package.js" as a sourcing import, but obviously I'm not going to bundle foreign assets into my app.
Yes - persons with death wishes having arbitrarily powerful consultation is the crux of it.
Apologies for the bad example. Replace w/ gain of function / whatever else, or just brainstorm with your local model, ect.
The logic, whose premises you can take or leave:
Even at the level of, say, Opus 4.5+, open weight models give a quick turnaround to every Joe and Jane on earth having easy access to pretty high quality improvised weapons design, cyber / auto-fraud capabilities, etc.
All the existing models (closed and open) put up decent resistance to participating in activities like this, and especially behind API walls with content monitoring and account bans.
But the published open-weight models can be fine tuned or abliterated into arbitrarily sharp-edged tools. EG, if it's physically feasible to build a nuke in your garage, it may soon be the case that more or less anyone will have competent guidance to do so.
A const prompt across all of Anthropic's subscribers could draw from a global cache rather than per-user?
Although saying that out loud makes me question it - each per-user chat and growing cache would need eventually to own its own ~contiguous memory block.
What access?
Push notifications.
Integration with OS features is what made the app ecosystem, because of utility.
This is true of some apps, like the beer-drinking one that uses the accelerometer / other orientation sensors.
It's not true of a large number of other apps, hence the "your app could have been a webpage" charge. This is distinct from "every app could be a webpage".
Great respect to site guidelines, and to you. Object on both counts.
1. The post was obviously bullish / optimistic on the technical capabilities. Not in the least dismissive.
2. The economics extrapolation is obvious. See current precedent for paid access for purchased screen-casts of dev work: https://pdoom.org/open_calls/04_crowd_cast.html
Any minute: wear it permanently to sell training data on LLMs. Take an audited IQ test to negotiate your rate.
Better than text-stripping the internet - this thing will soon be pulling the logits as well.
"The government has the right and responsibility to shut down sources of misinformation in the news and online."
Funny that I read this as AuthLeft coded (specific to Youtube suppression of Covid truthing). But obviously the alignment is just a function of whatever specific information is labelled "mis".
But in general: agreed, and this is a good list.
Is typst a good tool for something like a flyer (eg, printable respecting fold lines) or more generically one-page posters?
I see PDF as a blessed output, but it seems mostly in context of longer form typesetting-heavy workflows (books, papers), rather than design-heavy.
Very cool project. Thanks for sharing!
Very interested in this! Can you share more about the modelling method (eg, three js?), the task list, and outputs here?
I think there's probably some good juice to squeeze in terms of spacial awareness by doing a benchmark something like
- give 3d modelling task
- render and snapshot from a variety of angles
- feed to third-party vision model for a "what is this" type query
- grade on end-to-end accuracy
Bonus points for asking the vision model something like "how beautiful is this 1-10".
Long ago a friend of a friend described a job interview at an ice cream / chocolate shop in a local mall.
The interviewer asked something like "who is our competition here?", and the friend of friend listed off other places in the mall to get ice cream, candy, deserts, etc.
Wrong answer. The ice cream and chocolate store was in competition with every other store in the mall. Time or money spent at the GAP can't be time or money spent here.
---
Whether or not people are using LLMs for news specifically, any new, large eater of eyeball-time is going to hurt the business landscape for all other eyeball harvesters.
If package X is of sufficient public interest (user count, nature/sensitivity of user data, downstream distribution, etc), then the public interest + cryptographic credentials should permit access to best-available security auditing.
Your private fork doesn't meet the conditions described.
This is a credentials and access list oAuth style problem, and not really intractable.
For package X, I should be able to present my npm (homebrew, apt, nuget, etc) credentials with publishing rights for the package.
If package X is of sufficient public interest (user count, nature/sensitivity of user data, downstream distribution, etc), then the public interest + cryptographic credentials should permit access to best-available security auditing.
Yes, we still are trusting trust, that the owner of the package itself is not malicious, but that's not a sharp degradation from status quo.
I think that as simple as is doing a lot of work when the problem domain is all natural language (or more - all strings?) rather than some well specified DSA problem.
On-demand ciechanow.ski caliber articles are a pretty good AGI indicator. All the work on that site is wonderful.
I'm working on a framework for general purpose interactive tutoring systems. An SRS background process over a pluggable system of pedagogy protocols over a given curriculum. This is at https://github.com/patched-network/vue-skuilder, or https://patched.network/skuilder
With this framework, I'm making (among other things) an early literacy app at https://letterspractice.com. My aim here is to hit >= 75% efficacy of Mentava at <= 1% of the price.
The app is near to production readiness, and I'd be happy to share access now with anyone who has verbal but non-literate kids. Be in touch if interested at colin at letterspractice.com
People may perceive you to be cheaply mischaracterizing the argument.
Nobody believed or suggested that GPT2 could do longform or produce novel text that stood up against careful scrutiny as insightful or well informed. But because the capabilities were novel, people would have no strong alternative than to believe some person wrote it.
You current tripping over LLMisms is irrelevant. You have years of antibodies, both personal and herd-immunity (eg, the many, many articles and comments that describe LLMisms).
At that time, nobody believed a dead internet was technically feasible. Maybe this is hard to remember now.
The "danger" was in terms of spam / misinformation proliferation, not the same category of capabilities adjacent risks current discussed.
You can hold your own opinions on spam/misinformation as a problem, but to say there was no credibly anticipated outsized downside to a sudden jump in human-passing text generation feels pretty off to me.
... so the mechanic produced an invoice, itemized.
changing the CSS - $0.05
knowing which CSS to change - $30
Sure, live in shame, but don't let go of the humor in it all :)
Automation doesn't make operators more careful. It makes them forget how to be. The more reliable the system, the less ready the human.
The entire premise of a system is that it removes the need for careful attention.
system: signal lights tell me whether or not I can pass through an intersection, so that I do not have to attend to potentially high speed traffic from a variety of directions.
system: the side my knife blade sits on my arched guide fingers, so that I do not have to attend to the edge of the blade or the location of my fingers.
etc etc.
replication: https://en.wikipedia.org/wiki/Quine_(computing)
autonomous replication: https://en.wikipedia.org/wiki/Computer_worm
nb that writing your own quine remains in general terms a fun and challenging exercise in many programming languages, but not python.
I'm no decision theorist but I think they should wait for the rewards outweigh the expected harms in expectation rather than being statistically equal.
Once or twice I've experienced extreme pain, and it was downstream of a bright light shining on a wet rock for millions of years.
I try to imagine myself long ago, on the outside looking in, with someone explaining to me that extreme pain, wondrous art, hunger, triumph, and despair would all unfold in due time where the rocks were wet and the lights bright enough.
I can imagine myself calling this clear nonsense.
You're assuming that because Claude produces text that appears to express these qualities, Claude must have them.
Not to be confrontational, but the OP assumed no such thing. OP asserted that it's important for Claude to have the qualities - not that it's important for Claude to present as-if it had them.