They embedded a script that checks the victim’s host operating system and silently executes a remote payload.
Seems like this is becoming a recurring theme, similar story was on the front page last month.
HN user
spinning the web?
They embedded a script that checks the victim’s host operating system and silently executes a remote payload.
Seems like this is becoming a recurring theme, similar story was on the front page last month.
Firstly, this is the best TLDR I've read in a while and it convinced me to read the whole thing.
TLDR: I gain a lot of fulfillment by making things. I don't consider things built by others at my request to be made by me, and are therefore much less fulfilling. And then I feel sad. This article starts strong and then heads off into the weeds.
Secondly, maybe this conversations boils into: what level of abstraction are we comfortable working at?
I mean that to the tune of "To bake an apple pie, you must first create the universe".
Few people will create the universe (please introduce yourself, if you are one); some will buy the apples and the crust and put them together; some will make the crust from scratch; some will grow the apples from seed; some will follow a recipe; some will buy a frozen pie; some will simply buy the pie.
I think we can all agree that buying the pie is not baking it. But what about the others? To me, there's an argument that all of them are "baking the pie", just taking place at different levels of abstraction. And I think you can take pride in any of them, and hopefully more pride in whether the pie is delicious and well-shared.
As we become more disconnected from the work we do, let writings and art be created by AI, we forgo the meaning, emotions, and learning behind such work.
Excellent post, and thanks for the introduction to Wendell Berry's essay.
Link for others interested: https://classes.matthewjbrown.net/teaching-files/philtech/be...
Thank you.
Looking forward to it. I'm also a big fan of https://vincentwoo.com/2023/07/20/tunnel-vision-an-unauthori..., thanks for keeping the internet interesting.
From OP's prior work scanning Sutro Tower in SF: https://vincentwoo.com/2025/02/18/sutro-tower-in-3d/.
This scan is made possible by recent advances in Gaussian Splatting. This is an emerging technology that lets us quickly create very detailed models just from photographs. For this model (or splat, as we call them), my friend Daylen and I flew our drones around Sutro Tower at a respectful distance for an afternoon until we had collected a few thousand photographs.
I then aligned these pictures in free software called RealityCapture. Alignment is the process that teaches the computer that a bunch of points in different images all actually correspond with the same point in real life. Then I used another piece of free software called gsplat to produce the 3D model itself.
I'm assuming this was done similarly. Very cool.
Is there more information available? This article is effectively one sentence long:
On Sept. 30, 2025, a Flock vice president logged into the Dunwoody PD’s Flock system and looked at a camera in the children’s gymnastics room of a private community center here in Dunwoody.
And there is a lot being implied in this sentence.
It’s striking the extent to which Claude Code and Codex are proving to be quite sticky; whichever harness you start working with is likely to be the one you stick with, and that figures to be even more the case with non-technical users.
My experience has been quite the opposite. I was using Claude Code almost exclusively this winter/spring and swapped to Codex earlier this summer. It took no time whatsoever to switch. And before Claude Code, I was using Cursor. Same story.
[edit: Oh and there was also a brief interlude with Conductor, though I think they're more or less just serving the underlying Claude/Codex harness]
Not convinced that this slop measurement is useful.
I asked Codex to generate an article with a high score and then asked Codex to (reverse?) hill climb that score. The original generation scored 97% and then the optimized one scored 1%. Both are pretty bad and read like slop.
https://gist.github.com/wbew/8a2bd6686bf875210f2244ac8ea65bf...
Kimi Work is a Local Agent designed for deep workflows. It mounts your local folders, navigates the web autonomously via WebBridge, runs Python code in the background, and executes scheduled tasks.
It's clearly a dupe of Claude/Codex products (Codex especially, styling-wise), but I think Kimi's goal here is simply to appear on feature-parity with bigger labs. I doubt they'd want to invest much into designing a UI for "the future of agentic work" from the ground up.
Their bread and butter and claim-to-fame will continue to be low-cost, near-frontier models, for which they need to at least appear to have some app layer to plug into.
And, in some ways, duping is a great way to show people that even the app layer can be commoditized.
Agreed. And, it's really not that hard to organize a simple event.
I used to talk myself out of it all the time, but have recently just been going for it. It's been great.
Recurse is an incredible community. Thanks for everything you’ve done for the NYC and broader tech ecosystem!
Open source Fable/Sol challenger! Interesting to do a release product-first.
Inkling is not the strongest overall model available today, open or closed. Instead, a combination of qualities makes it a good open-weights base for customization: multimodal capabilities, efficient thinking, and availability on Tinker for fine-tuning.
Open base models that can be fine tuned on Tinker is a great business model IMO. You (i.e. an enterprise) can own your own model & have it perform frontier-or-better at your task at potentially much lower cost and Thinking Machines gets to be your essential infra/service provider in this world.
Also,
Inkling-Small matches or exceeds its larger sibling on many benchmarks — the result of improvements we made to the pre-training data and recipe for the smaller model.
Very cool! Excited to see the next generations of Thinky models.
The vulnerability was first identified by Mindgard on December 15, 2025. We reported it the same day and multiple times since. More than six months and 197+ new versions later, the issue remains present in the latest tested version of Cursor.
The report was initially closed as Informative and out of scope. After we challenged that determination, HackerOne reopened the report, reproduced the issue, and confirmed that the details had been delivered to Cursor. And then everything stopped. Requests for updates went unanswered, additional follow-ups received no response, escalation through HackerOne produced no meaningful engagement, and direct outreach to Cursor leadership yielded the same result: no response.
Really unfortunate. I don't understand why there's such a lack of response on the Cursor side.
Ah yea, I shouldn't have linked to Threads (it was just the first result I found to see the full photo). I replaced the link with a better one.
Here's a link to the actual photo. https://petapixel.com/assets/uploads/2026/06/09-Jupiter.jpg
[replaced a Threads link for a better one]
I’m an intergalactic space warrior and leader of the Recyclons from planet Sigma IX.
Ok you have my vote.
+1, I would love to stop reading AI slop.
What I don’t like is two things. One, this constant bullshit about some window closing, or the perpetual underclass, or falling hopelessly behind.
And two, this strawman jump from, oh hey, it’s a fancy autocomplete, smart compiler, better search engine, to it’s gonna like own the whole light cone bro like if you aren’t in SF and at the right parties there’s gonna be like a flash of light in the sky one day and you’re not even gonna know what happened but everything just Changed.
Haha, OP has a way with words.
In a way, both these emotional extremes (FOMO & the singularity) are just tools being used to continue driving the massive CapEx behind LLM improvement. Hate to love it? Love to hate it?
So good, thanks for compiling this list.
For this study, the Google Maps algorithm was modified to prefer alternative routes with similar travel times and segment types, effectively guiding trips away from the pre-selected congested segments
Over a six month period, we adopted a city-wide switchback (also known as crossover) experimental design, alternating between this treatment and the control (unaltered) routing algorithm over consecutive days to appropriately measure the effect of this intervention
Averaged across cities, we observe a median increase of around 2% in driving speeds on targeted segments, corresponding to a median decrease of 0.5% to 1.0% in fuel consumption rates
The cities were: Atlanta, Boston, Chicago, Los Angeles, Miami, New York, Philadelphia, Salt Lake City, San Francisco, Seattle.
The data and code is also available (https://github.com/google-research/google-research/blob/mast...) from the paper (https://www.nature.com/articles/s44284-026-00443-x).
Kudos Google! Nice to see this kind of work. That said, let's just build more trains?
I'm really enjoying Codex. I use the app, and I find the features are excellent (the in-app browser and annotations especially).
This is almost certainly a reaction to GPT 5.6, which IMO is better positioned than Fable from a cost perspective. It's crazy to see how the tides change between labs every few months.
What is the point of the AI enhancements toggle?
The key idea here is that your codebase is context that will be used for future changes. And context determines the model’s output, so it’s still worth having a well-designed codebase.
Easier said than done to be honest, especially if there are many people (and their agents) pushing code. It’s hard to keep up these days.
Thanks for including a section on Token Efficiency (https://x.ai/news/grok-4-5#faster-than-flash-models), hope to see this more prominently in all model releases.
Some stand out takeaways:
We assessed how reliable current measures are for trying to find microplastics in blood. And what we found is that lipids and fats will give you a false positive for polyethylene.
We worked with an architect, and we built the lab pretty much from scratch. [...] So we ended up going with stainless steel. It was the only way to not have any plastics.
I don’t think we’ve got really good evidence at all for what effects [microplastics particles on their own] might be having on human bodies. If we’re eating plastics, what size and what type of plastic can actually get into the bloodstream?
It’s designed to be extremely easy to self-host on your own infrastructure.
Kudos for this. Per the docs: https://docs.chatto.run/,
Chatto ships in a compact, self-contained binary
it uses NATS, a compact message broker that also ships with a built-in stream persistence engine [...] NATS is just as easy to provision as Chatto, and most of our examples will show you how.
you can also configure an external S3-compatible object storage for Chatto to store your files in, and we strongly recommend doing so...
The actual calls are powered by LiveKit (Apache-2.0), which you need to deploy alongside Chatto. As with NATS, the deployment examples show the required wiring.
...
And kudos for backing it up with real guidance. Great project.
Back in my undergrad, I took a Functional Programming class taught in Elm. It was primarily about functional data structures, but we also got to build a web app using Elm towards the end.
At the time, I didn't think much of it -- I was probably busy learning React and JavaScript and yada yada for employment purposes.
Now, having spent some time in industry and having used some gargantuan web frameworks, I find myself missing Elm. MVC in Elm is wonderfully straight-forward and easy to reason about.
Congrats on the road to 1.0! Glad to see Elm still active all these years later.
The title is misleading. This isn't an AI tutor so much as a practice quiz platform with an AI autograder.
constructed-response questions (CRQ) are graded by Claude Sonnet 4.6 against instructor-defined, question-specific rubric criteria
Crucially, LLMs make it feasible to grade formative CRQ against rubric criteria at scale, a capability that appears pedagogically significant rather than merely convenient.
They specifically call out that the "RAG chat assistant" part of Phosphor (the platform) wasn't used much.
I commend the effort here, but I don't think these results are particularly noteworthy. The conclusion is essentially that people who do practice quizzes will do better on exams.