How long was it trained for? How many tokens?
HN user
ericb
Do the two cards "share" their memory pool? Can work still be split across it? I'm wondering how it would do with something like fine tuning?
Do you get the speed of the 5080 with the memory of the 3090?
This is pretty cool! I would love to see an even larger models shrunk down.
If you got that into a couple gigs--what could you stuff into 20 gigs?
I'm not either. If this was GPT-voice, I'd be happy. It's concise, technical, with good emphasis but no drama or AI tropes.
A prompt injection solution that seems to benchmark better than any other approach out there, while not using hard-coded filters or a lightweight LLM which adds latency.
biased to hiring a slightly worse applicant
I understand your reasoning, but in practicality, I don't think this is true. This would be true if companies though with a coherent set of incentives. Instead, individual incentives are at-play here.
If a company is paying for a recruiter, it usually means:
- It isn't highly cash constrained - Values the time of its IC's, managers and HR more than the fee - Valuation for the role is not cost-based, but value-based - Only at the penny pinching startup stage is the recruiter fee a real factor in a multi-year investment that should be yielding a high return. Beyond that, the bias evaporates and the real incentives lie with individual incentives, and available budgets.
I mean, sure, but all those things I named don't seem to be scale induced? They seem to all stem from clueless regulation, which is as simple as not not signing silly laws? I'm missing where scale plays into the items I mentioned.
I see Massachusetts as sort of the non-insane liberal counterpoint to California.
Things work here and nobody seems to be passing the "oops my unintended side effects and clueless regulations messed things up horribly." Or, if they do, it is at something like 1/10th the level.
We didn't start warning label spam everywhere. We don't have weird propositions that are causing run-away housing prices. There aren't bar codes on our 3d printers, or cookie banner requirements on every website. Well, ok we do, but that nonsense all came in from other places.
We did pass laws to lower PFAS/PFOAS. That seems reasonable. Government can work.
Materialism is the outlier here, not the default, and it has never explained how subjective experience arises from physical processes.
Being an outlier doesn't make it wrong.
Materialism is the outlier here, not the default, and it has never explained how subjective experience arises from physical processes.
It's a pattern. The same way letters arise out of pixels on your screen.
From the screen's perspective, there are no letters, only pixels. It doesn't mean there is a "pixel soul."
I believe like the majority of humanity historically that I have a soul
It seems that your position is that the frequency of a belief across human history determines truth?
For large swaths of recorded history, earth was considered the center of the solar system. Given your reasoning, I should expect that is a belief you hold?
Is it possible that popularity of an idea is not a good measure for factuality?
What if "you" are a pattern of linear algebra at the core?
Nice! Can it open multiple files at a time?
What did you use to record the video on the home page, if you don't mind me asking? I need to do something similar. One tip I've seen is to record at a higher resolution than you need, then scale down. The demo is good, but looks a little grainy at points, FYI.
I'm not the OP.
Not everyone is paying for LLMs, even now. So I think it is perfectly reasonable to assume good intentions, here.
Someone spent their own tokens to ponder your code and thought they'd share the result. For anyone else looking, like me, I can see that this is probably going to come up relatively clean without having to spend my own tokens, or install it, and I'm more likely to, now that I can see that.
Sorcery - open source app and protocol that, together, let you share source code links that open in each user's favorite editor, right on the linked line.
Supports VS Code, Neovim, IntelliJ/JetBrains Family, Zed, etc.
About to do the first beta release this later this week.
The protocol is "srcuri" (pronounced, "Sorcery")
This site is: https://srcuri.com/
Source code: https://github.com/browserup/sorcery-desktop
Also--cool editor!
I took a look--it seems like you can pass a path on the command-line to open to. Can you pass a line number, also?
the tech is real and has great promise.
This was very true of the dotcom bubble. The entire "web" was new, and the promise was everything you use it for today.
Pets.com was a laughing stock for years as an example of dotcom excess, and now we have chewy.com, successfully running the same model.
Webvan.com, was a similar example of "excess" and now we have Instacart and others.
I looked up webvan just now--the postmortem seems relevant:
"Webvan failed due to a combination of overspending on infrastructure, rapid and unproven expansion, and an unsustainable business model that prioritized growth over profitability."
Some people treat politics like a tribal sport where "morally OK" is determined solely by which team did it.
Their mental model of the "other side" is someone who is similarly team-driven.
These folks get really confused when "whatabout your team?" falls flat on people who want to live by principles or morality, rather than hat color.
Not the op, but I think about that. Here's what I came to, for the moment:
* LLM's are lousy at bugs
* Apps are a bit like making a baby. Fun in the moment, but a lifetime support commitment
* Supporting software isn't fun, even with an LLM. Burnout is common in open source.
* At the end of the day, it is still a lot of work, even guiding an LLM
* Anything hosted is a chore. Uptime, monitoring, patching, backing up, upgrading, security, legal, compliance, vulnerabilities
I think we'll see github littered with buggy, unsupported, vibe coded one-offs for every conceivable purpose. Now, though, you literally have no idea what you're looking at or if it is decent.
Claude made four different message passing implementations in my vibe coded app. I realized this once it was trying to modify the wrong one during a fix. In other words, claude was falling over trying to support what it made, and only a dev could bail it out. I am perfectly capable of coding this myself, but you have two choices at the moment--invest the labor, or get crap. But, then we come to "maybe I should just pay for this instead of burning my time and tokens."
Injecting ENV variables into the template would be super useful.
gemini -p "Say hello"
Says hello, and just returns right away.
The gemini doc for -p says "Prompt. Appended to input on stdin (if any)."
So it doesn't follow the doc.gemini "Say hello"
Fails as it doesn't take any arguments.
For comparison, claude lets you pass the prompt as a positional argument, but it does append it to the prompt and then gives you a running session. That's what I'd want for my use-case.Feedback: A command to add MCP servers like claude code offers would be handy.
Sounds good on paper, but it has a game theory problem. If your efforts can always be out-raced by someone using AI to do autonomous work, don't you end up having to use it that way just to keep up?
I re-compress my thoughts during editing. That's how I write normally. First, a long draft, then a short one. Saving writing time on the long draft is helpful.
Slop is slop, whether a human or AI wrote it--I don't want to read it. Great is great. Period. If a human or AI writes something great, I want to read it.
Assuming AI writing will remain slop is a bold assumption, even if it holds true for the next 24 hours.
“I didn't have time to write a short letter, so I wrote a long one instead.”
- Mark Twain
Is the percentage meaningful, though? If an LLM produces the most interesting, insightful, thought-provoking content of the day, isn't that what the best version of HN would be reading and commenting on?
If I invent the wheel, and have an LLM write 90% of the article from bullet points and edit it down, don't we still want HN discussing the wheel?
Not to say that the current generation of AI isn't often producing boring slop, but there's nothing that says it will remain that way, and a percent-AI assistance seems like the wrong metric to chase to me?
Working gVisor Mac install instructions here.
https://dev.to/rimelek/using-gvisors-container-runtime-in-do...
After this is done, it is:
docker run --rm --runtime=runsc hello-world
runsc / gVisor is interesting also as the runsc engine can be run from within Docker/Docker Desktop.
gVisor has performance problems, though. Their data shows 1/3rd the throughput vs. docker runtime for concurrent network calls--if that's an issue for your use-case.
In ollama, how do you set up the larger context, and figure out what settings to use? I've yet to find a good guide. I'm also not quite sure how I should figure out what those settings should be for each model.
There's context length, but then, how does that relate to input length and output length? Should I just make the numbers match? 32k is 32k? Any pointers?
I'm a fan of TrilliumNext, which is open source, for this: