I live there in that city. There are hardly any homeless at all here. Not like other cities at least. I could see it being a major problem in other places.
HN user
deoxykev
meet.hn/city/us-Iowa City Interests: AI/ML, Cybersecurity, Entrepreneurship, Hacking, Philosophy, Music, Startups ---
How about LLM chat over DNS? https://github.com/accupham/llm-dns-proxy
And it typically works on captive portals too before payment.
Meta-commentary always leans nerdier.
Curious to hear what kind of work you do. Because there are definitely fields where productivity as 10x'd because of AI tools.
HTMX and shoelace is an awesome combo. Super fast to prototype things and tweak as needed. Being able to copy paste snippets and directly inject data in a straightforward way is a nice way of working. It limits cognitive overhead so you can focus on the domain logic rather than fight javascript dependencies.
Don't forget to finetune the reranker too if you end up doing the embedding model. That tends to have outsized effects on performance for out of distribution content.
Interesting, I had never heard about min-p until now. From what I understand, it's like a low-pass filter for the token sampling pool which boosts semantic coherence. Like removing static from the radio.
Do you have any benchmarks of min-p sampling with the new reasoning models, such as QwQ and R1?
Yeah, there is a clear bottleneck somewhere in llama.cpp. Even high end hardware is struggling to get good numbers. The theoretical limit should be higher, but it's not yet.
Benchmarks: https://github.com/ggerganov/llama.cpp/issues/11474#issuecom...
I don't think autoregressive models have a fundemental difference in terms of reasoning capability in latent space vs token space. Latent space enables abstract reasoning and pattern recognition, while token space acts as both the discrete interface for communication, and as a interaction medium to extend, refine and synthesize high order reasoning over latent space.
Intuively speaking, most people think of writing as a communication tool. But actually it's also a thinking tool that helps create deeper connections over discrete thoughts which can only occupy a fixed slice of our attention at any given time. Attentional capacity the primary limitation-- for humans and LLMs. So use the token space as extended working memory. Besides, even the Coconut paper got mediocre results. I don't think this is the way.
The fundemental challenge of using log probabilities to measure LLM certainty is the mismatch between how language models process information and how semantic meaning actually works. The current models analyze text token by token-- fragments that don't necessarily align with complete words, let alone complex concepts or ideas.
This creates a gap between the mechanical measurement of certainty and true understanding, much like mistaking the map for the territory or confusing the finger pointing at the moon with the moon itself.
I've done some work before in this space, trying to come up with different useful measures from the logprobs, such as measuring shannon entropy over a sliding window, or even bzip compression ratio as a proxy for information density. But I didn't find anything semantically useful or reliable to exploit.
The best approach I found was just multiple choice questions. "Does X entail Y? Please output [A] True or [B] False. Then measure the linprobs of the next token, which should be `[A` (90%) or `[B` (10%). Then we might make a statement like: The LLM thinks there is a 90% probability that X entails Y.
My take: the distills under 32B aren’t worth running. Quants seem to impact quality much more than other models. 32B and 70B unquantized are very good. 671B is SOTA.
8x 3090 will net you around 10-12tok/s
Have you hit any non-determinism errors keeping workflow state outside temporal?
Hey, I’m building agents on top of temporal as well. One of the main limitations is child workflows can not spawn other child workflows. Are you doing an activity for every prompt execution and passing those through other activities? Or something more framework-y?
Imhex is a really great frontend for Capstone. https://github.com/WerWolv/ImHex
Are you able to run 405B? 4Bit quant vram requirements are just shy of 192GB.
4 bit quants should require 85GB VRAM, so this will fit nicely on 4x 24G consumer GPUs, plus some leftover for KV cache optimization.
How does this compare to LayoutLMv3? Was it trained on forms at all?
Hi there, I would be interested in a chat about those back-office patterns and use cases. Could you send an email to a2V2aW4gQCBkZW94eSAuIG5ldA==
I use ansible to deploy and sync scripts, services, etc. I think it would work well for you use case as well.
Here is the list from the thread: --- 1. A Man Named Pearl
2. Once Upon a Time in Northern Ireland
3. Microcosmos
4. Crip Camp!
5. Keep The River On Your Right
6. All The Beauty And The Bloodshed
7. Harlan County USA
8. Stay on Board: The Leo Baker Story
9. Good Night Oppy
10. We Met in Virtual Reality
11. All That Breathes
12. Still
13. Can’t Be Stopped
14. The Amazing Jonathan Documentary
15. Searching for Sugarman
16. AKA Mr Chow
17. The Pigeon Tunnel
18. Little Richard: I Am Everything
19. I Know That Voice
20. Genghis Blues
21. Kings of Pastry
22. 20 Feet From Stardom
23. Stevie
24. Anvil: the Story of Anvil
25. Searching for Sugar Man
26. King of Kong
27. The Gleaners and I
28. Touching the Void
29. Cane Toads: An Unnatural History
30. 20 Days in Mariupol
31. Minding the Gap
32. How to Survive a Plague
33. Speciesism
34. Free Solo
35. Jim Allison: Breakthrough
36. Seven Up! series
37. Bowling for Columbine
38. The Fog of War
39. Rivers and Tides
40. Capturing the Friedmans
41. Spellbound
42. King of Kong
43. Crip Camp
44. Jesus Camp
45. Jiro Dreams of Sushi
46. Flee
47. Man on Wire
48. Queen of Versailles
49. The Thin Blue Line
50. Time Indefinite
51. The Gleaners and I
52. My Life as a Turkey
53. Happy People
54. Grizzly Man
A default key expiry of 3 days would be helpful here too, as it would mitigate the threat of someone compromising the endpoint and extracting secrets from browser history.
That’s a great idea, I think something like this would work very well as a installable PWA even.
This is cool. Do you mind me asking what videos you use it for?
Here’s a similar project, but for windows AD networks
I would like to know this as well.
I would love an API, and more flexible pricing options. A subscription just won’t work for my use case, which would be just once in a while for a hobby chatbot.
Looks like the hard-coded Windows internal GUID is found in C:\Windows\SysWOW64\Shell32.dll
If you search for the string: "User Choice set via Windows User Experience", you'll find the GUID used to calculate the hash.
Oh man, I wish there was something like this, but for eGPUs.