HN user

superluserdo

234 karma
Posts3
Comments59
View on HN

I basically implemented exactly this on top of whisper since I couldn't find any implementation that allowed for live transcription.

https://tomwh.uk/git/whisper-chunk.git/

I need to get around to cleaning it up but you can essentially alter the number of simultaneous overlapping whisper processes, the chunk length, and the chunk overlap fraction. I found that the `tiny.en` model is good enough with multiple simultaneous listeners to be able to have highly accurate live English transcription with 2-3s latency on a mid-range modern consumer CPU.

The real answer is it's completely domain-specific. If you're trying to search for something that you'll instantly know when you see it, then something that can instantly give you 5 wrong answers and 1 right answer is a godsend and barely worse than something that is right 100% of the time. If the task is to be an authoritative designer of a new aeroplane, it's a different story.

I wish I could go back to the days of doing almost anything at all without having to tell a server what a motorbike or traffic light is.

It is and has been for a while, but most of the more flashy and exciting developments in ML and AI don't have very much applicability to LHC event processing. To be able to state any kind of finding about some aspect of physics based on the scattering of particles in the accelerator and their decays in the detector, you need to take the background of all events and make multivariate discriminants on the data in order to enrich your signal as much as possible while throwing as little as possible away. This requires you to have a rigorous and verifiable statistical "paper trail" from start to finish, so you can say with confidence intervals how much signal and background you ought to have, vs how much you measure in your data after processing it. An overly broad black box doesn't really work for this kind of introspection.

I wouldn't write it off as a bubble, since that usually implies little to no underlying worth. Even if no future technical progress is made, it has still taken a permanent and growing chunk of the use case for conventional web search, which is an $X00bn business.

Speaking of mathematical missteps relating to bases, I've always been baffled by why we refer to a base system by the number above the highest representable single digit. Every base is "base 10" in that case! Why is binary referred to as "base 2", when the number 2 doesn't even appear? Wouldn't it make infinitely more sense to refer to our conventional number system as "base 9", binary as "base 1", unary as "base 0", and hexadecimal as "base F"? Or we could have used a more sensical word like "ceiling" or "roof" in that case, to convey that it's referring to the highest single-digit value in the system.

But day after day, over decades, surely the minuscule amount of heat being generated by activity on the surface has SOME cumulative effect.

No it doesn't, because the heat escapes out into space, so the minuscule effect is only an immediate one, not a cumulative one. That's why the greenhouse effect, being cumulative, is much stronger. If I put on an extra coat every day, it's not the heat output from my muscles in putting the coats on that is making me feel hot, it's the increasing number of coats that I'm wearing. Likewise, if I burn a tonne of coal, then that heat will have essentially disappeared overnight, but the global warming-causing CO2 will stick around in the atmosphere for another [very big number] years.

This very confident sentiment comes up in every comment section about IPV4/6

- "Updating the standard" is making a new protocol

- "forced suppliers to patch updates" - how?

- "USB backwards compatability that shiz. The fact I can take a modern usb device and plug it in a 1.1 gen port and it still just works. Why the hell isn't ipv4 like that for upgrades?" - because you're changing the address space of the protocol. If the new standard can address more than 2^32 things, then it won't be backwards compatible with v4.

- "Seriously is there any real technical hurdle why we didn't do it this way?" - Assuming you're talking about having a variable-length address from the start in IPV4, because I assume having a non-fixed packet header size would be much more computationally expensive and violate a lot of assumptions that you can make when the header is fixed (having a fixed region of the buffer that is known to always be the full header). You'd be much better having a fixed-length address that is enough to cover all possible nodes in the network - exactly what IPV6 has done.

- "Astounding they baked it in just xxx.xxx.xxx.xxx" - IPV4 was first deployed in 1982. Wikipedia tells me that the year before, there were just over 200 nodes on the ARPANET. I think you're doing a bit of a disservice to the people who designed this stuff by castigating them for not factoring a 20'000'000x increase in network size into their protocol.

His stance on compiler optimizations is another example: only add optimization passes if they improve the compiler's self-compilation time.

What an elegant metric! Condensing a multivariate optimisation between compiler execution speed and compiler codebase complexity into a single self-contained meta-metric is (aptly) pleasingly simple.

I'd be interested to know how the self-build times of other compilers have changed by release (obviously pretty safe to say, generally increasing).

Most users just want something that gets things done with as little friction as possible.

It's funny to read this as someone that always dreads having to get a new phone, or install some proprietary app, or reinstall a Windows VM, or sign up to some service, exactly because I know how much friction is going to be deliberately put in my way from every party involved.

"No you can't just change your phone's service provider", "No you can't unlock your phone's bootloader", "No you can't just boot into a new copy of the OS without signing into a bunch of things", "No you can't install this Windows image on a device without a trusted platform module", "No you can't just install a program from the web", "No you can't just look at all the files on your phone", "No you can't just have an app sync photos from the SD card", "No you can't create a login without giving us your phone number", "No you can't opt out of our 'telemetry'", "No you can't view the video you're paying for without the right OS/browser/monitor/cable".

Coming soon: "No you can't view the URL of the page you're on", "No you can't install an ad-blocker", "No you can't access this site without an attested, locked-down OS-browser stack", "No you can't install a different OS on this PC", "No you can't use a local, unlicensed generative AI model".

I think it's time we reframed the discussion and stopped dumbing users down and pretending that extremely anti-user behaviour is "user-friendly". Using technology used to be challenging because the hardware and software were still being bootstrapped to a point where it was fast, simple and bug-free to use. Now the experience of using technology is to have to navigate some corporate bureaucracy in every direction, which is totally independent of tech limitations.

Humane AI Pin 3 years ago

IMO "Solves iPhone addiction" is more or less a rephrasing of "people will quickly get bored of this".

It's just a smartphone, except you can't run third-party software, can't directly interface with it, and can't connect it to other machines. And instead of holding an N-million pixel, M-million-colour, extremely high-constrast display directly in your hand, you have to indirectly project (meaning extremely LOW contrast) a single-colour display onto your hand from a projector that's shaking around being clipped to your clothes.

The only single hypothetical upside I can see to this tech is that it might lower the two-second delay in looking at my phone caused by putting my hand in my pocket before raising my hand, but you could say that that goes against the goal of solving phone addiction.

Because people find the concept of a person talking to themselves in public a bit weird, and talking into a completely unthinking machine is basically that. Maybe perceptions could change when low-latency conversational AI is very widespread but I think for the medium term unless there's a second human involved, people will still instinctively see it as talking to yourself, not talking to "someone".

Humane AI Pin 3 years ago

The "laser ink display" looks a bit like the totally bunk display tech of the Cicret Bracelet "product" that VFX videomaker Captain Disillusion did a comprehensive takedown of a couple of years ago https://www.youtube.com/watch?v=KbgvSi35n6o.

While it looks like there are a few videos of apparent actual demos, I haven't seen one yet where the device (and more importantly, the recording camera's settings) are controlled by an impartial reviewer, and I'm extremely sceptical that this is usable in the real world. There's a demo by the founder where one of the inputs is to tilt your palm up, and even in the demo the projection struggles to compete with the indoor lights, nevermind the sun https://youtu.be/CwSeUV3RaIA?t=205.

The pitch of this seems to be "no more distracting screens, and no need to download and manage lots of apps and services". Except there is a (very poor) screen, it's your hand. And you're limited to just one service and set of apps, the one that comes with the device.

It's all well and good saying that the AI can do everything you want, but the real world (sadly) has copyright restrictions and content licensing agreements which an out-of-the-box service by a legit company will have to abide by. If the song I want to listen to isn't available on whatever music service this product is partnered with, could I transfer music files from my computer to this device? There's a lot of use cases like this where you very quickly start to want an actual screen, and actual methods of input more precise and domain-specific than conversational voice commands.

NAT is the problem that IPv6 fixes. Think about the parent comment

if you are making more than 4B addresses routable then any existing IPv4 device will not be able to route some addresses, so you will have caused a split in the internet

This has basically already happened. We've massively extended IPv4 by stuffing extra address bits into the router's port number, and it means that any two devices behind NATs can't directly route to each other.

Even despite this rats maze of proxies and NAT gateways we're still supporting virtually all the applications that consumers use

That's a tautology: "Despite the limitations of IPv4, we're still supporting all the applications that can work within the limitations of IPv4".

Lots of potential P2P applications (that might solve a lot of problems with have with the current centralised model of the internet) either don't make it past the drawing board because of NAT, or have to be encumbered with complex, expensive-to-develop, best-effort NAT-punching behaviour that burdens everyone involved (and can stop an application from being truly P2P by having to run things like STUN servers).

NAT seems to always get a bad rep because it inconveniences the very few that want to have an end to end experience

I think there would be many more that wanted this if it were trivially easy to do

but there has to be some sacrifice to keep the Internet running for the billions of users.

What's the sacrifice in using IPv6?

So in theory you could have a ridiculous computer that runs 1000GW of power through it without heating up.

That doesn't make sense. Power has to be consumed by something. All energy ends up as heat, so if the computer isn't heating up then the power consumption is 0W, not 1000GW.

Or a flying car.

A flying car with superconducting motors would still have to expend the energy to stay aloft by pushing air downwards. Even without thermal losses to electrical resistance, you'd still have losses to friction from the moving components.

https://tomwh.uk/blog/index.html

I've got a few posts on there (although most are still on my "been meaning to write that post for literal years" list). Some at random:

https://tomwh.uk/blog/posts/2021/12/07/temphost/ - Temphost: Host files quickly on a dumb HTTP host with optional time-to-live

https://tomwh.uk/blog/posts/2021/06/29/rsync-backup-restore-... - How to Backup and Restore root-owned Files Over the Network Using Rsync

https://tomwh.uk/blog/posts/2020/04/12/fun-with-decompiled-m... - Some Fun With Decompiled Super Mario 64

https://tomwh.uk/blog/posts/2020/03/28/fake-home-prison/ - Imprisoning Naughty Dotfiles in a Fake $HOME

https://tomwh.uk/blog/posts/2018/02/09/alt-useful-key-vim/ - The Most Useful Key In Vim (Not Escape)