HN user

khiner

252 karma
Posts10
Comments42
View on HN

Throwing in my vote - I wasn’t confused, saw your GH link and a “Zero to Hero” course name on RL, seems clear to me and “Zero to Hero” is a classic title for a first course, nice that you gave props to Andrea too! Multiple people can and should make ML guides and reference each other. Thanks for putting in the time to share your learnings and make a fantastic resource out of it!

One of the first programs I ever wrote - a program to find valid English crossword fills given a grid pattern with optional partial completions.

This project came to mind recently and I looked around on the Wayback Machine. Turns out I posted the jar on MediaFire and linked to it on an old blog on Jan 2, 2011. Luckily, there was [_one capture_ of the jar on MediaFire](https://web.archive.org/web/20240123154949/https://download1...) from oddly recently (Jan 23 2024). I downloaded and opened it on my 2023 MacBook Air, and it ran! Since it's Java, I'm guessing it runs on other computers, too :)

It was a delight to find it still working. Anyone else ever find an old program you thought you lost and get it running again?

Bun v1.0.0 3 years ago

It’s been so fun watching Bun progressing as quickly as it has! Truly incredible work, and the blog post is full of real value and time saved for future me from beginning to end - huge congrats on the 1.0.0 release, excited to see where bun goes from here!

I can think of a few directions for technology aiding in fact checking:

1. Much of finding out what’s true or false is about finding consistency amongst lots of observations. So this is the science direction. If you can analyze lots of data, say from first-hand direct measurements like from a scientific instrument, or analyzing second hand observational data, say from many news sources reporting on a political event. One could also imagine multimodal analysis combining these first and second-hand kinds of data to arrive at a consensus estimate of a “true” perspective. E.g. analyzing video and audio streams recorded at said political event, combined with many text reports of the event. So this point is about data mining, and jointly estimating semantic meaning from natural language and other kinds of data, in a way that’s consistent with everything else considered factual.

2. Provenance-tracking: Think block chain - if we can provably trace a piece of data back to its primary sources, tracking all its modifications along the way, this could help with establishing provenance, and verifying legitimacy of any modifications along the way.

3. Consensus, staking/voting, etc. A lot of deciding what’s true is about seeking consensus. One thing I’m generally optimistic about here is that, for any given fact “out there,” there are many more ways to describe it incorrectly than correctly. So even though it sounds scary for consensus to be an aspect of truth finding, it always will be, and at the very bottom it’s all we can hope for. One way that’s already getting traction to make consensus mean something, is the idea of staking. So you have to put something down on the table when you claim you believe something to be true. Software can (and already is) helping to build confidence behind some claims more than others by backing claims with value (money).

4. Humans are insanely bad at reasoning rationally, because of lots of reasons. Pick your favorite fallacy. We evolved to survive long enough to reproduce and rationality is a happy accident. One could imagine software being less susceptible to simple tricks, could be less incentivized to outright lie for personal gain or power seeking, or claiming to represent their actual beliefs when they are actually pursuing other goals by conveying something they don’t actually believe.

That is true, 3 would help steel my strawman. I agree that we’ll increasingly have capabilities to generate and publish garbage that’s _just_ good enough to generate clicks, and incentives to do this. In addition, I think we’ll increasingly have tools to produce content that is much more rich, imaginative, insightful, and factually correct in our future. Some more interesting questions to me are then: What will the ratio be? How will that ratio compare with what we see today? How easily will I be able to identify misinformation when I care about factual accuracy (again, compared with today)? How easily will I be able to avoid the garbage, vs find the good stuff?

The test use case of constructing a bio for yourself, hoping it accurately summarizes all the extremely low sample size data it happens to have of you in its web crawled training data, seems like one of the worst possible use cases for ChatGPT. It’s right there on the main page that it’s not to be trusted with factual information like this. ChatGPT will hallucinate details. It’s remarkable to me actually how often it will refuse to hallucinate, given that’s basically what its job is. I don’t find it interesting to find all these edge cases where ChatGPT produces empirically false data. It doesn’t even have the ability to look things up! If I were the OP and wanted help writing my bio, I would first write the draft myself, then use ChatGPT to help with the editing, prose, grammar, style, etc. You are the expert on the factual details of your own life, and if you’re surprised that a language model trained on web crawled data ending in 2018 is not, then all I’ve learned is that you don’t know much about what this thing is.

I also don’t buy these arguments of the form, 1. OpenAI’s public ChatGPT app is often factually inaccurate. 2. ChatGPT is an example of a ML system bootstrapped on web crawled text data. 4. Thus, the long term future of our distributed text-encoded knowledge base will be a cesspool of useless gobbledygook.

ChatGPT is a step forward in generative language modeling. It doesn’t preclude the development of other future systems to help us verify factual accuracy of claims, likely much better than humans can. We’ll be ok gang:)

Throwing a shameless plug in here for a set of Jupiter notebooks I made that follow each chapter of this book: https://karlhiner.com/jupyter_notebooks/mathematics_of_the_d...

Reproduces many of the interesting results in Python, providing charts and animations along the way. Hope it’s of use to someone! (I also have a similar series for two of the other three books in this series - Digital Filters and Physical Modeling.)

Shameless plug but I put together some Jupyter notebooks that walk through several of the fantastic books recommended in this thread: https://github.com/khiner/notebooks

I wanted to help myself and other folks develop better intuitions around the material, particularly focusing on short animations to develop better visual intuition along with working code examples of the material in the books.

Books covered (with a notebook for each chapter):

* Musimathics volumes 1 & 2 by Gareth Loy

* Introduction to the DFT by Julius Smith

* Introduction to Digital Filters by Julius Smith

* Physical Audio Signal Processing by Julius Smith

also a couple not directly about audio but helpful for the domain:

* Coding the Matrix by Philip Klein

* Accelerated C++ by Andrew Koenig and Barbara Moo

Hope someone gets some value from these - have fun!

Ah I see. I was on mobile and didn't get the left/right swooshing, but do on desktop. I also find the left/right swooshing from the outside pretty clunky and unnecessary. But the content is really nice, and I personally appreciate a bit of fun in experimenting with content presentation. Esp. given the topic.

For me, it works on both desktop and mobile, chrome & firefox. I don't get the left/right swooshing on mobile, but do on desktop. Mac/iPhone

Couldn’t agree more. I often read some top comments even before checking an article just to get a quick read on whether something is likely e.g. clickbaity. Reading the comments it seemed like there was some consensus that this was a page with mostly information that would be better presented in a flat form like a pdf rather than an interactive website... but it’s actually a super fun collection of interactive visualizations for some of the classic complexity examples, and it worked perfectly for me except for a kind of ostentatious fade-from-white effect on page load.

Colorize 6 years ago

I had so much fun with this. I actually like that it uses the mean rather than some clustered mode in a lot of the cases I tried out. For example, “space” is pleasantly brighter than the dark black/blue hue that would come from that approach. It’s a little warmer and brighter than that, and feels like a truer representation of the concept in my mind’s eye. Obviously it will end up muddy for some things with several dominant colors. The mean approach also allows for some fun experiments like “black and white” which returned a rewardingly exact gray for me. Which led me to “newspaper” which is gray with most weight to the red and some to blue, and looks just right. I just spent about 15 minutes on this. Simple fun and useful, I am totally going to use this for something, thanks!

Wow this looks incredible! I’ve been looking for a C++ plotting library suitable for real-time audio visualization but with a matplotlib-like api. This looks to fit the bill perfectly, and with great documentation. Well done, following.

I used to work for a company that transferred a huge amount of money from tenants to landlords. Fraud is a really big deal in that domain, and we were really excited to update our credit card entry forms to use the newer Stripe Elements forms that include this user behavior tracking, to feed into Stripe radar and in turn use Radar’s api to feed the fraud likelihood data in to our own fraud detection platform. This agent behavior tracking is proudly highlighted in Stripe’s documentation as the feature that it is - they’re not being sneaky about this in any way and this feature is incredibly useful to help companies stop fraudsters from doing serious financial harm. In other words, this is not news. If you could somehow prove that Stripe is selling this data, THAT would be a huge story. But as far as we know their explicit claims that they do not is true. Thanks for your awesome anti-fraud features, Stripe - they helped us help people to pay and accept rent safely.

I also recently took about 10 months off of work, specifically to focus on learning. It was incredible, and I don’t regret it financially. I would often get up at 6 in the morning or even earlier (which I never do) just from excitement about what I was going to learn about and accomplish in the day. Spending my time focused Only on what I was most interested in was incredibly rewarding. It’s going to be awhile before I am able to do it again financially. It is a life-impacting hit to the bank account, especially factoring opportunity costs, but I agree with the OP that it feels well worth it as an investment in myself. Can’t wait til I get the chance again.