HN user

putlake

994 karma

Created Diffen.com, SmashTheBubble.com and Hreflang.org. @thisislobo on Twitter.

Posts38
Comments207
View on HN
github.com 6mo ago

Microsoft releases VibeVoice-ASR, an open speech-to-text model

putlake
3pts1
jasuja.us 6mo ago

What I learned porting JustHTML to PHP with GPT 5.2 Codex

putlake
5pts2
jasuja.us 10mo ago

Money mistakes you didn't know you're making

putlake
34pts35
chromewebstore.google.com 11mo ago

Show HN: Follow/Block Users on Hacker News (My First Chrome Extension)

putlake
4pts0
blog.diffen.com 1y ago

I broke Grok3 and sent it into an infinite loop (without meaning to)

putlake
3pts0
blog.diffen.com 2y ago

The 5 building blocks of general intelligence

putlake
1pts0
blog.diffen.com 2y ago

AI Companies vs. the Open Web

putlake
1pts1
indianexpress.com 2y ago

New shoe sizing system proposed for Indians

putlake
1pts0
techcrunch.com 2y ago

Anthropic details "many-shot jailbreaking" to evade LLM safety guardrails

putlake
10pts2
www.theguardian.com 3y ago

US food pesticides contaminated with toxic ‘forever chemicals’ testing finds

putlake
12pts2
github.com 3y ago

PhotoGuard: Defending Against Diffusion-Based Image Manipulation

putlake
3pts2
www.youtube.com 3y ago

Why does this lady have a fly on her head?

putlake
1pts0
www.sandiegouniontribune.com 3y ago

VP allegedly duped Qualcomm into paying $150M to acquire tech it already owned

putlake
9pts0
news.ycombinator.com 3y ago

Does TaskRabbit get hacked every 2 years?

putlake
5pts5
www.youtube.com 4y ago

Zero-knowledge proofs explained in 5 levels of difficulty

putlake
2pts0
security.stackexchange.com 4y ago

Is Xcode vulnerable due to Log4j?

putlake
3pts0
blog.diffen.com 4y ago

Using Cloudflare Workers to prime Fastly's cache

putlake
4pts0
scholar.harvard.edu 4y ago

Investing in the Unknown and Unknowable [pdf]

putlake
3pts0
blog.diffen.com 5y ago

Using Cloudflare Workers to improve your Fastly cache hit rate

putlake
5pts0
github.com 5y ago

Chrome Flags for Tooling

putlake
2pts0
news.ycombinator.com 5y ago

Ask HN: Would you pay to warm your CDN cache?

putlake
2pts1
blog.diffen.com 5y ago

Priming my CDN cache using another CDN

putlake
1pts0
ncase.me 7y ago

How to Remember Anything Forever-Ish

putlake
3pts0
news.ycombinator.com 7y ago

Ask HN: Should browser extensions mess with Content-Security-Policy

putlake
2pts0
smashthebubble.com 7y ago

Show HN: Smash the Bubble

putlake
1pts0
minecraftlist.org 8y ago

The State of Browser Push Notifications

putlake
2pts0
github.com 8y ago

CSS Gridish: Auto-generate CSS grid code for your design

putlake
2pts0
www.npr.org 9y ago

As Congress Repeals Internet Privacy Rules, Putting Your Options in Perspective

putlake
3pts0
www.youtube.com 9y ago

The Future of Brain Stimulation (rTmS for Depression, OCD, Bipolar)

putlake
1pts0
scalegrid.io 9y ago

Cassandra vs. MongoDB

putlake
2pts0

I think it was when the LLM asked me a question at the end of its response. It felt like something other than a machine. Until then the pattern was me asking a question and ChatGPT giving me an answer, with or without hallucination. When it asked me a follow-up question it felt like talking to a being with agency. An entity that has thoughts or ideas or questions of its own.

DeepSeek v4 3 months ago

By "completely Nvidia-free" do you mean Nvidia wasn't used for training nor inference? Because if it's only inference, we know that Opus already can run on TPUs. Not to mention Gemini.

Thank you for the link. I tried to opt out. They sent me this email:

Equifax Workforce Solutions (provider of The Work Number) has received your employee request communication, but additional information described below is required to fulfill this request.

We will be following up with a secure email to obtain the below requested documents:

Proof of Identity:

Provide a copy of one of the following (must include current/legal name):

- Driver's License (must be current) - Paystubs (must be dated within 60 days) - State or Government Identification Card (must be current) - Social Security Card - Military Identification Card - Passport (must be issued from U.S.A. and be current) - W-2 or 1099 Form (most current year) - Birth Certificate

Proof of Address:

If you are requesting an Employment Data Report (EDR) or selected ‘Mail’ as your preferred method of contact,provide a copy of one of the following (must include current mailing address and be issued within the past 60 days)

- Driver's License (must be current) - Paystub - W-2 or 1099 Form (most current year) - Utility Bill (phone, water, gas, electric, trash or sewer, etc.) - Housing Rental Agreement or Mortgage document - your name must be listed on the document

For Identity Theft Block Requests, along with Proof of Identity and Proof of Address (if applicable), please provide your identity theft report and designation of items to be blocked:

- Identity Theft Report (police report, FTC Identity Theft Report, Police report, or United States Postal Inspection Service)

For Human Trafficking Victim Block Requests ONLY, along with Proof of Identity and Proof of Address (if applicable), please provide victim determination documentation (as described below), and designation of items to be blocked.

Victim Determination Documentation:

Provide a copy of one of the following victim determination documentation confirming that you were a victim of human trafficking, such as:

- Determinations made by federal, state, tribal, or local governments, government agencies, or law enforcement - Determinations by non-governmental entities or task forces authorized by a governmental agency to make such a determination - Self-attestation signed or certified by such governmental agency or non-governmental entity - Determination by court in a case where a central issue is whether you are a victim of human trafficking. (Court documents can be made up of several documents from the court case that together show that the court accepted as true or finding no genuine dispute that you were a victim of human trafficking.)

We will be following up with a secure email to obtain the requested documents.

Data Investigation Team Equifax Workforce Solutions

Mining isn't the only bottleneck with rare earths. There also the processing, which is an industry China has monopolized through sustained investments over decades. They have also improved processing efficiency through investments in technology. It's going to take a while for anyone else to catch up.

The way trademarks work is that if you don't actively defend them you weaken your rights. So Anthropic needs to defend their ownership of "Claude". I'm guessing they reached out to Peter Steinberger and asked nicely that he rename Clawdbot.

Claude in Chrome 7 months ago

The only model available for in-browser chat is Haiku 4.5. Is it just my account (Pro) or are others also restricted to Haiku?

1. There are penalties for underpayment of estimated taxes. And there are even interest charges on penalties. https://www.irs.gov/payments/underpayment-of-estimated-tax-b...

2. The triple tax advantage is not ridiculous. (1) and (3) are not the same thing. 401k for example is not taxed going in (you fund it with pre-tax dollars) but is taxed when taken out. When you withdraw money from your 401k in retirement, you owe taxes on the capital gains that have accrued in the account since you first put the money in. But if you take money out of HSAs for paying medical bills, there is no tax on the capital appreciation you have enjoyed in your HSA account.

Trusts are like life insurance. It's not about when you're old and your kids are all grown up and self-reliant. The point of the trust is when you have younger kids and need to plan for you and your significant other both dying unexpectedly. It's no use naming your 8 year old as a beneficiary because they won't be able to use any of the assets without trusted adults.

What you are saying is very reasonable but companies don't deduct enough in taxes at the time the RSUs vest. I don't know why that is. I'm sure there's some arcane reason. That's why it's important to keep track of this a little bit otherwise you will have an unpleasant surprise when taxes are due.

LLMs are famously bad at individual letters in a word. So something like this never works: Can you please give me 35 words that begin with A, end with E, are 4-6 characters long and do not contain any other vowels except A and E?

100% agree. TSV is under-rated. Tabs don't naturally occur in data nearly as often as commas so tabs are a great delimiter. Copy paste into Excel also works much better with tabs.

Code editors may convert tabs to spaces but are you really editing and saving TSV data files in your code editor?

In your "how to cook the perfect steak" video [1] there's a picture of various doneness levels of a steak. It's a fantastic picture. The creator of that picture will get jackshit from this. Phind gets value, the user gets value but the creator does not.

You're hyperlinking to the source, which is nice. But there's no reason for the user to click through so it won't really help the creator. The upshot of all this is that the open web will have less incentives for content creators. Someone's got to create new data for the AI borg. In future, these creators are less likely to be independent bloggers/photographers. Perhaps biased media outlets will be the only ones with the incentives to create.

[1] https://www.youtube.com/watch?v=cTCpnyICukM#t=55

I don't know why you are getting downvoted except that your opinion is unpopular in this thread. It's a legit counterargument though. However, speaking for myself as an N of 1, the reason I buy an iPhone is because of the assurance that it will receive updates -- especially security updates -- for several years. Android doesn't seem to hold that promise.

My last Android phone was an HTC that came out with this promise of delivering Android updates within 15 days -- a promise they did not keep. https://www.engadget.com/2016-08-25-htc-one-a9-android-updat...

As someone who runs nginx locally for web development, this is scary. One mitigation I can think of is to use this config for you Mac's local nginx:

  server {
    listen       80  default_server;
    server_name  _; # some invalid name that won't match anything
    return       444;
  }
And do the same thing for server_name localhost. For actual apps you are building, use a server_name like myapp.local rather than localhost. (edit: formatting)

As a user, I love Perplexity. As a web publisher, I would probably want to block it. The UX of AI overviews and summaries disincentives small web publishers from creating new content.

It's funny I posted the inverse of this. As a web publisher, I am fine with folks using my content to train their models because this training does not directly steal any traffic. It's the "train an AI by reading all the books in the world" analogy.

But what Perplexity is doing when they crawl my content in response to a user question is that they are decreasing the probability that this user would come to by content (via Google, for example). This is unacceptable. A tool that runs on-device (like Reader mode) is different because Perplexity is an aggregator service that will continue to solidify its position as a demand aggregator and I will never be able to get people directly on my content.

There are many benefits to having people visit your content on a property that you own. e.g., say you are a SaaS company and you have a bunch of Help docs. You can analyze traffic in this section of your website to get insights to improve your business: what are the top search queries from my users, this might indicate to me where they are struggling or what new features I could build. In a world where users ask Perplexity these Help questions about my SaaS, Perplexity may answer them and I would lose all the insights because I never get any traffic.

A lot of comments here are confusing the two use cases for crawling: training and summarization.

Perplexity's utility as an answer engine is RAG (retrieval augmented generation). In response to your question, they search the web, crawl relevant URLs and summarize them. They do include citations in their response to the user, but in practice no one clicks through on the tiny (1), (2) links to go to the source. So if you are one of those sources, you lose out on traffic that you would otherwise get in the old model from say a Google or Bing. When Perplexity crawls your web page in this context, they are hiding their identity according to OP, and there seems to be no way for publishers to opt out of this.

It is possible that when they crawl the web for the second use case -- to collect data for training their model -- they use the right user agent and identify themselves. A publisher may be OK with allowing their data to be crawled for use in training a model, because that use case does not directly "steal" any traffic.