HN user

ghm2180

81 karma
Posts1
Comments68
View on HN

This should be the way. Have a tiny burner phone for maps and any apps that you absolutely can't use without google(it should be a tiny set of < 10 apps hopefully) until you can fully de-google

My current de-google project is categorizing all my pictures on my local NAS to create the memories feature (where it shows historic pics on multiple theme axes). You can get really far with just a few hours of work a month to de-google and some off the shelf image embeddings.

The hero project in this category — what one cannot do trivially as an indie dev — is creating a great fresh PoI dataset. This is tough to do on a planetary scale because its a societal cooperation problem.

Understand Anything 3 months ago

The phrase going around the interwebs is "You can outsource your thinking but not your understanding". A phrase that can at times seem like this weird human<>llm endless loop; depending on what you think you understand and what the llm "thinks" to help you understand, it can seem like an LLM also understand. But it does not.

Its clear one can't really think about anything without building a basic understanding about it. Worth stating that these are distinct from learning. But, I would argue that it is important to know what you *have* to understand now and why is that important. An LLM can help you understand a great many things, you just need to know what you are looking for and that is something no artificial intelligence can really *do* for you. Trial and error, building a sense of self awareness, and talking to people is a better way to know what this is especially for fairly open ended problems.

Its already pretty easy to oneshot an extension aiding scraping and LI can do nothing about it. I've seen people build and install a local chrome extension in a couple of days and have an AI inject itself into devtools and scrape pretty much any website. And that was a few months ago. I don't think there is an easy way to defend against such things anymore. Its a matter of time that defensive programming measures like this become useless.

I use firefox with uBlock Origin's matrix turned on linked in and its cdn is explicitly black listed globally on it. I see links like ~`licdn` or some shit appear with a lot more frequency on webapps in the matrix now a days. I would recommend you all install it and block it actively.

Its disgusting.

About the "they asked us to view it and then fired us for it". Having worked in their RL division(I don't work at meta anymore) this story is quite weird for two reasons:

1. Meta AFAIR paid/compensated people — contractors or recruited via ads — to have them submit their data. There are strict privacy protocol and reviews in place to distinguish data use in these cases vs gen public. This is not to say the process is perfect, but if these users are gen public, I would be very shocked.

2. Hiring contractors to submit data is a more controlled environment VS recruitment of gen pub via ads to submit data, but the former has more well understood privacy disclosures than the latter. This means in practice asking contractors to wear glasses and "move around their surroundings naturally and do things" goes well with basically the privacy practice "the data your are submitting we can view and use all of it for purpose X and nothing but X". BUT this framing is with ad based recruited people — which are general users who willingly submit data — is much much harder. My suspicion is they are running ad based recruiting in general public and while those users may have signed a privacy statement it is very surprising that they did not tighten the privacy practices around the use of the data and who has access.

To be clear we are — in the service of speed — trying to bake data into compiled program output binaries, which is ostensibly faster because code-pages in memory when executed by CPU as instructions already have it?

I do use an email alias everywhere. But I don't believe you can do the same with phone numbers. I tried using my twilio rented number and there is a way systems use to figure out if that is a real number for a person or a VoIP one. Though it is sometimes successful in use for signups and hence spam reduction.

I've commented elsewhere about just having simple rate limits tied to oauth tokens. This should not be that hard.

There is one simple policy: Subscriptions are for use on human scale of comprehension. API Keys are for everything else.

Anthropic can have a machine/bot get rate limited and people can build workflows using `claude -p` or something even better (like an SDK) , all the while using their OAuth tokens for max/pro.

yeah rate limits are the way to go. I don't see how this is not really simple: A human can only type or tts the text in and read responses only so fast, anthropic can use this as a baseline. They can create a client that can back off(rate limit) and wait, like for something like when a user says something like "spawn 10 processes with claude -p to do X" this client can calculate the rate limit and place an event in a queue and developers can use this client's rate-limit and build workflows around it: e.g. a client queue that has a timer event that expires the rate limit that was set can wake up a daemon. There are a million ways to implement the queue <> daemon thingy, its just software at that point.

Since the subscription is hard linked to an OAuth token this should be easy to track too. What am i missing?

Interesting. What kind of context usage does it have when switching between the two providers? Like is it smart about using the # tokens when you go from claude -> codex or vice versa for a conversation?

How does ctx "normalize" things across providers in the context window ( e.g. tool/mcp calls, sub-agent results)?

This is doubly true in Machine Learning Engineering. Knowing what methods to avoid is just as important to know what might work well and why. Importantly a bunch of Data Science techniques — and I use data science in the sense of making critical team/org decisions — is also as important for which you should understand a bit of statistics not only data driven ML.

Any recommendations? I read designing data intensive applications(DDIA) which was really good. But it is by Martin Klepmann who as I understand is an academic. Reading PEPs is also nice as it allows one to understand the motivations and "Why should I care" about feature X.

It can't be that hard to just dump/export the entire JIRA in one day and migrate it to something else like linear.app? i was already exporting HTML dumps of the entire JIRA and using it in local tool calls to ground agents as far back as last year instead of wrestling with JIRA API to get it to work. This was before linear became popular.

The migration would take 1-2 engineering man-days I suppose. But its money well spent.

Let me put this in simpler terms: std::move is like putting a sign on your object “I’m done with this, you can take its stuff.”

and later:

Specifically, that ‘sign’ (the rvalue reference type) tells the compiler to select the Move Constructor instead of the Copy Constructor.

This is the best conceptual definition of what `std::move` is. I feel that is how every book should explain these concepts in C++ because its not a trivial language to get into for programmers who have worked with differently opiniated languages like python and java.

If you read Effective Modern C++ right Item 23 on this, it takes quite a bit to figure out what its really for.

I wonder when will there be something more rigorous on what works clearing house https://ies.ed.gov/ncee/WWC/Search/Products?searchTerm=AI&&&...

I am actually hoping someone there studies such interventions the way they did with CMU's intelligent tutor — which if I recall correctly did not have net strong evidence in its favors as far as educational outcomes per the reports in WWC — given the fall in grade level scores in math and reading since 2015/16 across multiple grades in middle school. It is vital to know if any of these things help kids succeed.

I have a noob question. How can developed economies where populations levels have plateaued continue to be expected to post positive GDPs (and therefore add net new goods and services) yoy?

Homes as assets should pass on, higher cost services of today would be replaced by lower cost which only temporarily would increase units sold(but should eventually plateau because #units/person is not going to change). No component of the GDP will move.

If the fertility rate of developed economies is less than 2.1 then there should be wealth and asset accumulation among the younger people over time. The demand for newer goods and services from the people who get richer is inelastic: most goods and service's prices dont matter to the rich they buy them any ways, and it continues to keep becoming more inelastic.

So within like a several years the demand should just collapse as wealth accumulates a lot and people work less and become more price insensitive. Immigration is set to remain low to developed nations for the next 3-5 years.

This seems to be quite evident in Japan and EU already(though in the EU if you adjust the productivity for work hours the GDP becomes same as America's).

So why do people assume developed countries would even buy this much more new stuff. 1.5 trillion$ worth of new things over the next 5 years?

So the bias is an issue can be handled in a variety of ways, one which I know to work is to use weights on your rarer class when training. You could also use larger margins to make sure you definitely don't mis-classify the rare class at the cost of mislableling your dominant class — presuming you are ok with it. An example is when doctors order breast biopsies, it happens a lot more than the cancer itself and based on a noisy technique of physical exam.

I would offer a stronger more pointed observation, ofen the problem in building a good classifier is having good negative examples. More generally how a classifier identify good negatives is a function of:

1. Data collection technique.

2. Data annotation(labelling).

3. Classfier can learn on your "good" negatives — quantitaively depending on the machine residuals/margin/contrastive/triplet losses — i.e. learn the difference between a negative and positive for a classifier at train time and the optimization minima is higher than at test time.

4. Calibration/Reranking and other Post Processing.

My guess is that they hit a sweet spot with the first 3 techniques.

Has someone done a survey to ask devs on how much they are getting done vs what their managers expect with AI? I've had conversations with multiple devs in big orgs telling me that Managers and dev's expectations are seriously out of sync. Basically its

Manager: Now you "have" AI, release 10 features instead of 1 in the next month.

Devs: Spending 50% more working hours to make AI code "work" and deliver 10.

.. that LLMs can only solve problems they have seen before!

This is a reductive argument. The set of problems they are solving are proposals that can be _verified_ quickly and bad solutions can be easily pruned. Software development by a human — and even more so teams — are not those kind of problems because the context cannot efficiently hold (1) Design bias of individuals (2) Slower evolution of "correct" solution and visibility over time. (3) Difficulty in "testing" proposals: You can't build 5 different types of infrastructure proposals by an LLM — which themselves are dozens of small sub proposals — _quickly_

Just as a note pressure cookers like the Instapot are rated at 1000wh but they don't use that all the time IIRC — only initially when building the pressure. In simple terms it's like using 1kw for 10 - 15 mins and not an hour. So on average very efficient.

Has someone measured the time for max usage and average usage? Could be a good alternative with a 1kwh lfp battery with a lot of juice left over after pressure cooking(no pun intended)