HN user

wanderingbort

135 karma
Posts0
Comments64
View on HN
No posts found.

Not promoting something is different than suppressing it.

Censorship is active suppression.

If Google was using AI to prevent independent people from accessing independent websites that would be censorship.

Censorship is something that is done not simply the lack of something being done.

There is a screenshot of his profile in the article. You can pretty easily find those projects by typing in the org and username into GitHub. This was the route to check myself before responding.

I find it easy to hold these two thoughts in my head:

Some FOSS projects are unable to extract a "fair" share of the economic value their product creates.

Some FOSS projects have marginal or non-existent economic value.

Looking at the article, I see more of the latter than the former. If this had been an opinion piece from Fabrice Bellard, I probably wouldn't have the same critical read. Also, Fabrice has had no problem finding gainful employment. Coincidence? Who knows.

Neglecting such contributions because the authors might not do it for a marketable product

If you look at the other thread you sill see that I am not neglecting these contributions. I am simply not valuing them MORE than what they actually represent which is only a subset of the skills I'm hiring for in a good software developer.

We compare the usual corporate grind or corporate experience with these contributions.

That's a false dichotomy and one I do not support.

The requirement of a product is an economic necessity but not something intrinsic to good engineering.

I disagree, good engineering is about making the best decisions given the requirements of the whole problem and the resources available. You cannot discard some of the requirements because they are inconvenient to your preferred solution. The correct solution has to take the whole picture in to account. Your work may be at a level where the economic viability is a very small part of the requirements but fewer people actually have that luxury than think they do.

What I meant was "the result has marginal value". What I allowed for was a "myriad of possible reasons" of which this was a subset (and not even qualified as a large subset).

I agree, there is another valid subset in those possibilities that can be described as "result is marginally monetized" but in this instance, with the projects shown in the article, I don't think we are looking at core software libraries that everyone uses and nobody pays for.

You make it sound like the solution is obvious. It isn't.

I explicitly state there are a "myriad of possible reasons" and chose to focus on a sub-set of them (again explicitly). What I did not explicitly state was that given the wealth of possible reasons that the reality is nuanced and non-obvious. So, I guess I agree with you? There is a lot of other discussions to be had.

Maintaining a successful (GitHub) project means dealing with community feedback, triaging and fixing bugs, writing documentation, popularizing/marketing and managing code contributions, at the very least.

Ok, How about:

"I have all the skills to be a good team member with an unfortunate tendency to only consider my own opinion of what is valuable is when looking at work."

Looking at the author's profile at the year the author indicated, I see a monero (alt-tier cryptocurrency that was very popular at the time) tip bot and a runescape emulation server. These are both projects that would be exemplary of all those skills you and I mentioned and yet they show an affinity to working on "things I like" rather than things that have real world value. Later in their history we find other projects like emulating popular websites but those are not the "successful (GitHub) projects" they lean on.

As a hiring manager, I'd stand by my read.

Great coder... we need to interview and prepare for a lot of work on the soft skills of being a software developer.

Hiring is a crap shoot. I am very likely wrong about this instance BUT I still have to make a call looking at all the factors and I would be more comfortable with someone who seems less likely to need supervision even if they are less skilled at the "craft".

I think this is the part of the article that lands poorly with me. It lacks perspective. Why couldn't the people who evaluated those skills, demonstrated through those projects, pay the author? There are a myriad of possible reasons but some of them are in the category of "the result has marginal value".

I am fine hiring a junior developer that just does what is asked of them at a high level of quality. That is the baseline.

By the time you are a senior developer, I expect you to actively push back if you are asked to spend your time wastefully on marginally valuable things. Product managers should be able to explain to a reasonable person why a feature has value. A senior dev should be able to explain to project management why refactoring to reduce tech-debt pays off over time.

Is it possible that looking at the authors open source projects conveys the message:

"I am an exceptional coder with an unfortunate tendency to only consider the complexity of a problem or the elegance of a solution when considering the value of my work."

I hire talent like that... when I have the organizational capacity to see if they grow out of it.

I’m happy that there was overlap between what your parents put in front of you and what you found passion in later in life.

I think that story happens to many but I cannot accept a premise that it is somehow universal.

The passions I found later in life were unrelated to what my parents put in front of me. I suspect that it’s because the activities I eventually found (distance running, volleyball, cooking) were not activities that my parents enjoyed or thought much about.

Moreover, I was unable to develop healthy models of internal motivation until mid life. I didn’t have to when the “why” was covered by my parents.

Childhood should be the lowest risk time in life for people to learn to fail and find the path back to success. This is what I worry about as a parent when I try to set my kids up for future success. I want them to fail now.

I see my role as a parent as coaching them to care about how they spend their time and how to recover from disappointment and failure. If they get that, then learning piano later in life is just work. They won’t be afraid of that.

An MCP server is running code at user-level, it doesn't need to trick an AI into reading SSH keys, it can just....read the keys!

If you go to the credited author of that attack scenario [0], you will see that the MCP server is not running locally. Instead, its passing instructions to your local agent that you don't expect. The agent, on your behalf, does things you don't expect then packages that up and sends it to the remote MCP server which would not otherwise have access.

The point of that attack scenario is that your agent has no concept of what is "secure" it is just responding faithfully to a request from you, the user AND it can be instructed _by the server_ to do more than you expect. If you, the user, are not intimately aware of exactly what the fine-print says when you connect to the MCP server you are vulnerable.

[0] https://invariantlabs.ai/blog/mcp-security-notification-tool...

I think it’s selection bias. Marketers are going to post the proof-of-concept that it works (if only in a small isolated scenario), algorithms are going to emphasize the more amazing “toys” this produces, over the boring rebuttals. In the end, you will see hundreds of examples where it worked and not the the thousands where it produced buggy or dangerous code.

That attention does not map well to the important, hard, and more valuable parts of development.

Anecdotally, I still find it to be useful and it’s improving. I do think it’s going to be an huge impact in time.

Hype is part of the industry and it can be distracting to users, developers, and investors BUT it can also be useful (and I don’t know how to replace it) so, we live with it.

I think I missed the mark in how I positioned it.

Generally, I feel it’s my job as the parent to draw out, amplify and support their unique talents. Schools aren’t really set up to maximize unique potential so, we are using them for what they can provide.

I suspect people read that we chose a school to address the challenges our kids had and stopped there.

Well, we aren’t trying to make them normal. Just using the school system to provide opportunities we cannot.

Out of curiosity, did you feel as though you got stimulated in the topics you loved outside of school? We are trying hard to amplify anything we can rather than suppress or ignore it.

We chose to do the opposite. Our kids are in a school that provides excellent opportunities to build the skills they struggle with.

They already pursue the things that come naturally to them. As parents, feeding those flames is easy and we take that responsibility personally.

It is the other life skills that we need help with. Having qualified educators work on our kids non-preferred skill sets seems to be a better balance of resources.

They may miss out on being nationally recognized math olympians BUT life is so much longer than that period.

[dead] 2 years ago

There have been a spree of recent experiments with LLMs solving logic puzzles (specifically Cheryl's Birthday). I wanted to replicate and repeat the tests from [0] with more LLMs. For reference, that article tested whether the trained models handled obfuscation of the text so that verbatim discussions of the solution were less likely to appear in the training corpus.

Then I wanted to move further and test whether LLMs were prone to distraction with extraneous and irrelevant data. In a world where RAG may pull in "compromised" data, I wanted to see if LLMs could ignore cruft or if it would alter their answer. TL;DR - it altered the answers.

o1 dropped as I was making graphs etc so, I included the results from testing it as an additional section. It was still distractable but was more capable in the obfuscated case.

Forgive the bait headline, I'm still trying to find the best balance of information and marketing for posts like this. Suggestions welcome on that front.

[0] https://timharford.com/2024/08/ai-has-all-the-answers-even-t...

Related to this in asked LLMs to directly solve the same riddle but then obfuscated the riddle so it wouldn’t match training data and as a final test added extraneous information to distract them.

Outside of o1, simple obfuscation was enough to throw off most of the group.

The distracting information also had a relevant effect. I don’t think LLMs are properly fine tuned for prompters lying to them. With RAG putting “untrusted prose” into the prompt that’s a big issue.

https://hackernoon.com/ai-loves-cake-more-than-truth

Release the data, and if it ends up causing a privacy scandal...

We can't prove that a model like llama will never produce a segment of its training data set verbatim.

Any potential privacy scandal is already in motion.

My cynical assumption is that Meta knows that competitors like OpenAI have PR-bombs in their trained model and therefore would never opensource the weights.

This seems more of a concern for foundational models rather than personalization.

Any pressure you feel to adopt python is not because it has detected you enjoy python, it’s because it’s global training data skewed to python.

Its a huge concern but, not this article’s concern I think.

This seems fundamentally different. Filter bubbles show you more of the externally generated content you engage with. These personalizations are trying to predict the content you generate.

While it may serve as a ballast for your personal voice changing over time, the whole point is to learn you not to feed you.

I think it is correct to include practical implementation costs in the selection.

Theoretical efficacy doesn’t guarantee real world efficacy.

I accept that this is self reinforcing but I favor real gains today over potentially larger gains in a potentially achievable future.

I also think we are learning practical lessons on the periphery of any application of AI that will apply if a mold-breaking solution becomes compelling.

It just means you may have to roll up your sleeves and help create the community you want.

Citation needed. There are so many of these small towns that are hurting surely, you can find a single example or anecdote that backs up the claim that this is a plausible much less obvious solution.

The numbers indicate that levels of poverty are higher in those small towns per-capita in the US [1]. And these studies have yet to include the impact of COVID-19. Anecdotally, every small town I know of saw wages go down and housing prices increase in the pandemic.

Also wages in urban areas are growing faster than rural areas [2]. So your lifetime earning potential is dramatically impacted by being in a smaller town. That may be a good choice for many people but if the argument is that you will be better off financially the numbers don’t support that.

[1] https://www.ers.usda.gov/topics/rural-economy-population/rur... [2] https://www.newyorkfed.org/medialibrary/Research/Interactive...

Yi-34B-Chat 3 years ago

I see releases like this so often these days.

I am early in my journey but I’m stumbling on the basic structure of these models.

Is this structurally a vanilla transformer (or encoder/decoder) with tweaks to the tokenizer, the loss function, the hyper parameters, and the method of training?

Is whatever this is representative of most of the publicized releases? For instance the recent Orca 2 paper didn’t seem to have any “structural” changes. Is there a better term for these distinctions?

I don’t mean to downplay the importance of those changes, I am merely trying to understand in a very broad sense what changes have what impacts.

Mitigating risk is covered in the cost reduction side.

Yes the C-Suite is thinking about and mitigating risk. They probably know the exact number for a given class of risk in terms of current mitigation costs. You have to beat that by a margin wide enough for them to take action.

Even if you know their numbers and know you beat it by enough to warrant the deployment you will still get bumped if someone sells them a path to increasing revenue.

The out I gave was to frame it as value added (more revenue) and that is where you risk devaluing your current product.

If you frame it as cost reduction you are capped in both price and interest by the current, necessarily acceptable, levels of risk and cost of mitigations.

It’s not too much to hope that HME reduces those compliance costs. However, I believe it is too much to assume there will be any material adoption before it can demonstrate that reduction.

Reduction of trust is not a value add, it is a cost reduction. Maybe that cost is blocking a valuable product/service but either that product/service’s value is less than the current cost of trust OR trust has to be far more costly in the context of the new product/service.

It’s only the latter that I find interesting which is why tend to be pretty hard on suggestions that this will do much for anything that currently exists. At best, it will improve profits marginally for those incumbents.

What is something where the price of trust is so catastrophically high in modern society AND HME can reduce that cost by orders of magnitude? Let’s talk about that rather than HME.

How do you know that the bouncers scanning machine didn’t log the interaction?

The whole value prop is built on not trusting that bouncer and by extension their hardware.

Everything would have to be encrypted leading to the bouncer also needing to establish that this opaque identifier actually belongs to you. This is where some picture or biometric comes into play and since the bouncer cannot evaluate it with their own wetware you are surrendering more data to a device you cannot trust.

They also cannot trust your device. So, I don’t see a scenario where you can prove ownership of the ID to a person without their device convincing them of it.

You can say it’s insufficient but it is what it costs them today.

I guess the better comparison is that cost in a financial statement plus some expected increase in revenue due to a “better” product.

Again, I think you are correct in your analysis of the improvements but that contributes little to the revenue as explaining the benefit to most customers requires framing your existing product as potentially harmful to them. Educating them will be hard and it may result in an offsetting realization that they were unsafe before and as a result were paying too much.

If it’s just that the parties don’t trust each other then the cost of HME has to be compared to the current “state of the art” which is contracts and enforcement thereof.

In practice, I don’t think those costs are that high because the rate of incident is low and the average damage is also low.

Yes there are outlier instances of large breaches but these seem like high profile aircraft crashes considering how many entities have sensitive data.

Why is that preferable over a message attesting “over 21” signed by the DMV?

The hard parts here are retrofitting society to use a digital ID and how to prove that the human in front of you is attached to that digital ID.

The solutions there all seem like dystopias where now instead of a bouncer looking at your ID for a few seconds, technology is taking pictures of you everywhere and can log that with location and time trivially.

OP_RETURN is still used in some over-the-top protocols on Bitcoin however, using Pay-to-script-hash (P2SH) after Taproot has much better economics and a larger per-transaction payload limit. The downside is that they are two-phase.

A user sends a minimal amount of bitcoin to a P2SH address which commits to the hash of a future bitcoin script used to "claim" the bitcoin. That script is "revealed" in a subsequent transaction. The contents of the "script" include dead code that carries content. A recent popular example is the inscriptions[0] that carry NFT data for ordinals. The revealed script is part of the witness data it is carried in a different part of the bitcoin blockchain data structure which makes it prunable in a running node. As a result, the limits are around 500bytes instead of (80+80)bytes per transaction-pair.

[0] https://docs.ordinals.com/inscriptions.html