HN user

lbill

89 karma

Hi, I am a test manager. I code sometimes, to automate some actions because I am lazy.

I also do some sound engineering, video production and music composition on my free time.

lucienbill.github.io

Posts6
Comments44
View on HN

The only way to make check whether a LLM output is true is to do the work (to have it dkne by a real person).

For tasks that are trivial to verify, it's ok: a code compiler will run the code written by a LLM. Or: ask a LLM to help you during the examples mapping phase of BDD, and you'll quickly be able to tell what's good and what isn't.

But for the following tasks, there is a risk: - ask a LLM to make a summary of an email your didn't read. You can't trust the result. - you're a car mechanic. You dump your thoughts to a voice recorder, and use AI to turn it into a textual structured report. You'd better tripple check the output! - you're a medical doctor, attempting to do the same trick: you'd have to be extra careful with the result!

And don't count on software testing to make AI tool robust: LLM are non deterministic.

"Microsoft Bing Copilot has falsely described a German journalist as a child molester, an escapee from a psychiatric institution, and a fraudster who preys on widows.

Martin Bernklau, who has served for years as a court reporter [...] asked Microsoft Bing Copilot about himself. He found that Microsoft's AI chatbot had blamed him for crimes he had covered."

[dead] 3 years ago

I stumbled upon this piece and found out some nasty stuff about the Brave web browser.

It's quite hard to create something that works optimally on a first try. I think I never managed to do so! The first schema was slow, but hey, it worked, good job! That is a huge first step. Modifying it so it runs much faster can definitely come next.

Yep, you still have to deal with all kind of people when you do OS, and that isn't always nice.

I once made a PR to a project I use a lot at work. The PR was about the documentation of the project, so no change to any feature, no new thing, just an improved version of the tutorial... Or so I thought. The maintainers refused my PR. My first instinct was to be upset about it, but then I thought "this might be a natural reaction because they said 'no', but maybe, just maybe I can avoid acting like a kid and use my brain". So I re-read their reply: it contained a valid reason and provided me with a good solution to my problem, I thanked them and moved on. They were just right to refuse my PR.

I work as a tester for a some websites: I test the GUI and the REST APIs used by the smartphones apps. These tools make my work a lot easier (they might not be the best for your specific needs, they are just the ones I use):

- Talend API Tester: a Chrome extension to test (you probably guessed it) APIs! I can automate many things with it and use regex in my tests. If I had to do all that manually, I'd have become crazy by now. - Watir: a Ruby library that is essentially a wrapper around Selenium. Watir is easy to use, and I found Ruby very easy to learn. Also, bundler makes the process of keeping my libraries up to date really painless (when my webdriver tells my that it can't communicate with the browser because it has become outdated, a simple 'bundle install' fixes the issue)

[] Note: I don't need to do a fancy program, I just need to write some farily simple scripts that automate actions a verify a few things on web pages.

I don't mind the "-gates" : issues are bound to arise. The only way to avoid "-gates" is to stop producing and selling anything at all. Modifying a manufacturing process to correct a flaw is not always easy and can take some time, so I understand the "We'll fix it in a later iteration" stance.

I think the real issue here is the failure of Apple to acknowledge the technical flaw. That, and the the poor customer service. If it's a design issue, it should be covered in the warranty : it would be costly on a short term perspective, but in this specific case I think it would be worth it, as it would keep customers happy with the brand.

Bad news: Apple takes us for fools, again. this sucks.

Good news: many businesses and influencers (iFixit, Louis Rossman, Linus Tech Tips ...) are talking about it. This is bad press for Apple. The more noise we make about these issues, the better the chances are that Apple improves on its flaws.

I'm not overly optimistic about it: spreading the word about Apple's bad habits might be useless. But trying and failing is in my opinion better than not trying at all : at best, Apple makes better product ; At worst, they just keep on doing what they're doing.

That is a compelling argument. Be careful with the terms of use though, they can hide a few surprises. Example: in Vectary's terms of use[1] I interpret the section 4.6 as "backing-up your data is your responsibility. If our servers burn, we might not be able to recover your data and we won't be responsible for it".

That being said, the limits of Vectary's services seem to be clearly stated, so one can adapt to them (by using some storage drives from OVH, Digital Ocean or any such provider to back-up the data, for instance)

[1] https://www.vectary.com/terms-of-use/

I agree with the author : start simple, then use more complex tools when and if you need to.

I believe it is true about infrastructure, and about features and code as well.

When my team needs to release a functionality for the "product"[1] that we maintain, we have a very simple strategy : we go step by step. The users tell us what they need, and we start to deliver as soon as we can. Our first release only solves a portion of the user's problem, but the second release solves more, and so does the next, and so on. I devised this strategy with the users to make sure they are fine with it: they expect a perfect product... eventually. But in the mean time, they having only some portions of their needs met, and we avoid the feeling of being overwhelmed by a great apparent complexity.

[1] The context is a bit complicated... and irrelevant. We don't exactly maintain a product, but it's close enough for the argument here.

I had my doubts about Rust. I knew it was a powerfull tool, but I thought it was too hard to learn, to complex, and that it was reserved for heavily resources-constrained environments, or any other situation where pure performance was one of the prime concerns.

Then I attended a brilliant conference/demo about Rust, in which the speaker proved that one could build something with Rust without giving much thought about the complex principles of the language. You could get familiar with the language by using it, and then have a much better chance to understand the core concepts of Rust.

The conference is here (in french) : https://touraine.tech/talk/o9KtP8ZrZ130zLZPzdqn/

Does this mean that United Airlines is still using the inadequate system described in the article? In my opinion, public shaming is the last resort: when you tried everything and failed to make your legitimate concerns about cyber-security heard by the company, you go public and hope that the bad press creates some kind of PR issue... But what if it doesn't? What if the public shaming proves useless? What can be done then?

Like many (many many many) people, I stream. I don't get a ton of viewers, but I very rarely get 0. Here are the "magic" tricks I use: - First trick: I didn't start alone. I belong to a small group of friends that were already involved in amateur video productions on YouTube, and had some other friends hanging out and enjoying our videos. We started a twitch channel as a group a established a schedule: 90% of the time, when a member of our group launches a stream, the other members watch and interact on the chat. And so do some of or friends (not 90% of the time though, but it's all right). - Second trick: we associated ourselves with another streaming structure. We quit a while ago because we both evolved toward different directions, but in our time there we met other streamers, because very good friends with some of them and became part of their community... and they became part of our's as well! We do occasionally organize some common twitch events with them. Not because it might bring us some viewers, but because we like playing with them. - Third trick: since we are a group, it is very easy for us to keep an active presence on tweeter and facebook. When something really funny happens on our live-stream, there is usually someone available to do a clip and publish it. - Last trick: we all have jobs, we all have lives. We don't rely on twitch to get revenue nor to get social interaction. In other words: if we fail on twitch, we'll be quite all right.

When we stream we talk, play and have a ton of fun... and sometimes a random viewer discovers us! And when we're lucky, he/she decides to become a regular. While that's nice, it doesn't really matter: we're doing it out of passion, not for the numbers. If success finally comes, good. If boredom comes before, we'll just stop and move on.

(note: I won't post our channel name here, it isn't the point of this post. Besides, we are french, and it turns out that most people in this strange world don't speak french)

You are right! I work on test automation for end-to-end testing. I dare say my work is very useful and has prevented a lot of bugs from hitting our prod. But at my workplace we also all agree on the fact that it is often quite painful to locate the cause of the bugs that I report, precisely because we don't have enough unit tests on our old code.

Since we live in a world of limited budget and time to spend, I agree with the conclusion of the article: "use unit tests where it makes sense". It is what we try to implement on our new code (and when refactoring legacy code).

In a way, the "luck" factor makes the game more realistic. You plan your actions and send your best troops for the jobs, betting on the high chances of success of the mission... but just as well as in real life, anything can happen! It makes me feel more involved, and this "shit can happen" aspect is precisely what I love the most about Wesnoth.

EDIT : typo

I get that Intel feels threatened by AMD. They are trying to impress the consumers... but bullshitting a demo is a very bad move! When a consumer decides to build a new PC, the characteristics of the product matter, but so does the reputation of the company that manufactures it. Right now Intel is putting too much effort into sketchy marketing practices: it undermines the actual work being done on their processors by some very talented people.

Presenting it as an extreme overclocking demo would have been a much wiser option.

I really like Blizzard's solution : {$Username}#{$number} is very practical! It complicates thing a bit when you want to share your contact info to a friend, or when you try to remember a specific battletag, but it solves the uniqueness problem. And to be honest, on most sites I end up using numbers at the end of my username anyways, such as "Username0037".