HN user

hdkrgr

153 karma
Posts1
Comments54
View on HN

As a corporate customer, the main point for me in this is Microsoft now retaining (non-exclusive) rights to models and products after OpenAI decides to declare AGI.

The question "Can we build our stuff on top of Azure OpenAI? What if SamA pulls a marketing stunt tomorrow, declares AGI and cuts Microsoft off?" just became a lot easier. (At least until 2032.)

I agree, that seems reasonable!

I was referring more to the users of such a system (what the AI Act would call a 'deployer'). They may have significantly less expertise but could still be required to track real-time energy use. Of course, simply referring to the 'energy label' by the provider could be a viable solution.

I agree and share this concern in principle.

But... have you seen the state of GDPR enforcement? Anyone who made an honest effort is fine. I don't know of any GDPR enforcement action where the indicted company wasn't blatantly and willfully violating or ignoring the law.

FWIW, everything I've seen from regulators and the legislators involved in the nitty-gritty of the act seems to suggest that most of them are really smart people who know what they don't know. They know that AI is quickly evolving and the draft of the law goes out of its way to _not_ be too specific about _how_ to comply. E.g., I would not expect the EU (or national regulators) to bring down 'one right way' to report energy consumption.

The fact that they COULD still bothers me.

The AI Act EXPLICITLY enumerates all use cases that will be considered 'high-risk' (Annex III). If your use case is not on the list (or on the 'prohibited' list), then you're good to go. There's no mechanism where someone opposed to your model can argue you should be high-risk because of supposed harms perceived or dreamt-up by some political group. (Caveat: The list of high-risk use cases will probably be able to be amended by the Commission unilaterally after the regulation is enacted.)

Since there's some confusion about this:

- The AI Act regulates both 'high-risk AI systems' and 'foundation models' and applies different requirements for them.

- 'foundation models' are essentially defined in the act as "very large scale and expensive generative ai models that will probably only be offered via API" (my words). The reason the act wants to regulate them is so that USERs of foundation models have a chance to make their downstream use case complaint if that use case is high-risk. For example, if I'm a health insurance provider and I'm using a chatbot enabled by GPT4 in my health insurance sign-up flow, then my system may be high-risk and needs to be compliant. I need access to some information aobut GPT4 (e.g. expected error modes, potential biases etc) to do that.

- The wording of the act makes a point of highlighting that your run-off-the mill open source generative AI project will not constitute a 'foundation model'. The exact scale at which a project will become a regulated 'foundation model' is not yet clear, but it can be assumed that it will be at least tens of millions of dollars. If you can spend that much on compute an researchers, I think you can spend a few k on becoming compliant.

- The technomancers article confuses requirements for High-risk systems with those for foundation models. (It also gets some of the high-risk requirements completely wrong, but that's another discussion.)

- The stanford HAL website does a great job with the facts! I really value seeing thoughtful contributions to the discussion like theirs. (Especially from an American institution!)

it goes further than that. the technomancers blogpost gets a lot of the actual requirements completely wrong (for example the supposed requirement for third-party or government "licensing". Which is nowhere in the Act).

What really frustrated me about this whole discussion is seeing some SV heavyweights quoting this article uncritically and screaming about how stupid the EU is again, while referring to supposed requirements that are nowhere to be found in the act. I would assume these people have access to the best information in the world, yet they don't seem to have had any of their staff actually read the draft. :/

FWIW, I quickly wrote up some of my thoughts about what the technomancer's article gets wrong at the time, but then didn't get around to polish and publish them. If you're interested, here are my notes: https://gist.github.com/heidekrueger/bdee0268ecdad5f6b56f557...

Edit: I want to emphasize that I DO share some of the concerns that the blogpost raises about the current draft of the act. I just wish we could have a meaningful discussion about it rather than namecalling and fearmongering.

Sure... but maybe the GPU is sitting idle 40% of the time while still consuming 200W. Should I have to break this idle energy consumption down onto actual use (assuming the server/gpu is only used for this one model)? I guess it would make sense, but... WHO should do this and then continually update the model documentation when idle rates or the hardware changes?

I do think this will be a useful metric, and it seems obvious that the hyperscalers will have a feature helping you keep track of energy use and emissions of the resources you rented. But why demand this on the level of an individual model/product? For these foundation models, I think it's reasonable to assume they will all be trained on hyperscaler-provided gpu-clusters, so there'll likely be an off-the-shelf funcitonality by AWS/Azure/GCP to report this number, but the draft of the EU AI Act also demands tracking energy use for other 'high-risk' AI systems which companies may plausibly train and/or deploy on-prem. Good luck tracking the per-token energy use of your model that's running on some on-prem server on last-gen GPUs.

I had the chance to talk to a staffer of one of the MEPs leading the political negotiations in EU parlament committee a few weeks ago. His take was that the pro-tech/pro-business parties conceded the 'AI users must track their energy use'-point to the Greens in the latest draft (which is the parliament's counterproposal to earlier drafts by the commission --think EU executive-- and council --think governments of the member states--) because it's so unrealistic in practice that it's likely to be stricken out of the law again during the final negotiation round between parliament and council negotiators.

I really hope that'll be the case. FWIW, I believe companies _should_ be required to keep tabs on their (and their supply chain's) emissions, but demanding that this be done at model/system level by data scientists is just ridiculous.

edit: grammar

There are also significant consequences and side effects to having a 'free market', they are called externalities and in the case of fossil fuels they lead to huge unconsidered costs.

I'm putting 'free market' in quotes, because even beyond externalities, the current system with zonal uniform pricing is only 'free' for a very narrow view of what constitutes a commodity or a marginal cost.

Consider two households: (a) in northern Germany close to the Danish border, located in a small town with lots of local wind turbine capacity, (b) in Southern Bavaria with very limited renewable energy production close by. There's also no adequate power transport infrastructure to get the renewable energy from the North (or elsewhere in Europe) to household (b), mostly because local politicians in Bavaria oppose putting up any visible infrastructure, whether power lines or wind turbines).

Now, during peak wind hours, a significant portion of the wind turbines in the north will go offline because the network cannot handle the load, while household (b) still needs to get their power from gas. [This is not a contrived example, but reality in Germany.] Yet (a) and (b) both pay the same price -- that of the gas producer. How is this a 'free' market?

But to your point, the good news is that people are taking electricity market design very seriously, not lightly. (Section 6 of this white paper is a good read that outlines many of the current market inefficiencies (Disclaimer: my former academic advisor and a few former colleagues are coauthors)): https://synergie-projekt.de/wp-content/uploads/2021/12/Elect...

In the linked Netzpolitik article (in German) it's pretty clear that the 90% refers to Precision (i.e. 90% of flagged instances are True Positives), not accuracy. Still, the article (and probably the primary documents as well) do a horrible job of differentiating such concepts.

[For clarification: Don't understand as my post as an argument for Chat Control, please :)]

When I was studying abroad, I lived in a temporary student dorm that was placed in an industrial district with a special permit from the government.

I tried to order a textbook online and my transaction got flagged as suspicious, so I had to call a support person, and he wasn't having it. - foreign credit card - address marked as non-residential area - sketchy email-address using their company name

Had to take the bus to a bookstore.

For me, it's a Shawn in Colorado who goes to bible study, renovates houses, and signs (me!) up for every Republican newsletter he can find.

Also: When I use Facebook's feature "show data that others have uploaded about you" (or similar), it is full of this guy's stuff that was provided to facebook (and attributed to me) by businesses this guy has relationships with.

Nothing I can do to remove it.

To be fair: As a patient, except for them scanning my insurance card, I see very little evidence that would suggest that most of data exchange isn't being done via fax, snail-mail, or people talking into phones.

Why in the world do I get a piece of paper from my doctor that I'm supposed to mail to my insurance provider (or scan and upload if you're lucky) when I'm being diagnosed with something?

Doctor's offices are the least digitized businesses around.

There's first signs of this getting better, but I can't wait for things to change...

I wouldn't say this is true for "the majority user's of R" at all.

But for "the majority of professors who tangentially use R code in classes on statistics/bioinformatics/economics/finance (anything not explicitly about R and/or Data Science best practices)"? Absolutely.

The R code you see in industry (or academic labs where someone cares about modern R) looks vastly different from those script examples in college that are most people's first impression of the language.

This makes it sound more dramatic than it is.

Let's assume the "real" share of infectious people in the population is 10/100k (That's twice as much as the most recent reported 7-day incidence (5.1/100k) of new infections for Leipzig, the area where the study was performed.)

Further let's assume a PCR Test has a 30% FNR, and a 10% FPR (numbers completely made up, I don't have a source).

Then out of 1500 People who have tested negative, we'd expect 0.05 to have the virus.

This has been exactly my experience. As a data scientist (not a lawyer!) I had to ensure that some of our existing data processing pipelines complied with GDPR (and make sure we could comply with its reporting requirements.)

I found the Articles well-structured, easily understandable, and overall plainly reasonable. In my experience, those who complained about the 'bureaucratic overhead' of making their pipeline compliant were those who were in charge of processes that clearly violated the spirit of the law, trying to press them into the letter of the law somehow.

1. They absolutely can and do, but many sell themselves short.

2. Big caveat: There are absolutely many conservative managers who insist on speaking perfect German / not switching the team language because of the new junior dev, so there will definitely be job openings where immigrants are discriminated against at those big corporates, but IME that doesn't apply to the companies as a whole.

The reverse is obviously also true: If you want to work in a young, international, open culture, you might prefer startups, but most of them offer lower salaries.

R 4.0 6 years ago

Wholeheartedly agree.

I used to do mostly data analysis in my day-to-day work and R was my go-to and absolute favorite language for years in terms of usability for data analysis. Doing actual software development in R is quirky at best, to be honest.

Nowadays I write code for research that requires 'actual' software development, so I've been using python almost exclusively (with pytorch under the hood, which I love.) No doubt, python is a better language for software engineering.

Nevertheless, for analysis I just cannot warm up to numpy/pandas/matplotlib _at all_. When it's time to analyze results of my experiments or produce publication level graphics, I write my python results to disk and use the tidyverse as a last mile solution.

R 4.0 6 years ago

No aliases (afaik), but convention is that you explicitly call mypackage::mean in cases where names might even hint at being ambigous.

FWIW, the German federal government (of all people), is organizing a distributed hackathon this upcoming weekend: https://wirvsvirushackathon.org/

Seems like it's supposed to be (mostly) in German language as well. If they do and market this right this might actually be an opportunity to get non-tech people to contribute important PROBLEMS that can be solved by tech right now, so we don't just get another few dozen news aggregators and case tracker dashboards.

Generally, the parents work in different places of course.

As parental leave is a legal right (not given by the employers), employers simply have to comply with the parents' wishes. In the past employers often frowned upon men taking parental leave, but the younger generation has absolutely normalized this behavior.

It should also be noted that during leave, the government picks up (part of) your regular paycheck - so you don't cost your employer anything while you're not there. (except administration overhead etc)

The parental leave doesn't have to be taken in one block and you can also convert it into 'parental part-time'. A somewhat common pattern that double-earning professional parents choose nowadays that I've seen with some of my team members is something like: 1. simultaneous leave for both partners in the 1-2 months after birth 2. leave of one partner for a few months after that while the other partner works full-time 3. a few months of simultaneous part-time (e.g. 3/days week) where on any given day, one partner is at home 4. full-time work of both partners for a while once the kid is old enough for day-care 5. another month or so of simultaneous parental leave after 1-1.5 years that's used for a vacation.