HN user

KvanteKat

88 karma

Recovering academic; Network engineer. Interested in statistics, combinatorics, mathematics, hpc, distributed computing, programming, ML, & AI.

Posts0
Comments31
View on HN
No posts found.

Given that you're citing Wikipedia on this, the issue of detecting and fighting auto-generated slop in articles is actually quite fascinating.

There was a really interesting talk given by Mathias Shindler (long time editor of German Wikipedia) at the 39C3 conference about this topic a few months back that is worth a watch for anyone interested in the issue: https://youtu.be/fKU0V9hQMnY

Sure, but the alternative is not really any better: if the choice is between being the guy who got it wrong vs. being the guy who got it wrong _and_ being the guy who persisted in throwing good money after bad, surely the former is prefereable. As far as I see, the fact that they keep going indicates that they genuinely still believe Copilot could pan out and become profittable in the long run.

You can think of it like this:

- The characteristic function of a random variable X is defined as the function that maps t --> ExpectedValue[ exp( i * t * X ) ]

- Computing this expected value is the same as regarding t as a constant and integrating the function x --> exp( i * t * x) with respect to the distribution of X, i.e. if X has the density f, we compute the integral of f(x) * exp( i * t * x) with respect to x over the domain of f.

- on the other hand: computing the Fourier transform of f (here representing the density of X) and evaluating it at point t (i.e. computing (F(f))(t) if F represents the Fourier transform) is the same as fixing t and computing the integral of f(x) * exp( -i * t * x) with respect to x.

- Rearranging the integrand in the previous expression to f(x) * exp( i * -t * x), we see that it is the same as the integrand used in the characteristic function, only with a -t instead of a t.

Hope that helps :)

For those interested in looking slightly more into the characteristic function, it may be worth pointing out that the characteristic function is equal to the Fourier-transform (with the sign of the argument being reversed) of the probability distribution in question.

In my own experience teaching teaching probability theory to physicists and engineers, establishing this connection is often a good way of helping people build intuition for why characteristic functions are so useful, why they crop up everywhere in probability theory, and why we can extract so much useful information about a distribution by looking at the characteristic function (since this group of students tends to already be rather familiar with Fourier-transforms).

The variable n comes out of nowhere in theorem 3.3, and they do not refer to it in the proof itself as far as I can tell. Is this just an editing error (I think the formula 3.4 needs the variable n if f is multidimensional and we are integrating over R^n, but since f is in L^1(R) I'm not sure what it signifies. I am however worried that there's something I'm missing).

How is "The circumferance of an idealized circle divided by its diameter" not a finite expression of π? Saying something cannot be expressed finitely in an integer-based numeral system, and saying that it admits no finite representation are two radically different statements.

Despite it being a non-starter from a pragmatic standpoint, we could for instance easily imagine a novel numeral type that encodes the set S = {a + b·π where a and b are integers} (we can encode integers quite easily and all we need to reposesent such a number in silico is to encode a and b). Using such a numeral type, we are able to do exact arithmetic if our operations are restricted to addition and subtraction (and if we are content with fractional representation of numbers as being considered "exact", we can also do division and multiplication although we would have to work within the larger set S' = { (a + b·π) / (c + d·π) where, a, b, c, and d are integers and c·d ≠ 0} rather than within S).

Unrelated to the article in question, but using ℝ over 𝕽 for the reals is more of a modern development. If you read older articles and textbooks, many will use 𝕽 rather than the sleeker ℝ (most in my experience, but results will be heavily affected by your cut-off for 'old').

I don't have direct evidence for my speculations, but I presume the reason [fraktur](https://en.wikipedia.org/wiki/Fraktur) was more common in mathematics back in the days is largely down to articles having to be type-set using movable type. If you insisted on using ℝ over 𝕽, you were likely to make life considerably harder for your printer (which in turn meant higher printing costs), since they would be considerably more likely to have to cast new types. As printing was modernized and movable type was replaced by more flexible printing-technologies, this pragmatic reason for preferring one glyph over the other went away. Another explanation/contributing factor is that the switch seems to have occurred in tandem with the barycenter of mathematics switching away from continental Europe and towards the US in the post-WWII period (at least if we disregard Soviet mathematics which also flourished in this period, but which was largely published in Russian). The average American would probably be less familiar with 𝕽 and other fraktur glyphs than the average German.

Part of the issue is that after a while we tend to forget that the cases that turned out to be true were dismissed as conspiracy theories at the time. In recent memory for example, a lot of claims about the capabilities and application of the US signal intelligence apparatus abroad (especially in allied coununtries) were dismissed as conspiracy theories prior to Snowden. If you talk to a lot of people today, they will tend to remember it more as a "we kinda' always knew, but just didn't have confirmation" situation than a "I'm sure the NSA is doing _something_, but there is no way it would be this extensive" situation.

I suspect OP may have been going for a variation on the old "Programmer returns with zero eggs and 12 gallons of milk after having been asked to get one gallon of milk and if they have eggs to buy a dozen"-joke, but it falls flat in this instance since it relies on an interpretation bordering on deliberate misconstrual (i.e. applying the modifier "for each year of service" to the whole phrase "16 weeks plus two additional weeks" rather than just to the latter fragment "two additional weeks").

Without knowing the exact approval history in the US, I doubt that it is a lax approval process as much as it is an absence of better alternatives.

There are not really any known effective clinical interventions (be they in the form of medicine, therapy, or things like more exercise), and a 10% improvement is better than nothing (clinical therapy (notably Cognitive Behavioral Therapy) is generally believed to be more effective but not by that much and it is financially out of reach for many people who are able to afford medication).

The problem with treating depressed people is to get them to actually do the things that will help them. That can be incredibly difficult without medication.

The problem with talking about doing "things that will help them" is that we don't really have a lot of effective clinical interventions. Even the most common interventions (SSRIs, Congnitive Behavioral Therapy, physical exercise, etc.) are not really that effective at treating depression and have little to no effect in large parts of the affected population. That being said, these treatments _do_ work for some people so they should definitely not be dismissed out of hand (although it may in some cases be regression to the mean more than anything else, i.e. if you get better after a while on your own but have undergone treatments of one form or another you may erroneously believe your most recent treatment was effective).

For anyone who wants to get really into the weeds, here are all the articles in the sequence of discussion papers:

Main Paper (Jager and Leek): ttps://doi.org/10.1093/biostatistics/kxt007

Response papers:

  - Yoav Benjamini and Yotam Hechtlinger: https://doi.org/10.1093/biostatistics/kxt032

  - David R. Cox: https://doi.org/10.1093/biostatistics/kxt033

  - Andrew Gelman and Keith O'Rourke: https://doi.org/10.1093/biostatistics/kxt034

  - Steven N. Goodman: https://doi.org/10.1093/biostatistics/kxt035

  - John P. A. Ioannidis (the spicy response): https://doi.org/10.1093/biostatistics/kxt036

  - Martijn J. Schuemie, Patrick B. Ryan, Marc A. Suchard, Zach Shahn, and David Madigan: https://doi.org/10.1093/biostatistics/kxt037
Jaeger and Leeks' rejoinder to the responses: https://doi.org/10.1093/biostatistics/kxt038

edit: fixed some formatting and link to main paper

For anyone who likes academic drama (or who is interested in the underlying methodological disagreements among academic statisticians), it is worth pointing out that Jager and Leek's 2014 paper was a discussion paper, and that Ionnidas was one of the people invited to write a response to be published alongside with the original paper. He did (https://doi.org/10.1093/biostatistics/kxt036) and the response is extremely critical of Jaeger and Leeks' methodology and his contempt for the authors is not hard to read between the lines.

Given how common and wide-spread misattribution of code is on GitHub, I'd say there is a strong argument (moral rather than legal--I'm not an IP lawyer and will leave judgements regarding legal liability up to the professionals) that they can be held responsible for this mess exactly because it is such a well-known issue and that rolling out copilot without addressing this (most likely as you suggest by actually spending more resources on vetting projects and tidying up training data) amounts to gross negligence on the part of GitHub since there is good reason to believe this will exasperate this problem significantly.

As the linked article points out, practices can vary widely across fields (and in fields where preprints are available, maintaining authour anonymity in a double-blind setting furthermore relies on revieweres not having come across a preprint of the paper that is being reviewed which in turn can easily happen in smaller research-communities).

Yes, but not by "growing the pie"; MBAs increase profits in the sense that they reduce the share of profits going to workers (at least according to this paper; I'm not familiar with the rest of the literature on this subject).

From the abstract of the paper: "...Exploiting exogenous export demand shocks, we show that non-business managers share profits with their workers, whereas business managers do not. But consistent with our first set of results, these business managers show no greater ability to increase sales or profits in response to exporting opportunities..."

It doesn't involve cryptography, but mastodon has for at least a couple of years supported link-verification in profiles (it basically checks if a link back to your mastodon profile exists on a page linked on your profile), so a linking to a page that only you credibly control (say, a personal website) is the de-facto system of decentralized user-verification on mastodon.

Edit: supported since 2018 https://github.com/mastodon/mastodon/pull/8703

Dear Chess World 4 years ago

The fundamental asummetry at play here is that cheating is a lot easier than detecting cheaters in modern chess. As such, I'm not sure insisting on people shutting up unless they can provide ironclad evidence won't just lead to professional chess becoming rife with cheaters (which is presumably the "existential threat" to proffessional chess that Carlsen reffers to in his statement).

I admit that it is a trickly problem, and I agree that Carlsen's behavior here is not beyond reproach. Withdrawing from the tournament only after having lost a game makes the statement much less impactfull since one can not discount the possibility that he's just being a rather sore looser. I would personally have respected his decision much more if he had followed his impulse (again, referred to in his statement) to withdraw as soon as Niemann had been invited to the tournament in the first place.

Dear Chess World 4 years ago

Surely he means just that; i.e. Hans Niemann is known to have cheated in past games (these were online games and it must be noted that Niemann maintains that he has never cheated before or since the and never in an "over the board" tournament).

Another useful (to Google) aspect of hiring an ethical AI team is also that is allows Google to shut down ethical concerns raised in other parts of its own organization by going "we agree that ethics is an important, but please focus on your own job", i.e. an internal AI team can act as a sort of "controlled opposition". The fact that the ethical AI team is now being undermined within the company after trying to do their job and making clear that they are not content to be relegated to such a role is pretty consistent with this motive.

While I share your hesitation here, I think two points are important to keep in mind:

1. There is quite a difference between compulsory auditing (what the post you reply to refers to) and the government directly controlling industry.

2. In other industries this is quite commonplace and hasn't led to government takeover of industries (banking comes to mind. In their regulatory implementation on the Basel III accords developed in response to the 2008 financial crisis, both the UK and EU mandate government audits to ensure compliance with stress-testing and and leverage requirements; the US is also a signatory to these accords, but I am less familiar with their implementation into US law).

I'm not personally a huge fan of this approach, but I don't find the argument that government oversight is a slippery slope to totalitarianism that persuasive. In my opinion, a much a stronger critique of mandatory government audits is that they are often not that effective at preventing the negative outcomes they set out to prevent but still massively increase the legal complexity of operating in (or entering) a given industry without falling afoul of the law.

A big issue is also that power companies in much of Europe (don't know much about how things look in the US) do not have an infrastructure capable of handling the additional load if EV-charging starts replacing gas-stations. As such, many are not exactly enthusiastic about rolling out superchargers at an accelerated pace since this would only further strain their networks.

Bitcoin Is a Ponzi 6 years ago

Aren't you just describing an input-based theory of value in your first statement (as opposed to a Riccardian theory of value where price is determined by demand relative to supply)? I was under the impression that a labour theory of value is a rather specific kind of input theory since it would not treat energy as a legitimate input but instead require the value of energy be determined by the human labor required to produce that energy--i.e. the value of bitcoin should be determined by considering the amount of labour required to produce the energy required for mining plus the labor directly involved in running and the Bitcoin network. Consider a contrasting input-based theory of value which instead of using labour uses some finite resource (or collection thereof) like barrels of oil as its unit of value: here the value of Bitcoin would be determined by the number of barrels of oil required to produce the energy required to mine bitcoin. Even though it uses energy as a means of determining the value of bitcoin, this second theory of value does not involve considerations of labour at any stage.

Also (as an aside irrelevant to my above question), I think you greatly overestimate the number of people who subscribe to a labour theory of value. Personally, I don't recall encountering people who subscribed to it other than far-left shitposters on twitter (who are definitely a minority, although a rather vocal one).

I mean, isn't that (one of) the point(s) of the corporation as a legal entity: to allow people with capital assets and a joint enterprise in mind to coordinate their resources and act as a single entity for the purposes of pursuing that enterprise [edit: which of course includes purchasing labor in all but the smallest undertakings]. One of the original motivating theories behind the drive for labor unions organized within companies (rather than earlier forms of labor organization like guilds and professional societies which predate modern economies) was that the unification of capital interests within a company necessitated that labor similarly unify, since the alternative would be that the individual seller of labour would lack any kind of bargaining power when negotiating with their employer.

Unless we abolish corporations as a concept, how exactly do you effectively propose that we "outlaw all cartelization" in a way that isn't just outlawing unions while keeping capital interests unified?

tl;dr: unions complement the inherrent concentration of capital interests in modern ecconomies and are no more inherently like cartels than the corporate form itself.