HN user

jpeloquin

568 karma
Posts4
Comments239
View on HN

Post-mandate, I've been submitting to closed access journals and getting OA on the side for free due to the mandate. Pre-mandate, I only submitted to paid OA journals, and paid ~ $3k each time for it.

The article claims the solution is "every government grant should stipulate that the research it supports can’t be published in a for-profit journal. That’s it! If the public paid for it, it shouldn’t be paywalled." That's an equivocation fallacy. Whether a for-profit journal publishes the work at some point is orthogonal to whether it is available un-paywalled, which it now must be.

You say that publishers replaced subscription fees with APC charges, but I haven't seen this happening when I've submitted papers recently. Journals need new submissions or they lose mindshare. Authors are price-sensitive and will shop around. Starting a new journal isn't that hard (it can been done as a side project) so high margins will likely be undercut. I have no idea why the author chose to pay a $12k APC; they probably didn't need to. Finally, closed-access journals will have residual subscription income from their closed-access archives for many decades; if the author wants to kill that income stream off, their proposed solution will not do it. So while I agree with the article's condemnation of the publishers, who are certainly no friends of science, I think it's wildly off-base on pretty much every other point.

Author-pays APCs are even potentially a good thing as long as they aren't much higher than the cost of publication. Universal APCs would provide some pressure against publishing many low-value papers that aren't really worth the time it takes to read them. The paper spam is kind of getting out of control.

Evaluating a function using a densely spaced grid and plotting it does work. This is brute-force search. You will see the global minima immediately in the way you describe, provided your grid is dense enough to capture all local variation.

It's just that when the function is implemented on the computer, evaluating so many points takes a long time, and using a more sophisticated optimization algorithm that exploits information like the gradient is almost always faster. In physical reality all the points already exist, so if they can be observed cheaply the brute force approach works well.

Edit: Your question was good. Asking superficially-naive questions like that is often a fruitful starting point for coming up with new tricks to solve seemingly-intractable problems.

From main text:

Discussions with different stakeholders suggest that many currently perceive systematic fraudulent science as something that occurs only in the periphery of the “real” scientific enterprise, that is, outside OECD countries. Accumulating evidence shows that systematic production of low quality and fraudulent science can occur anywhere.

From supplement (section about the output of the "ARDA" paper mill):

We obtained 20,638 documents and were able to impute country of authorship for 13,288 documents (64.4%). Of these documents, more than half were solely from India (26.4%), Iraq (19.3%), or Indonesia (12.2%).

The identity and reputation of the authors, and the publication venue, is (for now) still a strong signal when evaluating the credibility of an article.

The article is spot-on though in that there is a real risk of paper mills infecting formerly reliable journals, and this is not helped by the publishers' commercialism. For example, it used to be easy to ignore Hindawi journals (they are characteristically low quality); then Wiley started publishing them under its own brand. The good is now mixed with the bad under the same label. Practicing scientists can fall back on whether they know the authors personally but that doesn't really help non-practicing professionals or the general public.

Industry research is generally R&D (applied science, engineering research), not basic research (basic science). Not to disparage either; both are needed, but they are quite different and a person may be suited to one but not the other. It can be hard for someone looking for work to determine where an organization's focus is, as an outsider.

Multiple comparisons and sequential hypothesis testing / early stopping aren't the same problem. There might be a way to wrangle an F test into a sequential hypothesis testing approach, but it's not obvious (to me anyway) how one would do so. In multiple comparisons each additional comparison introduces a new group with independent data; in sequential hypothesis testing each successive test adds a small amount of additional data to each group so all results are conditional. Could you elaborate or provide a link?

Publications with public funding have already escaped the paywall, partially as of 2013 and completely as of this year:

https://par.nsf.gov/

https://pmc.ncbi.nlm.nih.gov/

https://ospo.gwu.edu/overview-us-policy-open-access-and-open...

https://www.nih.gov/about-nih/who-we-are/nih-director/statem...

https://www.coalition-s.org/plan_s_principles/

The intent of the Bayh-Dole Act was to deal with a perceived problem of government-owned patents being investor-unfriendly. At the time the government would only grant non-exclusive licenses, and investors generally want exclusivity. That may have been the actual problem, moreso than who owned the patent. On the other hand, giving the actual inventors an incentive to commercialize their work should increase their productivity and the chance that the inventions actually get used.

Once something has a predictable ROI (can be productized and sold), profit seekers will find a way. The role of publicly funded research is to get ideas that are not immediately profitable to the stage that investors can take over. Publicly funded research also supports investor-funded R&D by educating their future work force.

The provided examples do not clearly support the idea that industry can compensate for a decrease in government-funded basic research. Bell Labs was the product of government action (antitrust enforcement), not a voluntary creation. The others are R&D (product development) organizations, not research organizations. Of those listed, Xerox PARC is the most significant, but from the profit-seeking perspective it's more of a cautionary tale since it primarily benefited Xerox's competitors. And Hinton seems to have received government support; his backpropagation paper at least credits ONR. As I understand it, the overall deep learning story is that basic research, including government-funded research, laid theoretical groundwork that capital investment was later able to scale commercially once video games drove development of the necessary hardware.

The median sample size of the studies subjected to replication was n = 5 specimens (https://osf.io/atkd7). Probably because only protocols with an estimated cost less than BRL 5,000 (around USD 1,300 at the time) per replication were included. So it's not surprising that only ~ 60% of the original biomechemical assays' point estimates were in the replicates' 95% prediction interval. The mouse maze anxiety test (~ 10%) seems to be dragging down the average. n = 5 just doesn't give reliable estimates, especially in rodent psychology.

Prioritize work by ROI and alignment with the institution's mission, communicate the prioritization to relevant decision-makers in your management chain, and actually follow through on it unless coerced otherwise.

Edit: There's no guarantee this will work out positively—nothing is guaranteed—but it's worth considering if the alternative is giving up / changing careers.

Well, obviously the part-time thing will bring a reduction in my institutional teaching and admin duties. I have to say there is uncertainty about how much relief will arise in practice

As someone who has tried something similar, institutional bureaucracy expands to fill all available time. People engaging in bureaucratic empire-building will still happily consume your personal unpaid time. And splitting attention between multiple lines of work creates some legitimate additional overhead, which doesn't help.

I'm not sure what the winning strategy is. I think it is necessary to either get out entirely or play the bureaucrats' "system-game" to some degree, but not on their terms and not fairly. When the bureaucracy demands useless work, maximize their costs and minimize yours. Many academics constitutively cannot make themselves do a lazy, poor job, but is a useful skill to deploy defensively so that you can fulfill your education and research responsibilities. Often you'll find that the bureaucracy only cares about the superficial appearance of compliance; the actual actions performed are irrelevant to them. Shift responsibility to some other part of the bureaucracy and use LLMs to generate boilerplate. If the bureaucracy never attempts to punish you in any way, that may indicate you're being more compliant than necessary. This approach is safest if your retirement plan is fully funded and you don't truly need to keep the job. It is also helpful if at some part of what you do is visibly important to someone who does have power; this helps deflect consequences when you accidentally step a over a line. Everything depends on context and execution though; I hope the part-time approach works out for you.

Securing the funding often means writing a research plan that is interesting and convincing enough to be selected for funding in a competitive review process (80–97% rejection rate). The plan usually represents a substantial intellectual contribution that serves as the foundation for derived papers. As long as the people who join the project later don't freeze the planner out of the paper writing process, they'll usually meet all authorship criteria.

Well, the authors' byline says they're from Northeastern University, Shenyang 110169, China and Faculty of Management and Economics, Dalian University of Technology, Dalian 116024, China. I've heard that Chinese universities sometimes have explicit publication quotas or offer cash bonuses for publications. Having never worked at a university in China, I cannot personally verify that though.

MDPI also spams anyone it thinks might be willing to submit a paper or serve as an editor, and does not seem to care much about the quality of the submissions it receives.

Right, "sharing" here must mean DNA that was cloned from the same ancestral DNA strand, not merely that it shares the same informational content. I got lost in the analogies that frame things in terms of what's "better" for the organism and lost sight of this.

The most important thing from the perspective of replication of a DNA strand is the number of copies of DNA passed to the next generation, and future generations, right? Which would be 0.75 * (mean marginal increase in next-generation sisters) + 0.5 * (mean # offspring). The probability that these next-generation individuals actually get to reproduce in turn would also factor in somewhere.

What's also interesting is that if we take the point of view of the queen (through whom the altruistic genes must pass), the queen's reproductive strategy is relatively few children + hordes of sterile helpers + killing her own sisters. So are we talking about a fitness advantage of altruistic traits (maximizing # sisters), or a fitness advantage from selfish traits [maximizing P(fertile child survival) I guess, since # children is small] that produce hordes of sterile helpers?

Edit: Circling back to the organism perspective, in the sense of "I would gladly give up my life for two brothers or eight cousins.", how many bees is it worth giving up one's own life for in that specific sense? We do have a common ancestor after all and thus a non-zero R-factor.

Based on the information I found, the % difference between two randoms humans in terms of base pairs (including non-coding DNA) is even less than the difference in terms of genes, so the distinction does not materially alter the discussion. Also the article framed its explanation in terms of genes, not base pair sequence.

"Between any two humans, the amount of genetic variation—biochemical individuality—is about .1 percent." https://www.ncbi.nlm.nih.gov/books/NBK20363/

https://book.bionumbers.org/how-genetically-similar-are-two-...

Forensic comparisons are mostly about comparing the number of short tandem repeats at handful of loci, a very small part of the the whole genome.

If you have any information that indicates the DNA similarity between people is less than 98–99% I would love to hear it. I have not personally analyzed the sequences from the 1000 genome project to check, and am relying on summaries written by other people.

The concept of indirect fitness must be more complicated than explained here. The article explains it as a worker bee sharing 75% of her genes with her sisters, but only 50% with a child, so there is selection pressure for workers to be sterile and self-sacrificing. But few genes actually differ between individuals, so the percentages are much higher. E.g., I share ~ 99% of my genes with each one of you reading this. Assuming honey bees' genetic variation is not much more extreme than human variation, we're talking about 99.5% vs. 99.75% sharing, which sounds more like an explanation of why altruism would be preferred in general rather than uniquely affecting bees.

The article does eventually circle around to acknowledge this, but it's easy to miss and very underdeveloped compared to the discussion of kin selection: "So why do bees die when they sting you? Perhaps because they're disposable parts of a larger super-organism which has evolved by multi-level selection."

Shifting the topic from research misconduct to good laboratory practices, I don't really understand how someone would forget to take pictures of their gels often enough that they would feel it necessary to fake data. (I think you're recounting something you saw someone else do, so this isn't criticizing you.) The only reason to run the experiment to collect data. If there's no data in hand, why would they think the experiment was done? Also, they should be working from a written protocol or a short-form checklist so each item can be ticked off as it is completed. And they should record where they put their data and other research materials in their lab notebook, and copy any work (data or otherwise) to a file server or other redundant storage, before leaving for the day. So much has to go wrong to get to research misconduct and fraud from the starting point of a little forgetfulness.

I mean, I've seen people deliberately choose to discard their data and keep no notes, even when I offered to give them a flash drive with their data on it, so I understand that this sort of thing happens. It's still senseless.

Oh Shit, Git? 2 years ago

Recipes like these aren't useless, but yes, they really need to be prefixed with whether they expect to start from a clean work tree and empty staging area. Or describe what they'll do to uncommitted changes, both staged & unstaged. Otherwise they pose a substantial risk of making the problem worse.

Or, better yet, the gay satanic-panic currently gripping half the country, and the insane culture war being waged around it. You can't actually believe that all those people who have strong opinions about it have been somehow personally wronged by homosexuals.

Or the satanic panic over Dungeons & Dragons in the 1980s. One of the cops ("school resource officers") in the middle school I went to still believed in that nonsense and it was the early 2000s by that point.

If anyone is curious, the following seems to work:

0. Install the Tree Style Tab extension (or whatever vertical tabs extension you prefer).

1. Enable userChrome.css: set toolkit.legacyUserProfileCustomizations.stylesheets=true in about:config.

2. Set browser.tabs.inTitlebar=0 in about:config so the title bar buttons (and, on some OS's, the title bar itself) remain visible.

3. Create =chrome/userChrome.css= in your Firefox profile folder and write the following to it:

  @namespace url("http://www.mozilla.org/keymaster/gatekeeper/there.is.only.xul");

  /* hides the native tabs */
  #TabsToolbar {
      visibility: collapse;
  }

Is it considered unethical to do medical experiments on yourself without any oversight like you would find in a typical human subject trial?

Note there are different contexts at play here. When someone says "ethics" in a scientific context, it may encompass scientific integrity, avoidance of questionable research practices, reproducibility, etc., as well as medical and moral ethics. The speaker may not even be fully aware of these distinctions, since the subject is often taught with a rule-based perspective.

Experimentation on oneself is often _scientifically_ unethical (i.e., when done with the intent to make a scientific discovery) because:

1. The result is often too contaminated by experimental integrity issues to have scientific value. As another comment in this thread notes: "sample size of 1, confirmation bias, amped-up placebo effect, lack of oversight, conflict of interest when the patient is the investigator". Lack of oversight means no one is checking the validity of your work, it's not a permission thing. Every issue that is blamed for the so-called reproducibility crisis is worse.

2. Due to publication pressure, abandoning the cultural prohibition against self-experimentation amounts to pressuring everyone to self-experiment to grow their CV by a few quick N = 1 studies, or do something risky when their career flags. Obviously, oversight to ensure that self-experimentation proceeds only in cases of terminal disease mitigates this concern.

In practice, journal editors currently provide oversight addressing point #2, which is why work like what we're discussing here still gets published. See also Karen Wetterhahn's valuable documentation of her (accidental) dimethylmercury poisoning (https://en.wikipedia.org/wiki/Karen_Wetterhahn).

Experimentation on oneself in an attempt to cure your own illness by any means at your disposal, provided you do not harm others, is not _morally_ unethical IMO. It just rarely has a scientific role.

According to the original account, the pencil/pen thing wasn't about an audit trail, and both the IRB and hospital admin were equally silly.

IRREGULARITY #3: Signatures are traditionally in pen. But we said our patients would sign in pencil. Why?

Well, because psychiatric patients aren’t allowed to have pens in case they stab themselves with them. I don’t get why stabbing yourself with a pencil is any less of a problem, but the rules are the rules. We asked the hospital administration for a one-time exemption, to let our patients have pens just long enough to sign the consent form. Hospital administration said absolutely not, and they didn’t care if this sabotaged our entire study, it was pencil or nothing.

https://slatestarcodex.com/2017/08/29/my-irb-nightmare/

In my opinion, archive the data that was actually gathered and the code's intermediate & final outputs. Write the code clearly enough that what it did can be understood by reading it alone, since with pervasive software churn it won't be runnable as-is forever. As a bonus, this approach works even when some steps are manual processes.