HN user

riskassessment

124 karma
Posts0
Comments37
View on HN
No posts found.

Fair, the important distinction is agent-agnostic rather than open-source. There are other risks to using a closed source editor but those are mostly orthogonal to this discussion.

I was surprised people were so willing to jump to closed source IDEs just for access to coding agents. The trade-off you pay for tight integration between the IDE and the coding agent is lock-in because the barrier to switching IDEs is nontrivial.

Your coding environment stands a lower chance of disruption when you use an open source IDE with a CLI agent. Yes it's slightly annoying to separate the agent from the IDE but the benefit is that it's much easier to switch between Claude Code, Codex, Gemini CLI (now antigravity CLI), etc which means you can more easily benefit from pricing and coding performance differences which seem to change monthly.

Stealthily degrade the model or stealthily constrain the model with a tighter harness? These coding tools like Claude Code were created to overcome the shortcomings of last year's models. Models have gotten better but the harnesses have not been rebuilt from scratch to reflect improved planning and tool use inherent to newer models.

I do wonder how much all the engineering put into these coding tools may actually in some cases degrade coding performance relative to simpler instructions and terminal access. Not to mention that the monthly subscription pricing structure incentivizes building the harness to reduce token use. How much of that token efficiency is to the benefit of the user? Someone needs to be doing research comparing e.g. Claude Code vs generic code assist via API access with some minimal tooling and instructions.

NaN Is Weird 4 months ago

Nor is that inequality an oddity at all. If you were to think NaN should equal NaN, that thought would probably stem from the belief that NaN is a singular entity which is a misunderstanding of its purpose. NaN rather signifies a specific number that is not representable as a floating point. Two specific numbers that cannot be represented are not necessarily equal because they may have resulted from different calculations!

I'll add that, if I recall correctly, in R, the statement NaN == NaN evaluates to NA which basicall means "it is not known whether these numbers equal each other" which is a more reasonable result than False.

They teach us Scientific Realism in school.

I'd argue the opposite is true for anyone who has studied statistics which is largely built on Instrumentalism (think George Box: 'All models are wrong, but some are useful') and Popperian falsification (Null Hypothesis testing). We are absolutely taught to treat models as predictive tools rather than metaphysical truths.

I don't understand this reasoning. Randomizing people to AI vs standard of care is expensive and risky. Checking whether the AI can pass hypothetical scenarios seems like a perfectly reasonable approach to researching the safety of these models before running a clinical trial.

I was expecting a system like Leibniz notation, Boolean Algebra, Begriffsschrift, or the notation system in Principia Mathematica

R is perhaps the closest, because it has data.frame as a 'first class citizen', but most people don't seem to use it, and use e.g. tibbles from dplyr instead.

Everyone in R uses data.frame because tibble (and data.table) inherits from data.frame. This means that "first class" (base R) functions work directly on tibble/data.table. It also makes it trivial to convert between tibble, data.table, and data.frames.

Google Antigravity 8 months ago

html

Would be willing to bet this is the issue. Adding html files to context for gemini models results in a ton of token use.

Gptel has been working great for me. I'd be interested in checking this out but I only have so much time to set up and test new tools. What features would make it worthwhile to switch from gptel?

Gemini 2.5 Deep Think 12 months ago

I'd be interested in tests involving tasks with large amounts of context. Parallel thinking could conceivably useful for a variety of specific problem types. Having more context than any specific chain of thought can reasonably attend to might be one of them.

hardly forever. Given the age of the company you're citing, they can only estimate retention out to 1 year.

I said I was skeptical of there being a precise pattern of rapid aging. I never said I was skeptical that rapid/non-linear aging can occur. If you did experience rapid aging in the way the paper measured this from 38-40 that is more evidence in support of my point that there is some broad random distribution of when rapid aging occurs and this paper and blog post overintepret the data to mean rapid aging occurs precisely in your mid-forties and at 60.

I read the paper before I made my original comment. They fit a clustering algorithm and then hand waved at intepreting the clusters. 'Omics papers get away with a lot of hand waving. Yeah they did some peak detection and found peaks, but you are going to find peaks in a random walk.

They didn't test the theory that rapid aging occurs at those two specific time points in an independent hold out set.

Most importantly even if these peaks exist this paper does not prove they are biological. They could correspond to common socially driven changes in behavior

Sure but I think it's an oversimplification to say that universities are unfocused and that this lack of focus is a problem because 1) the alternative already exists (smaller colleges are typically focused primarily on education with less research and less sports) 2) Universities doing a lot of things (including offering undergrads connections to research labs) is one of the main reasons that they are an attractive option relative to colleges.

Probability of dying in a random accident is the probability of a random accident occuring times the probability of dying from a random accident conditional on one occurring. I am not convinced that fitness would be entirely unassociated with both of these probabilities, particularly the latter. Meaning that the extent to which this outcome is a good negative control is overestimated.

The value of good editorial staff may very well be less important in some fields. In the biomedical field subject-matter expertise still matters a lot in terms of discerning good research from seemingly good research. Which also means that prestige journals won't make research any more groundbreaking but if functioning properly should enrich for papers that are less likely to be junk science. I'd also argue that generic journals like Nature and Science are so unfocused that their staff probably provides little to no additional expertise and they rely entirely on peer review. Whereas staff at more specialized journals with lower but still very respectable impact factors are probably doing more informed work to select for quality science.

Moving forward the removal of the embargo. But my point is that access to federally funded science was free prior to anyone coming up with a plan to remove the embargo. You just needed to wait up to a year before a paper was put on pubmed central. This removal of the embargo is hardly a meaningful change in terms of access but one that erodes the institutions that ensure peer review happens. It is easy to say peer review is largely based on volunteers, but if journals ceased to exist tomorrow I doubt anyone here would volunteer to do the task of what the journals do now. At least you can put peer reviewer on your academic CV. The paid journal staff do much less glamorous work but still serve a role in keeping peer review running.

Infrastucture is not the same as software. I was mostly referring to the human infrastructure (although editorialmanager is not free and until someone makes an open source alternative the supscription fees do support that license). And I would argue that the existence of a small number of prestige journals with scientific staff makes the entire system worth it, even if it means we have to deal with the existence of Elsevier and the like.

I specifically said journals subscription fees support peer review infrastructure. Yes peer reviewers are unpaid but peer review also would not exist in anything resembling its current form in the absence of journal staff moving papers through the peer review system. Associate/deputy editors are unpaid but the main editor of the journal is often paid and does provide scientific oversight and review, particularly at the margin of acceptance/rejection. The main editor of course is also responsible for recruiting associate editors who in turn are responsible for finding appropriate peer reviewers, so having a good editor who can recruit and maintain quality deputy/associate editors is key. Some journals even have staff scientific reviewers which act as a check on the occasional oversights of unpaid peer review.

My reading of this press release is that they are just removing the 12 month embargo period before the already mandated free-access (untypeset) versions of grant-supported manuscripts can go on pubmed central. The prior policy of a 12 month embargo period allowed publishers to have a small value add over the free version. This value add justifies subscription fees which support, among other things, infrastructure necessary to support peer review and possibly some in-house staff scientific editing and review. I do wonder whether it is worth it to make all papers available immediately if indirectly may make peer review even less supported than it is now.