HN user

acidbaseextract

576 karma
Posts4
Comments154
View on HN

The normal way? You can implement whatever kind of index you like — b-tree index, bitmap index, hash index are all useful and conceptually simple if you're familiar with the backing data structures.

For example, if you want to index a "foreign key" id stored in each "record" in a JSON array of objects, you build a hash table from the FK id values to the JSON array indices of the objects that have that id. It can be as stupid simple as an `fk_index = defaultdict(set)` somewhere in your program, to use a Pythonism.

Now when someone wants JSON objects in that array matching a given FK id, they can just O(1) look in the index to know the position of records that match. Much better than an O(N) scan of every item in the array.

Of course you have to to maintain the index as writes to the JSON happen, but that's not bad once you understand how things work. No real secret sauce.

I don't regret putting it off until my 40s because it let me find a person I want to have a family with

Thank you for including this tidbit. I'm a man in my early 30s interested in having kids and recently discovered that I've poorly vetted my 5+ year SO's interest in having kids. Her "yeah, I'm hypothetically interested" has become "hard not interested".

The requirement to break up with her if I want kids is deeply painful, but the real source of my dread is the feeling that I won't then be able to find someone good in time. I'm very glad you found someone good to have a family with.

I'll give you my answers to your questions as someone writing production JS/TS. Answering your questions is precisely illustrative of why hooks issues are hard to catch.

unexpected null/undefined values

This is an issue and has led to bugs. Switching to TS and making the compiler flag them fixed it.

mistyping variable or attribute names

Not an issue. Code obviously breaks if you have incorrect names.

using the wrong number of equal signs

This is an issue but causes few bugs. Linting generally catches it, though cute boolean punning still bites us.

failing to catch errors and handle rejected promises, etc.

This is an issue and has led to bugs.

Pretty much anything where there's some implicit details that the compiler or linters can't reason about programmers find a way to get wrong. One thing I like about the hooks linter setup is that what it encourages you to do by default will prevent most bugs, only lead to potential performance issues, unnecessary rerenders, unnecessary refetches.

I'm curious if you've used useDeepCompareEffect, the use-deep-compare-effect npm package? I've found that it is pretty reasonable foolproofing for many of these identity questions. I'm well aware of Dan Abramov's objections to the deep equality checking [1] but I still find it a bit easier for me and other devs to reason about when doing things like data fetching.

[1] https://twitter.com/dan_abramov/status/1104414469629898754

Wut? I just make websites man.

I don't want to be dismissive, but I hate deploying my applications for this reason. I'm an application developer. I'm not averse to infrastructure as code, and containerization, and I'm happy to do ops for my preferred stack. But I can't learn all this stuff too.

Gassée also started Be Inc which created BeOS. Beyond being great in and of itself, BeOS was also slated to be the successor to Mac OS 9. If you're an operating system or file system nerd, you will find it very worth looking into BeOS. Gassée overplayed his hand and Apple's acquisition of Be Inc fell through.

Apple ultimately went with the NeXT / Steve Jobs combo, quite wisely, but for a long time there was a whole gang of BeOS fanboys lamenting that decision.

I've always had a soft spot in my heart for PHP style "raw query in the template". Even as we've moved away from that I feel like GraphQL recapitulates a lot of the reasons why having queries tightly bound to views is convenient.

It feels like "put the template in the query" is a great Uno reverse to "put the query in the template" and the associated issues it brings, but I'll have to think on the consequences more. Thought provoking article!

It's been a long time since I was there, but I thought there were plenty of 15,000 call site refactorings done in a single final CL. Not that the Linux kernel should do the same!

No shenanigans for any library-ish code with few dependencies that targets relatively modern Python.

I've had one or two fundamental version conflicts with a 5+ year old application with 100+ dependencies and a decent amount of legacy stuff. They were a pain in the ass, and the sdispater's stance on not allowing overrides is a pain in the ass. We ended up forking the upstream libraries to resolve the version conflict.

With all of that, poetry is amazing and a huge step forward. I'd advocate it wholeheartedly.

I think the GitHub GraphQL API docs are particularly badly organized and basically show the API as a bucket of stuff, rather than as CRUD entities (which is how it's actually organized). I don't think this a GraphQL issue.

FTA, though the misspellings are frustrating:

In a GraphQL API, tools such as Dataloader allow you to batch and cache database calls. But in some cases, even this [isn't] enough and the only solution is to block queries by calculating a maximum execution cost or query [depth]. And any of these solutions will depend on the library you’re using.

ClickHouse, Inc. 5 years ago

As a sidenote, I saw your talk on Clickhouse to the CMU database group [1] back when and was extremely impressed with your deep technical knowledge yet down-to-earth presentation. Still haven't had an opportunity to use Clickhouse for production work, but would welcome it.

[1] https://www.youtube.com/watch?v=fGG9dApIhDU

Been working with React since 2014. That is the damndest thing ever, especially that it doesn't even throw a warning.

It seems like an identity confusion issue where the VDOM diff is ambiguous, and React resolves it in the "wrong" way. Adding keys to each `LabeledInput` resolves the issue, but I'm surprised that the runtime doesn't complain when you create the inputs without keys.

I wonder if this is why the checkbox that's checked moves, but stays in the same relative position (the second checkbox in the list): https://medium.com/@ryardley/react-hooks-not-magic-just-arra... or if it's just the ambiguous VDOM diff causing that.

The time series feature (TSFeature) extraction module in Kats can produce 65 features with clear statistical definitions, which can be incorporated in most machine learning (ML) models...

I'd be curious about the performance of these. A time series featurization library I've liked the look of but haven't used for real is catch22: https://github.com/chlubba/catch22

In particular I like catch22's methodology:

catch22 is a collection of 22 time-series [that are] are a high-performing subset of the over 7000 features in hctsa. Features were selected based on their classification performance across a collection of 93 real-world time-series classification problems...

I overall agree with the sentiment on delivery and needing to deal with a variety of issues, but one nit:

If it were my business or my team I would want an ML Engineer to relentlessly knock down any barriers to getting models into production.

My favorite is when they knock down the question of whether a certain project even needs ML to focus on getting ML into production.

Many teams fail at ML because there was an essential task that nobody wanted to do that didn't get done.

Many teams fail at ML because there was a task that didn't need ML (or at least anything more than a linear model) that is made opaque to anyone but an ML engineer after they implement it.

there is zero evidence of a lab leak besides circumstantial.

First, there are credible, credentialed virologists saying that the lab leak hypothesis has not been ruled out, and that it has been inadequately investigated [1].

Second, there are real anomalies in the Covid genome [2] that seem unlikely to have occurred naturally:

however, several characteristics of SARS-CoV-2 taken together are not easily explained by a natural zoonotic origin hypothesis. These include a low rate of evolution in the early phase of transmission; the lack of evidence for recombination events; a high pre-existing binding to human angiotensin-converting enzyme 2 (ACE2); a novel furin cleavage site (FCS) insert; a flat ganglioside-binding domain (GBD) of the spike protein which conflicts with host evasion survival patterns exhibited by other coronaviruses; and high human and mouse peptide mimicry.

In particular, the furin cleavage site is extremely interesting because it's exactly the type of genetic manipulation done in gain of function research that was ongoing at the Wuhan Institute of Virology. More on the FCS:

Because the presence and coding sequence of a FCS is important for pathogenesis, host range, and cell tropism (Nagai et al. 1993; Millet et al. 2015), the addition of a FCS into viruses has been an active area of gain-of-function research. A FCS can be easily inserted using seamless technology (Yount et al. 2002; Sirotkin and Sirotkin 2020) without any need for cell passage, as previously performed in experiments on virulence and host tropism (Cheng et al. 2019). Insertions to change the properties of SARS-r CoV viruses are documented by Ren et al. (2008) and Wang et al. (2008). Considering that natural mutations have a very low probability to result in a stretch of 12 amino acids coding for an optimized FCS without any known intermediate form in Sarbecovirus, an artificial insertion of the FCS in SARS-CoV-2 may provide a more parsimonious explanation for its presence than natural evolution.

In summary, the FCS confers SARS-CoV-2 enhanced human pathogenicity and has never been identified in another Sarbecovirus. At the same time, FCSs have been routinely inserted into coronaviruses in gain-of-function experiments, and we provide a hypothesis through which the specific amino acid sequence of SARS-CoV-2′s FCS may have been generated through cell culture.

[1] https://science.sciencemag.org/content/372/6543/694.1

[2] https://link.springer.com/article/10.1007/s10311-021-01211-0