HN user

trombonechamp

264 karma

http://maxshinnpotential.com

Posts2
Comments71
View on HN

This is something a bit different, the effect in the paper comes from the eigenvectors of the data matrix. I'm no expert on kernel density estimation, but since it is basically just convolution, I would guess the effect you are describing comes from multiplication in Fourier space.

Maybe a nitpick, but Aaron Swartz was probably quoting Richard Hamming:

And I started asking, "What are the important problems of your field?" And after a week or so, "What important problems are you working on?" And after some more time I came in one day and said, "If what you are doing is not important, and if you don't think it is going to lead to something important, why are you at Bell Labs working on it?" ... If you do not work on an important problem, it's unlikely you'll do important work.

It is the main idea behind his lecture/essay "You and your Research", which is worth reading: https://www.cs.virginia.edu/~robins/YouAndYourResearch.pdf or watching: https://www.youtube.com/watch?v=a1zDuOPkMSw

I did my PhD on this topic.

One classic result from biophysics is that, if all of your decisions have a fixed level of difficulty, then what the author suggests is (mathematically) proven to be suboptimal: if a decision feels more difficult, you actually should spend more time on that decision. (Keywords to search for: drift diffusion model[2], sequential probability ratio test[3])

The technical term for what the author is suggesting one should do (not spending so much time on decisions with equal outcomes) is an "urgency signal", or just "urgency" for short. If you are using an urgency signal, then you will spend less time on difficult decisions (i.e. ones with near equal outcomes) than you would without an urgency signal, but still more than you would easy decisions. (For easy decisions, you spend approximately the same amount of time regardless of whether you have an urgency signal.) In the extreme case, for an infinitely strong urgency signal, you will spend equal time on both easy and difficult decisions. (See Paul Cisek's work, e.g., [1].) Conceptually speaking, you need time to detect that the current decision has near-equal outcomes.

It was only recently (mathematically) proven that if you have different levels of difficulty in your decisions, it is optimal to use an urgency signal [4,5]. So since most sequences of decisions aren't all equally difficult (as in the study referenced here), in practice, people will use an urgency signal in decision-making.

Of course, both the desired accuracy and the urgency signal depend on how the decision is being evaluated: if your goal is to make an accurate decision (e.g. buying a house), then you will have a weaker urgency signal and require more evidence before you make a decision. By contrast, if you prioritise decision speed, you will show more urgency signal and require less evidence to make a choice. (The technical term is "speed-accuracy tradeoff".)

Current work suggests that when you are making some decision for the first time, you will not use an urgency signal, but as you become an expert at making those decisions, you will gradually develop an urgency signal [6]. This makes sense conceptually: if you know approximately how hard you can expect these decisions to be, you will recognise the situation where they have approximately the same utility and adjust your strategy accordingly.

[1] https://www.jneurosci.org/content/29/37/11560

[2] https://en.wikipedia.org/wiki/Two-alternative_forced_choice#...

[3] https://en.wikipedia.org/wiki/Sequential_probability_ratio_t...

[4] https://psycnet.apa.org/fulltext/2017-16730-001.pdf

[5] https://link.springer.com/article/10.3758/s13423-017-1340-6

[6] https://www.sciencedirect.com/science/article/abs/pii/S00100...

Darktable 4.0 4 years ago

The name "Darktable" is a play on Adobe Lightroom, since "lightroom" is a combination of "light table" and "darkroom", both concepts from film photography.

FYI, similar technology is already mature and commercially available, e.g., through Inscopix (no affiliation): https://www.inscopix.com/nvista Looks like the innovation here is getting reliable and lightweight two-photon recordings (allowing more cells and multiplane acquisitions) while the animal is moving around. Still very cool!

And responding to the other commenters, no, this kind of technology is definitely not (and never will be) a toy...

Once I needed a single LAPACK function (tridiagonal matrix multiplication) in a C Python extension. I spent an eternity trying to link to LAPACK reliably across platforms until giving up and deciding to implement it myself.

My implementation ended up being about 10 lines of code and ran slightly faster (!!!) than the LAPACK function. (I think this is because LAPACK provided numerical stability for special cases which didn't apply to my use case.)

Oh wow, this is a great idea. How do you deal with the lossy compression? There must be a lossless codec which uses the redundancy better than deflate?

Disclaimer: georgewfraser is my academic cousin.

Neuroscience research is huge and moving fast. Researchers have trouble keeping up with the bleeding edge in their sub-sub-field, let alone neuroscience as a whole. Conferences are one major way that most of us find out about new work.

Pre-covid, conferences were pretty exclusive and unwelcoming to non-researchers. However, now that conferences are all online due to covid, one thing I would recommend is going through the videos of old conferences to find something you are interested in. Then you can pause the video and look up any words or concepts which are unfamiliar. It may take a while to get through a talk, but you'll learn a lot this way. For example, Neuromatch is the big "mostly computational neuroscience" online conference: https://www.youtube.com/channel/UCcBKrxkfNv04R9PXLovjf5w/vid... Another formerly-in-person conference is Cosyne: https://www.youtube.com/channel/UCzOTbZTHTubFNjANAR33AAg/vid...

Heck, registration for Neuromatch is cheap so you can attend this year's conference if you want to: https://conference.neuromatch.io/

One of the major reasons I hold the medical and bio community is such low esteem - they don't know what they don't know yet act as if they do.

On one hand, this is due to poor science journalism. The media loves nutritional research because telling people what they are doing wrong generates clicks. It is very easy to sensationalize. "A small correlation between celery and colon cancer in a specific population from a specific region" (or whatever) in a scientific paper becomes "Celery slowly killing you, Harvard researchers prove".

One the other hand, this is how science works. We do experiments, get the results of those experiments, and try to make the best use of the data we have. Usually, the truth is much more complicated, but there are always bigger and more complicated experiments which need to be done to understand this more complicated truth. This happens all the time in all fields of science, but isn't as visible to the broader public because most research isn't all that relevant or interesting to non-scientists.

It is REALLY hard to design good (ethical) experiments in nutrition. The scientific consensus now is very different than what it will be in 10 years, because we will have more data and higher quality data. This will allow researchers to continue to incorporate more complexities and nonlinear effects. So you're right in that "linear thinking" is not ideal, but a linear model is better than no model.

There are ways we can measure perseveration in the lab, and so yes, it can be higher or lower in different people. If you want to look into it more, a popular way to measure it is the Wisconsin Card Sort Test (WCST), or the version for children, the Dimension Change Card Sort task (DCCS). However, like all cognitive tests, these are imperfect measurements of perseveration - they may be indicators of higher or lower perseveration, but scores may be influenced by a multitude of other factors as well. Likewise, there may be aspects of perseveration not captured by these tests - like you mention, there are different ways that one can be perseverative, and one single number can't ever represent the complexity of human behavior.

My PhD was (partially) on this. The inability to inhibit actions when facts change on a short timescale (sub-second) is thought to be biological. According to the most widely accepted theory, think of the brain as having a slow system and a fast system. The slow system is good for complex processing, but of course it is slow. There is also a fast system for quickly responding to things. If something changes while the slow system is working, usually the fast system is pretty good at stopping the slow system from acting, allowing the brain time to incorporate the new information into its plans. But if you are already planning on doing something and getting ready to do it, there is a limit to how quickly the fast system can interrupt the slow system. It tends to be on the order of 1/10 to 1/5 of a second.

On a longer timescale (minutes to days), there is a clinical symptom called "perseveration" whereby people can't let go of previously held beliefs in the face of changing information. It is common in, e.g., patients with schizophrenia.

Glad you like Paranoid Scientist, and thanks for the link to your project!

Yes, you're right that it only works on immutable data types. At first I had implemented something which copies the object, but this used way too much ram and ended up being quite buggy with a lot of edge cases.

I didn't invent the term hyperproperties (sadly). If you do end up implementing hyperproperties in runtime checks, you'll definitely want to look into to doing a statistical approach. The key to making this work is reservoir sampling (https://en.wikipedia.org/wiki/Reservoir_sampling) - there are some more details of how I did this in the Paranoid Scientist paper (https://arxiv.org/abs/1909.00427).

I love this. I think people often underestimate the value of dynamic runtime checks. For certain applications, e.g. many types of scientific/exploratory software, dynamic checks and static checkcs have nearly identical utility. But static checks can get really tricky to work with as a developer.

I built a library for doing modular runtime verification in Python (https://github.com/mwshinn/paranoidscientist) and evaluated it for scientific software. In the end, it was pretty effective, but not perfect, and there are some major changes I would make if I were to do this again. One problem was that some of the most important cases to check were important because they were difficult to check. (E.g. some function arguments change the meaning other arguments - this is extremely common in major frameworks like numpy/scipy.) By contrast, the flashiest feature in my package was runtime checking of hyperproperties (i.e. checking properties like monotonicity or concavity that depend on relationships between multiple executions of the function), but this was rarely used in practice.

The two most common criticisms I hear about runtime checking are (a) it is just an assert statement under the hood, and (b) the performance hit is unacceptable and the only solution is static checks. Regarding (a), sure, they may reduce to assert statements, but most idioms in programming also "reduce to something under the hood". The question is whether dynamically-checked (refinement) types/predicates are a useful abstraction, and in my experience, yes they are. Regarding (b), you probably don't want to use runtime checks on software intended to be run primarily by people other than the developers. But many classes of problems, software is written as a means of discovery rather than as a tool for someone else to accomplish a particular task. For these problems, runtime checks and static checks are approximately equally useful. Static checks are nice to avoid because they can get you into deep water really quickly, so their scope can be quite limited. Also, people tend to overestimate the performance penalty of dynamic checks. Even complex checks often incur no more than a 10% performance penalty.