HN user

sarosh

876 karma
Posts73
Comments126
View on HN
news.ycombinator.com 2y ago

Ask HN: Best Board Games of 2023?

sarosh
3pts1
tv.apple.com 4y ago

Foundation (Apple TV)

sarosh
3pts0
www.unofficialgoogledatascience.com 5y ago

Why model calibration matters and how to achieve it

sarosh
1pts0
arxiv.org 5y ago

Underspecification Presents Challenges for Credibility in ML

sarosh
2pts0
www.nytimes.com 5y ago

SpaceX’s ‘Resilience’ Lifts 4 Astronauts into New Era of Spaceflight

sarosh
3pts0
www.amazon.com 6y ago

Amazon Echo Loop – Smart Ring with Alexa

sarosh
1pts0
www.wired.com 6y ago

The Evidence That Links Russia’s Most Brazen Cyberattacks

sarosh
4pts0
sunnyday.mit.edu 7y ago

Engineering a Safer World (Leveson)

sarosh
1pts0
jeffe.cs.illinois.edu 8y ago

Algorithms, Etc. (2015)

sarosh
66pts6
arxiv.org 8y ago

Emergence of Invariance and Disentangling in Deep Representations

sarosh
3pts0
arxiv.org 8y ago

Deep Learning: A Critical Appraisal [pdf]

sarosh
85pts18
paulgraham.com 8y ago

YC Patent Pledge

sarosh
4pts1
www.supremecourt.gov 9y ago

SCOTUS: Packingham V. North Carolina [pdf]

sarosh
1pts0
memory-beta.wikia.com 9y ago

Ferengi Rules of Acquisition

sarosh
8pts1
wccftech.com 9y ago

AMD Ryzen Threadripper 1950X 16 Core, 32 Thread CPU at 3.4 GHz

sarosh
16pts3
www.youtube.com 9y ago

PBS Idea Channel – This Episode Was Written by an AI

sarosh
3pts2
research.fb.com 9y ago

Learning to Segment

sarosh
1pts0
medium.com 9y ago

Stop Pushing Pixels

sarosh
2pts0
www.mitpressjournals.org 9y ago

Computational Linguistics and Deep Learning

sarosh
3pts0
www.anandtech.com 10y ago

The Nvidia GTC 2016 Keynote

sarosh
1pts0
en.wikipedia.org 15y ago

Math: Voronoi Diagram

sarosh
2pts0
www.remindblog.com 16y ago

Coloring a Graphic Novel – Part 2

sarosh
2pts0
agtb.wordpress.com 16y ago

Auction Algorithm for Bipartite Matching « Algorithmic Game-Theory/Economics

sarosh
4pts0
arxiv.org 16y ago

[1004.1001] The Graph Traversal Pattern

sarosh
4pts0
blog.heroku.com 16y ago

Background Jobs with DJ on Heroku

sarosh
15pts7
blog.tannerburson.com 16y ago

Multiple Sinatra .90 applications in one process

sarosh
1pts0
arxiv.org 16y ago

Information-Sharing and Privacy in Social Networks

sarosh
1pts0
arxiv.org 16y ago

Is It Real, or Is It Randomized?: A Financial Turing Test

sarosh
3pts0
arxiv.org 16y ago

PageRank: Stand on the shoulders of giants

sarosh
1pts0
arxiv.org 16y ago

Physics and Five Problems in the Philosophy of Mind

sarosh
7pts10

But why does, as you explain "training goes brrr"?

Francis Bach, the author, makes a good faith effort to explain exactly why this material is beneficial (see https://francisbach.com/my-book-is-out/):

"Why yet another book on learning theory? ...the main reason is that I felt that the current trend in the mathematical analysis of machine learning was leading to overly complicated arguments and results that are often not relevant to practitioners. Therefore, my aim was to propose the simplest formulations that can be derived from first principles, trying to remain rigorous without overwhelming readers with more powerful results that require too much mathematical sophistication."

From my own reading and experience on the mathematical analysis approach of this "training goes brrr" approach, I thought the material in Chapter 12, Overparameterized Models, was interesting and coherent with 12.2.4 Linear Regression with Gaussian Projections being an especially elegant explanation. It would be interesting to hear if you had read/skimmed/purused this section and found it wanting etc.

This is the PDF of the following 2011 book focused on FFTs and fast arithmetic for both real numbers and finite fields. https://www.amazon.com/Matters-Computational-Ideas-Algorithm... The author is Jörg Arndt: born 1964 in Berlin, Germany. Study of theoretical physics at the University of Bayreuth, and the Technical University of Berlin, Diploma in 1995. PhD in Mathematics, supervised by Richard Brent, at the Australian National University, Canberra, in 2010.

Interesting that the underlying model, a LoRA fine-tune of Qwen2.5-Coder-32B, relies on synthetic data from Claude[1]:

  But we had a classic chicken-and-egg problem—we needed data to train the model, but we didn't have any real examples yet. So we started by having Claude generate about 50 synthetic examples that we added to our dataset. We then used that initial fine-tune to ship an early version of Zeta behind a feature flag and started collecting examples from our own team's usage.

  ...

  This approach let us quickly build up a solid dataset of around 400 high-quality examples, which improved the model a lot!
I checked the training set, but couldn't quickly identify which were 'Claude' produced[2]. Would be interesting to see them distinguished out.

[1] https://zed.dev/blog/edit-prediction [2]: https://huggingface.co/datasets/zed-industries/zeta

Defer to other experts, but (briefly) normalizing flows are a method for constructing complex distributions by transforming a probability density through a series of invertible transformations. Normalizing flows are trained using a plain log-likelihood function, and they are capable of exact density evaluation and efficient sampling. See:

Danilo Rezende and Shakir Mohamed. Variational inference with normalizing flows. In ICML, 2015. Link: https://bigdata.duke.edu/wp-content/uploads/2022/08/1505.057...

Laurent Dinh, David Krueger, and Yoshua Bengio. Nice: Non-linear independent components estimation. In ICLR Workshop, 2015. Link: https://arxiv.org/pdf/1410.8516

And for your direct question, the following paper "Efficient Bayesian Sampling Using Normalizing Flows to Assist Markov Chain Monte Carlo Methods" appears upon a superficial glance to be relevant. Link: https://arxiv.org/pdf/2107.08001

PhD Simulator 3 years ago

Looking at https://research.wmz.ninja/projects/phd/rulesets/default/eve... provides most of the key 'game loops', i.e.

# idea -> prelim -> major -> 2 figures -> submitted paper

interesting to see the hypothesis about reading more papers being borne out:

# increase the success rate as the player reads more papers probability: 0.60 + player.readPapers / 100 - itemCount('idea') / 20

Also interesting to see that passing the qualification exam provides the largest player.hope boost (+10)

Was fun to see the TooManyIdeas random event - now to actually get it to trigger.

The author, David Mumford, is known for his distinguished work in algebraic geometry, and was awarded the Fields Medal in 1974. [2]

Pattern theory was formulated by Ulf Grenander to describe knowledge of the world as patterns. [3]

Prof. Mumford explains that "[s]everal essential ideas brought me to realize how Grenander's Pattern Theory was the right way to understand almost all cognitive skills and especially vision. One was the emphasis on pattern synthesis as well as pattern analysis." [0]. Second, "was that natural signals given by functions f vary not only by random additive perturbations but often by composition with random rearrangements of their domain. The resulting probability distribution in the vector space of signals is nothing like Gaussian. Its support is usually a twisted snakey submanifold. This puts the lie to all simplistic Gaussian pattern recognition algorithms." [0]. And third, "graphical structures were everywhere in the representations of ideas in cognitive domains" [0].

His most recent work seems to be "Pattern Theory, the Stochastic Analysis of Real World Signals" [1]

[0] https://www.dam.brown.edu/people/mumford/vision/pattern.html

[1] https://www.amazon.com/dp/1568815794/ref=sr_1_1?ie=UTF8&qid=...

[2] https://en.wikipedia.org/wiki/David_Mumford

[3] https://en.wikipedia.org/wiki/Pattern_theory

An interesting take from Mike Pondsmith given his heavy involvement in the venture at the end of (at least in my opinion) well-written article: "[C]omparing the tabletop experience with its video-game incarnation, he noted that the latter doesn’t really compare to the former when it comes to self-expression. “You could be you in a tabletop game and bring all the stuff that you wanted to bring into it,” he said. “A tabletop game is limitless. A video game, by its very nature of how it’s designed, has some limits.” "

The paper itself is here: https://arxiv.org/abs/2111.09259 with the key conclusion that "Examining the evolution of human concepts using probing showed that many human concepts can be accurately regressed from the AZ network after training, even though AlphaZero has never seen a human game of chess, and there is no objective function promoting human-like play or activations" and "[t]he fact that human concepts can be located even in a superhuman system trained by self-play broadens the range of systems in which we should expect to find human-understandable concepts"

A nice summary of the current understanding (from a lay perspective) on manifold "sameness".

From the article: "But Freedman’s [1981] proof left open the “smooth” four-dimensional Poincaré conjecture, which says that any four-dimensional smooth manifold that is homotopy equivalent to the four-dimensional sphere is also diffeomorphic to the four-dimensional sphere. This is an even stronger statement than the one Freedman proved — since a diffeomorphism is a stronger form of equivalence than a homeomorphism — and one that mathematicians today have no idea how to settle.

This leaves them in the strange position of being unable to perform one of the most basic classification tasks of all: recognizing when a smooth four-dimensional manifold is really a sphere."

There is already a nice writeup on the current incident from Cloudflare at https://blog.cloudflare.com/october-2021-facebook-outage/

They key observations:

"Due to Facebook stopping announcing their DNS prefix routes through BGP, our and everyone else's DNS resolvers had no way to connect to their nameservers. Consequently, 1.1.1.1, 8.8.8.8, and other major public DNS resolvers started issuing (and caching) SERVFAIL responses.

But that's not all. Now human behavior and application logic kicks in and causes another exponential effect. A tsunami of additional DNS traffic follows.

This happened in part because apps won't accept an error for an answer and start retrying, sometimes aggressively, and in part because end-users also won't take an error for an answer and start reloading the pages, or killing and relaunching their apps, sometimes also aggressively."

This Economist article is from October of 2020. Peter Turchin is not a historian. He is instead in the Department of Ecology and Evolutionary Biology at the Universtiy of Connecticut. He advocates for a field of 'cliodynamics' which tries to apply math to meaningfully describe and predict social trends, especially large ones such as collapse. The Nature article from about decade ago: https://www.nature.com/articles/454034a A large project he directs and uses to study this: http://seshatdatabank.info/seshat-about-us/

Some descriptive work such as: https://www.pnas.org/content/115/2/E144.full

Sone of the work [0] has been criticized on methodological grounds [1] and some recent (2017) related work [2]. Finally, in his own words, a comparison of Psychohistory and Cliodynamics (2012) [3]

[0] https://www.nature.com/articles/s41586-019-1043-4 [1] https://github.com/babeheim/moralizing-gods-reanalysis [2] https://www.pnas.org/content/114/30/7846.full [3] http://peterturchin.com/cliodynamica/psychohistory-and-cliod...

There may be issues in their legal structure the prohibit direct investment or their Board feels more comfortable offloading the technology risk to someone else. Alternatively, they may want certain liquidity features (or other options) that may not be directly available.

Probably the most interesting piece is "Box G: “Low-For-Long” Interest Rates and Implications for Financial Stability" ... " A longerterm risk is how the market participants’ exposures to greater levels of duration risk affect financial stability when rates eventually increase. The 2013 Taper Tantrum is an example of this potential dynamic, although the wider financial stability implications of that episode were limited. The potential risk here is that unexpected increases in rates negatively affect the balance sheets of financial institutions in such a way that leads to financial instability. Banks without adequate capital buffers could face solvency issues, while pension funds and insurance companies could experience liquidity problems related to losses on derivatives positions or increases in early liquidations. Additionally, with valuations in both equity and credit markets relatively high by historical standards and likely to become further stretched in a low-for-long environment, the risk of a sharp correction becomes more likely, especially in conjunction with high levels of leverage or excessive reliance on short-term wholesale funding. Even small changes to expectations of far-in-the-future cash flow may have a disproportionate effect on current valuations when interest rates are low. As a result, such rate changes can lead to sharp adjustments in valuations. The potential negative effects that an unexpected increase in rates would have across a variety of market participants make this longer-term risk worth monitoring. Adequate guidance on the timing and pace of any such policy-related increase will likely reduce this risk."

The 2013 Taper Tantrum refers to the 2013 collective reactionary panic that triggered a spike in U.S. Treasury yields, after investors learned that the Federal Reserve was slowly putting the breaks on its quantitative easing (QE) program. See https://www.reuters.com/article/us-usa-fed-2013-timeline/key...

"The U.S. economy was in the midst of the longest post-war economic expansion, with historically low levels of unemployment, prior to the onset of the COVID-19 pandemic earlier this year. The global pandemic not only brought about a public health crisis but also caused a contraction of economic activity at an unprecedented pace. Initially, the pandemic reduced consumer spending, slowed manufacturing production, and led to widespread business closures. The unemployment rate surged from 3.5 percent in February to a record high of nearly 15 percent in April. Since then, extraordinary measures undertaken by policymakers have succeeded in arresting the decline in economic conditions, initiating a recovery and lowering the unemployment rate to 7.9 percent as of September. However, a protracted virus outbreak poses downside risks that can slow the recovery and even prolong the economic downturn."

...

"With cash flows impaired due to the COVID-19 pandemic, many businesses may be challenged to service their debt. Since March, nearly $2 trillion in nonfinancial corporate debt has been downgraded, and default rates on leveraged loans and corporate bonds have increased considerably. The growing number of bankruptcy filings could stress resources at courts and make it harder for firms to obtain critical debtor-in-possession financing. It could also prevent many firms from restructuring their debt in a timely fashion, potentially forcing them into liquidation."

...

"Money market funds (MMFs) offer shareholders redemptions on a daily basis while holding many short-term assets that are less liquid, especially in times of stress. Stresses on prime and tax-exempt money funds in March revealed continued structural vulnerabilities, which led to increased redemptions and, in turn, likely contributed to the stress in STFMs. Among institutional and retail prime MMFs, outflows as a percentage of fund assets exceeded that of the September 2008 crisis. Outflows abated after the Federal Reserve announced support for the CP market and MMFs."

...

See also W. Brendel and M. Bethge. "Approximating CNNs with bag-of-local-features models works surprisingly well on ImageNet." https://arxiv.org/abs/1904.00760 (which is referenced in this paper) and explains "[t]his suggests that the improvements of DNNs over previous bag-of-feature classifiers in the last few years is mostly achieved by better fine-tuning rather than by qualitatively different decision strategies."

Sad DNS Explained 6 years ago

The article explains the novelty as: “[T]he researchers found a very clever trick: they leverage ICMP rate limits as a side channel to reveal whether a given port is open or not. ICMP rate limiting was introduced (somewhat ironically, given this attack) as a security feature to prevent a server from being used as an unwitting participant in a reflection attack.“ The suggested mitigation is Linux kernel upgrade to roll out unpredictable ICPM rate limiting.

Peter Turchin has a blog post (http://peterturchin.com/cliodynamica/the-mad-prophet-of-conn... ) expressing his own views about the Atlantic article, explaining: "I cringed in a number of places as I read his article. Yes, I propose a fairly ambitious program of testing theories about historical processes by translating them into explicit models and then testing model predictions with large datasets. But no, I don’t think of myself as a Hari Seldon. In fact, the fictional Hari Seldon had no appreciation of nonlinear dynamics and mathematical chaos (because Asimov wrote the stories before the discovery of chaos). And cliodynamics is not psychohistory."