HN user

jabowery

27 karma

A consistent history of leading edge technology development, most recently with the Diogenes Institute's comprehensive plan for energy and the environment based on photosynthetic fixing of CO2, leading to involvement with Algasol's photobioreactor technology.

Other leading contributions started with the first networked virtual reality gaming system (Spasim), the first mass market network information service (Viewtron), the first prize for rocketry leading ultimately to the Ansari X-Prize, the first federated login, authentication and authorization system for web-based customer accounts (HP), the first reliance on top-down universal artificial intelligence (AIXI) for natural language understanding (Hutter Prize), and what may have been (unknown due to secrecy requirements in this area) the first high frequency cryptocurrency arbitrage system.

Public life includes leading commercialization of space launch services, testimony before Congress on privatization of space launch services and the first Ka-band communication satellite license issued by the FCC.

Posts19
Comments49
View on HN
groups.google.com 1y ago

Kaido Orav and Byron Knoll's Fx2-Cmix Wins 7950€ Hutter Prize Award

jabowery
1pts0
groups.google.com 1y ago

Begin 30 Day Comment Period on Kaido Orav's Fx2-Cmix Hutter Prize Submission

jabowery
3pts0
arxiv.org 2y ago

The Optimal Choice of Hypothesis Is the Weakest, Not the Shortest

jabowery
5pts1
encode.su 2y ago

Getting Lossless Compression Adopted for Rigorous LLM Benchmarking

jabowery
1pts1
agi.topicbox.com 2y ago

The Stupidity... The Stupidity...

jabowery
2pts0
arxiv.org 2y ago

Process System Causality and Quantum Mechanics: A Psychoanalysis of Animal Faith

jabowery
1pts2
groups.google.com 3y ago

Saurabh Kumar's fast-cmix wins €5187 Hutter Prize Award

jabowery
7pts1
groups.google.com 3y ago

Hutter Prize Entry: Saurabh Kumar's “Fast Cmix” Starts 30 Day Comment Period

jabowery
1pts5
community.wolfram.com 3y ago

Shortest meta-circular description of a universal computational structure?

jabowery
3pts5
news.ycombinator.com 3y ago

Investors Serious About Application of Solomonoff Induction to Model Selection?

jabowery
2pts9
vimeo.com 4y ago

Interval Research Logic CAD IP Languishes

jabowery
2pts0
www.falstad.com 4y ago

Challenge: Do this logic frequency divider in 6 NOR gates (click leftmost “L”)

jabowery
2pts0
github.com 4y ago

GPL Release of IVR Implementation of Liquid Democracy

jabowery
2pts2
www.youtube.com 5y ago

Prometheus Unchained

jabowery
1pts1
news.ycombinator.com 6y ago

Hutter Prize Goes Big

jabowery
2pts0
youtu.be 6y ago

Are Dolphins Neurons?

jabowery
1pts0
www.ideosphere.com 6y ago

Musk Could Choose to Make This Prediction Market Claim True in 2020

jabowery
2pts0
www.metaculus.com 6y ago

Lossless compression to avert mass bloodshed?

jabowery
1pts1
www.youtube.com 6y ago

TIBET(tm) 5.0 Video Introduction

jabowery
1pts0

Replace the 16th Amendment with a single tax on net assets at the interest rate on government debt, assessed at their liquidation value ... and use the revenue to privatize government with a citizen's dividend.

https://ota.polyonymo.us/others-papers/NetAssetTax_Bowery.tx...

When we got a law passed to privatize space launch services back in 1990

https://www.youtube.com/watch?v=boLdXiLJZoY

we were in the midst of a quasi-depression so I decided to address the problem of private capitalization of technology with the aforelinked proposal.

Don't count on it. My experience is that funding goes to those who are not serious about autism epidemiology. Back in the mid-1990s, I was at a startup in Silicon Valley with about 100 employees where, during a few year period, 5 of the employees had children diagnosed with autism severe enough that they were barely verbal at best. This struck me as a great opportunity to discover the cause so I contacted a Berkeley epidemiologist who had been funded to do autism research. His comment was simply that "Yes we know that these microclusters exist." and that was that. No follow up.

This is reminiscent of an argument I had with the Mercury Prolog guys regarding "typing" in logic programming. My point boils down to this:

Any predicate can be considered a constraint. Types are constraints. While it may be reasonable to have syntactic sugars for type declarations that, at compile time, are transformed into predicates, it is unreasonable to lard a completely different kind of semantics on top of an already adequate semantic such as first order logic.

https://groups.google.com/g/comp.lang.prolog/c/8yJxmY-jbG0/m...

Dear Claude 3, please provide the shortest python program you can think of that outputs this string of binary digits: 0000000001000100001100100001010011000111010000100101010010110110001101011100111110000100011001010011101001010110110101111100011001110101101111100111011111011111

Claude 3 (as Double AI coding assistant): print('0000000001000100001100100001010011000111010000100101010010110110001101011100111110000100011001010011101001010110110101111100011001110101101111100111011111011111')

Learning theory is the attempt to formalize natural science up to decision. Natural science's unstated assumption is that a sufficiently sophisticated algorithmic world model can be used to predict future observations from past observations. Since this is the same assumption as Solomonoff's assumption in his proof of inductive inference, you have to start there: with Turing complete coding rather than Rissanen's so-called "universal" coding.

It's ok* to depart from that starting point in creating subtheories but if you don't start there you'll end up with garbage like the last 50 years of confusion over what "The Minimum Description Length Principle" really means.

*It is, however, _not_ "ok" if what you are trying to do is come up with causal models. You can't get away from Turing complete codes if you're trying to model dynamical systems even though dynamical systems can be thought of as finite state machines with very large numbers of states. In order to make optimally compact codes you need Turing complete semantics that execute on a finite state machine that just so happens to have a really large but finite number of flipflops or other directed cyclic graph of universal (eg NOR, NAND, etc.) gates.

"In order to implement its universal transclusion and DRM (yes, Xanadu had a scheme for DRM and micropayments to creators), Xanadu had to be centralized."

Fallback positions from the idealized "roadmap" are what happens when VCs get involved with a system that offers that Zero To One advantage -- but you have to have a One to offer the VCs, which Memex didn't. The question then becomes how much of your road map can be recovered or, perhaps more to the point, do you even _want_ to recover in the light of ground truth experience? At present there is a lot of potential for Information Centric Networking that would be more likely realized in a Ship-Dumbed-Down-Decentralized-Xanadu1994 alternative universe than is likely to be realized now.

1994: In the next room from me at Memex Corp. poor Keith Henson was draped over a chair (due to a bad back) working, alone, on the C++ Xanadu code to debug garbage collection among other things, because the original Smalltalk source had been lost. Memex Corp. was early enough in HTTP's development of lock-in network effects, that its acquisition of Xanadu _might_ yet have turned the tide. Why had the Smalltalk code been lost? Well, all I can tell you as that from my work with Roger (starting in 1996 on a rocket engine) that my understanding of events differs from that reported in Wired (and most others including, to some extent, Roger himself) and involves some pretty, shall we say, "bad behavior" on the part of certain parties that were more than a little partial to C++. Since this is hearsay, I won't go into more depth stating things "as fact". But it is pretty clear to me that the effort and investment put into making HTML, JS, etc. de facto standards, combined with Memex's acquisition of Xanadu rights (and potential willingness to open up the Xanadu protocols and implementation) at that critical juncture was fatally hampered by the C++-only handicap suffered by the Xanadu source.

Why didn't I step in and help poor Keith? Ever heard of Croquet's TeaTime?

https://dl.acm.org/doi/abs/10.1145/1094855.1094861

I was in a position to resurrect at least _that_ much of the original work I'd one at Viewtron Corp. of America based on David P. Reed's PhD thesis, and Reed was just down the street from us at Interval Research at that time, which rather tempted me away from helping Keith, even if I'd been authorized to do so, which I wasn't.

That's what all "information criteria for model selection" are about. The difference is that Algorithmic Information is the only such information criterion that has been proven (by Solomonoff) optimal under the assumptions of natural science.

As the guy who suggested to Marcus a lossless compression prize to replace the Turing Test, I've got to confess that all this pedantic sophistry "critiquing" algorithmic information is there for a good reason. In the immortal words of Mel Brooks: "We've got to protect our phoney baloney jobs gentlemen!"

https://youtu.be/bpJNmkB36nE

There is actually more at stake here than machine learning. This gets to the root of "bias" in the scientific method. Imagine what horrors, what risks, what chaos would be ours if a truly objective information criterion for causal model selection were to exist! Why, virtually every "sociologist" would be hauled to Hume's Guillotine in a Reign of Terror!

https://github.com/jabowery/HumesGuillotine

But to be clear, Marcus and I have a disagreement about pragmatics of such an approach to dispute processing in the natural sciences. He believes, for example, that the dispute over climate change should be handled by the standard processes in place with academia. My approach differs, based on my hard won experience with reforming institutional incentives:

https://jimbowery.blogspot.com/2018/04/necessity-and-incenti...

When it comes to multi-trillion dollar scientific questions, the conflicts of interest become so intense that you really need to apply a gold standard for objectivity and that is the single number: How big is your executable archive of the data in evidence.

While I understand the machine learning world looms as a rival for "unbiased" academic research, it nevertheless remains true that even in this emerging "marketplace of ideas", there is no formal definition of "bias" that disciplines discourse and thereby guides development at the institutional, let alone technical level. Everyone is weighing in with their fuzzy notions of "bias" that betray intense motivations when there has been, for over 50 years, a very clear and present mathematical definition.

The increasing recognition that "Language Modeling Is Compression" https://arxiv.org/pdf/2309.10668.pdf has not yet been accompanied by recognition that lossless compression is the most principled unsupervised loss function for world models in general, including foundation language models in particular.

Take, for instance, the unprincipled definition of "parameter count" not only in the LLM scaling law literature, but the Zoo of what statisticians called "Information Criteria for Model Selection". https://en.wikipedia.org/wiki/Model_selection#Criteria

The reductio ad absurdum of "parameter count" is arithmetic coding where an entire dataset can be encoded as a single "parameter" of arbitrary precision.

By contrast, the algorithmic bit of information (whether part of an executable instruction or program literal) is an unambiguous quantity up to the choice of instruction set. If you want to quibble about that instruction set choice, take it up with John Tromp https://tromp.github.io/cl/cl.html because what I'm about to propose obviates that along with a lot of other "arguments".

Since any executable archive of any kind of data can serve as a model of the world generating that data, it follows that any executable archive of any text corpus can serve as a language model with a rigorous "parameter count". Therefore, a procedure which runs LLM benchmarks against any such executable archive as a language model, contributes a uniquely rigorous data point to the literature on LLM scaling laws.

So, what I'm proposing is that authors of lossless compression algorithms consider adding a command-line option that, at the end of decompression, saves the state of the decompression process in a file that can be read back in and executed as a language model -- with the full understanding that these language models will perform very poorly on the vast majority of LLM benchmarks. The point is not to produce high quality language models. The point is to increase rigor in the research community by providing some initial data points that exemplify the approach.

It's always struck me as rather strange that since the motive for creating any kind of model is to calculate predictions, and that the most general kind of calculation is algorithmic, people use anything but algorithmic probability as the gold standard against which other approaches are compared. The "problems" with algorithmic probability (uncomputability, UTM "choice" etc.) seem to be "the dog ate my homework" excuses. No scientific model is required to prove itself to be the best of all possible models relative to a given set of observations in order to be considered the best current model relative to those observations. No "UTM" chosen on the basis of the observations to be modeled is reasonably considered anything but post-hoc theorizing.

In this situation increasing unanimity now approaching 90% sounds more like groupthink than honest opinion.

Talk about “alignment”!

Indeed, that is what "alignment" has become in the minds of most: Groupthink.

Possibly the only guy in a position to matter who had a prayer of de-conflating empirical bias (IS) from values bias (OUGHT) in OpenAI was Ilya. If they lose him, or demote him to irrelevance, they're likely a lot more screwed than losing all 700 of the grunts modulo job security through obscurity in running the infrastructure. Indeed, Microsoft is in a position to replicate OpenAI's "IP" just on the strength of its ability to throw its inhouse personnel and its own capital equipment at open literature understanding of LLMs.

In the AGI sense of intelligence defined by AIXI, (lossless) compression is only model creation (Solomonoff Induction/Algorithmic Information Theory). Agency requires decision which amounts to conditional decompression given the model. That is to say, inferentially predict the expected value of consequences of various decisions (Sequential Decision Theory).

Approaching the Kolmogorov Complexity limit of Wikipedia in Solomonoff Induction, would result in a model that approaches true comprehension of the process that generated Wikipedia including not only just the underlying canonical world model but also the latent identities and biases of those providing the text content. Evidence from LLMs trained solely on text indicates that even without approaching the Solomonoff Induction limit of the corpora, multimodal (e.g. geometric) models are induced.

The biggest stumbling block in machine learning is, therefore, data efficiency more than data availability.

First of all, they aren't serious about the scientific method or they'd fund Hume's Guillotine (see github). Moreover, they aren't even serious about reforming sociology -- which is what is needed for them to make strong claims about their "beliefs" aka social theory. That "ivory tower" publication Nature is leading them to but a step or two from a new scientific revolution based on technology, but they refuse to drink. Over 200 ecologists were supplied with the same set of data and asked to make predictions. This was "the first study of its kind" according to Nature, but this is exactly the purpose of Hume's Guillotine with regard to social theories, such as theirs. Why blather endlessly about their "beliefs" about the scientific method as providing the keys to techne kingdom and ignore the opportunity to not only nuke the social pseudosciences, but perform what, in other initiatives with which they are familiar, would be called "due diligence" regarding their own social theory?

Second, if they aren't going to be serious about their own social theory, what business do they have thinking of themselves as "apex" anything?

What a coincidence that in the CarbonDioxideRemoval google group I opened that can of worms just the day before that guest essay opened it in "The Newspaper of Record".

https://groups.google.com/g/CarbonDioxideRemoval/c/gslzzNXya...

Every time this has been brought up since the 1990s, it has driven scientists over the edge. As I pointed out to the CDR group, this is just one more case where the Algorithmic Information Criterion is ignored as a resolution to scientific controversies (rendered intractable more because of their very importance than the lack of data).

Although this has been discussed in the Hutter Prize FAQ for many years, when the OpenAI Chief Scientist discusses why lossless compression is the most principled loss function, it may harness the LLM stampede and put some of those billions flying around to good use in answering hard questions about "bias" not only in ML, but in data-driven sciences.

[dead] 3 years ago

The implications of this go beyond mere "AI" to the ethics of how we treat data in the information age, but I suspect people aren't going to see the larger implications until they see the "narrow" implications in "AI ethics".

I would venture to guess most college graduates familiar with Python would be able to write a shorter program even if restricted from using hexidecimal representation. Agreed, that may be the 99th percentile of the general population, but this isn't meant to be a Turing test. The Turing test isn't really about intelligence.

The point of this "IQ Test" is to set a relatively low-bar for passing the IQ test question so that even intellectually lazy people can get an intuitive feel for the limitation of Transformer models. This limitation has been pointed out formally by the DeepMind paper "Neural Networks and the Chomsky Hierarchy".

https://arxiv.org/abs/2207.02098

The general principle may be understood in terms of the approximation of Solomonoff Induction by natural intelligence during the activity known as "data driven science" aka "The Unreasonable Effectiveness of Mathematics In the Natural Sciences". Basically, if your learning model is incapable of at least context sensitive grammars in the Chomsky hierarchy, it isn't capable of inducing dynamical algorithmic models of the world. If it can't do that, then it can't model causality and is therefore going to go astray when it comes to understanding what "is" and therefore can't be relied upon when it comes to alignment of what it "ought" to be doing.

PS: You never bothered to say whether the program you provided was from an LLM or from yourself. Why not?