HN user

concurrentsquar

95 karma
Posts1
Comments25
View on HN

My team can't even evaluate Ball (as a tool for physical simulation of spherical cow-like objects): https://github.com/nate-parrott/ball/issues/9

Currently, we are using Unreal Engine 5 to do our hundreds of architectural physics simulations - the major issue is that UE5 is very slow on *the EC2 instance* (we only have one 2048 core EC2 instance shared between the entire office; we used to use Vercel and Cloudflare but we had to sell our homes to suddenly subscribe to Cloudflare Enterprise (the CF sales guy told us that we would not be allowed to run a CF Worker for more than 30 days without it, even though we had a CF worker run for 37 years, and many of our CF workers have been running before the creation of CF (nobody knows why)) and a giant spike in our Vercel Cuda Function Invocations (for GPGPU compute on the Edge, allowing architects to view the collapse of their buildings with only ~53 ms of latency (compared to ~53 ms without Next.js))). Ball seems much faster (it can run on a Macbook Air), potentially allowing us to save at least several tens of millions of dollars per year on AWS costs.

Calculating the cost per kilogram for LEO with Starship gives me a new startup idea: small business (or even personal) interplanetary postal service.

It only costs $150 per kg in the near future to send objects into space with Starship; so I could, for example, send a Raspberry Pi (47 grams) into LEO for ~7 dollars (as long as I also had 149 tons of other objects from other people to send). A more useful use case would sending fully automated manufacturing facilities (probably either for semiconductors (https://www.nasa.gov/general/the-benefits-of-semiconductor-m...) or crystals (https://uofuhealth.utah.edu/newsroom/news/2017/07/proteinxl))

100k Stars 2 years ago

Great visualization, though (ironically, as one of the first Chrome experiments) the music no longer works on Chrome by default (go to site settings > sound and set it to "Allow" to hear it), and it is somewhat outdated now (for example, it states that no exoplanets have been discovered orbiting Proxima Centauri (and that the 'proposed' JWST is required to find these planets)).

OpenAI could either hire private testers or use AB testing on ChatGPT Plus users (for example, oftentimes, when using ChatGPT, I have to select between 2 different responses to continue a conversation); both are probably much more better (in many aspects: not leaking GPT-4.5/5 generations (or the existence of a GPT-4.5/5) to the public at scale and avoiding bias* (because people probably rate GPT-4 generations better if they are told (either explicitly or implicitly (eg. socially)) it's from GPT-5) to say the least) than putting a model called 'GPT2' onto lmsys.

* While lmsys does hide the names of models until a person decides which model generated the best text, people can still figure out what language model generated a piece of text** (or have a good guess) without explicit knowledge, especially if that model is hyped up online as 'GPT-5;' even a subconscious "this text sounds like what I have seen 'GPT2-chatbot' generate online" may influence results inadvertently.

** ... though I will note that I just got a generation from 'gpt2-chatbot' that I thought was from Claude 3 (haiku/sonnet), and its competitor was LLaMa-3-70b (I thought it was 8b or Mixtral). I am obviously not good at LLM authorship attribution.

One distinguishing feature of "deluxe-chat": although it gives high quality answers, it is very slow, so slow that the arena displays a warning whenever it is chosen as one of the competitors

Beam search or weird attention/non-transformer architecture?

Reddit may have told OpenAI to pay (probably a lot of) money to legally use Reddit content for training, which is something Reddit is doing with other AI labs (https://www.cbsnews.com/news/google-reddit-60-million-deal-a... ); but GPTBot is not banned under the Reddit robots.txt (https://www.reddit.com/robots.txt).

This is assuming that lmsys' GPT-2 is retained GPT-4t or a new GPT-4.5/5 though; I doubt that (one obvious issue: why name it GPT-2 and not something like 'openhermes-llama-3-70b-oai-tokenizer-test' (for maximum discreetness) or even 'test language model (please ignore)' (which would work well for marketing); GPT-2 (as a name) doesn't really work well for marketing or privacy (at least compared to the other options)).

Lmsys has tested models with weird names for testing before: https://news.ycombinator.com/item?id=40205935

You don't (you have to use real-valued inertial 'latent weights' during training): https://arxiv.org/abs/1906.02107

(there is still a reduction in memory usage though (just not 24x):

"Furthermore, Bop reduces the memory requirements during training: it requires only one real-valued variable per weight, while the latent-variable approach with Momentum and Adam require two and three respectively.")

This is a related official press release (I can not find the actual source for this article): https://www.anl.gov/article/new-international-consortium-for...

The Trillion Parameter Consortium (TPC) brings together teams of researchers engaged in creating large-scale generative AI models to address key challenges in advancing AI for science. These challenges include developing scalable model architectures and training strategies, organizing, and curating scientific data for training models; optimizing AI libraries for current and future exascale computing platforms; and developing deep evaluation platforms to assess progress on scientific task learning and reliability and trust.

More new information from Swisher:

"More scoopage: sources tell me chief scientist Ilya Sutskever was at the center of this. Increasing tensions with Sam Altman and Greg Brockman over role and influence and he got the board on his side."

"The developer day and how the store was introduced was in inflection moment of Altman pushing too far, too fast. My bet: He’ll have a new company up by Monday."

[source: https://twitter.com/karaswisher/status/1725702501435941294]

Sounds like you exactly predicted it.

Being able to compress well is closely related to intelligence as explained below. While intelligence is a slippery concept, file sizes are hard numbers. Wikipedia is an extensive snapshot of Human Knowledge. If you can compress the first 1GB of Wikipedia better than your predecessors, your (de)compressor likely has to be smart(er). The intention of this prize is to encourage development of intelligent compressors/programs as a path to AGI.

- Marcus Hutter, http://prize.hutter1.net/

This competition was from 2006.

The consensus among AI experts (really just everyone) during the 1990s/2000s was that:

- AGI could be achieved by giant GOFAI (usually expert systems/knowledge bases) projects (like OpenCog and Cyc).

- ... or that AGI development is limited by lack of knowledge about key insights (mostly symbolic/rational) into intelligence, not computation/data. (IE AGI was viewed similarly to proving that NP = P or other very-hard math/computer science/psychology/philosophy problems).

- ... or that AGI will be achieved through brain scanning/connectomics.

- ... or that AGI is impossible.

Nobody (except for LeCun and Schmidhuber) paid much attention to neural networks until AlexNet (2012) showed that they could be ran and trained at fast speed and beat the symbolic competition. In the 2000s, only a real, actual psychic would be able to tell you that LLMs would be a valuable path for AGI research.

Here is a list of various expert (and "expert") perspectives on AGI during the 1990s/2000s (notice how nobody is talking about neural networks, and they are definitely not talking about anything remotely close to a LLM or transformer):

Copycat is a computer program designed to be able to discover insightful analogies, and to do so in a psychologically realistic way. Copycat's architecture is neither symbolic nor connectionist, no was it intended to be a hybrid other two (although some might see it that way); ... [describes a very symbolic system to our modern day eyes, though it was not really symbolic to 1990s AI researchers]

- Douglas Hofstadter and Melanie Mitchell, Fluid Concepts and Creative Analogies (Chapter 5), 1995

Interviewer: Are you an advocate of furthering AI research?

Dennett: I think that it’s been a wonderful field and has a great future, and some of the directions are less interesting to me and less important theoretically, I think, than others. I don’t think it needs a champion. There’s plenty of drive to pursue this research in different ways.

Dennett (cont): What I don’t think it’s going to happen and I don’t think it’s important to try to make it happen; I don’t think we’re going to have a really conscious humanoid agents anytime in the foreseeable future. And I think there’s not only no good reason to try to make such agents, but there’s some pretty good reasons not to try. Now, that might seem to contradict the fact that I work on a Cog project [sic] with MIT, which was of course is an attempt to create a humanoid agent, cogent, cog, and to implement the multiple drafts model of consciousness; my model of consciousness on it.

Dennett (later): [Cog is intended as a] proof of concept [for AGI]. You want to see what works but then you don’t have to actually do the whole thing.

- Daniel Dennett, Daniel Dennett Investigates Artificial Intelligence, Big Think, 2009

[Context: Marvin Minsky had a speech where he talked about how expert systems don't work, because they do not have any common sense (and the only solution seems to be to create a giant AGI project (without automatic data gathering)).]

Only one researcher has committed himself to the colossal task of building a comprehensive common-sense reasoning system, according to Minsky. Douglas Lenat, through his Cyc project, has directed the line-by-line entry of more than 1 million rules into a commonsense knowledge base.

- Mark Baard, AI Founder Blasts Modern Research, Wired, 2003

Section 1 discusses the conceptual foundations of general intelligence as a discipline, orienting it within the Integrated Causal Model of Tooby and Cosmides; Section 2 constitutes the bulk of the paper and discusses the functional decomposition of general intelligence into a complex supersystem of interdependent internally specialized processes, and structures the description using five successive levels of functional organization: Code, sensory modalities, concepts, thoughts, and deliberation. Section 3 ... [yada yada, this is old, wrong stuff]

- Eliezer Yudkowsky, Levels of Organization in General Intelligence, 2007, Machine Intelligence Research Institute

I could list more examples, but I have spent way too long on this post. What I will say is that Hutter probably had the most correct idea of how modern semi-general AI would work (from the 2000s). He figured out that compression is a extremely important component of intelligence > 10 years before everybody was doing LLMs. That is impressive.

I should probably write a blog post over this.

How about adding a arXiv Atom feed viewer for LK-99?

It should be easy to implement: https://export.arxiv.org/api/query?search_query=all:LK+AND+9... is the Atom feed I am using to track all current replications (ie none) - and it may be exciting (suspense) and useful.

Assuming that the livestream is using OBS - there have been many forum and reddit posts suggesting ways to add RSS feeds to OBS livestreams:

- https://obsproject.com/forum/threads/rss-feed-solution-for-o...

- https://www.reddit.com/r/obs/comments/g23rqr/can_i_put_in_a_...

- https://www.reddit.com/r/obs/comments/mvrkaj/rss_live_ticker...

I really need someone to bring me down a notch. This is too exciting!

It's on arXiv, which is a preprint journal, which means it has no peer-review; and is therefore generally less trustworthy (especially when the paper has no connection to a technical conference or is not being published elsewhere, and is in a non-computer science or mathematics field).

In addition to this, claims of room-temperature superconductors have been mired in controversy or otherwise proven false:

- http://www.superconductors.org/roomnano.htm (2004)

- https://www.nature.com/articles/nature.2012.11443 (2012)

- https://www.scientificamerican.com/article/a-superconductor-... (2018)

- https://www.quantamagazine.org/room-temperature-superconduct... (2020)

- https://forbetterscience.com/2023/03/29/superconductive-frau... (2022-2023)

Considering that many fraudulent claims of room-temperature superconductivity have gotten into Nature and other top-tier publications, I would wait for multiple independent recreations of the results in the paper.

I always wanted a Windows CE clamshell laptop, but I just have not had the time to look at ebay. The small laptops were always interesting to me.

Is there a 3D printable case for the Raspberry Pi that is like one of these? I looked at the DevTerm but I heard that it's not that good, and anyways it's not exactly the same as this.