HN user

antimatter15

4,769 karma

Kevin Kwok MIT'17 http://antimatter15.com

[ my public key: https://keybase.io/kkwok; my proof: https://keybase.io/kkwok/sigs/5qGM4MGK91WsEYGjQRq6h0oHCn4i8PZaN8gq4FonTRM ]

Posts52
Comments247
View on HN
datatracker.ietf.org 9mo ago

Proposed DNS RFC 8767: Serving Stale Data to Improve DNS Resiliency (2020)

antimatter15
3pts0
www.youtube.com 11mo ago

A Cheeky Pint with Anthropic CEO Dario Amodei [video]

antimatter15
2pts0
github.com 1y ago

Show HN: Vibechat – A chatroom for people bored waiting for Claude

antimatter15
10pts0
dynomight.net 1y ago

DumPy: NumPy except it's OK if you're dum

antimatter15
28pts1
github.com 2y ago

React 19 Breaks Async Composability

antimatter15
88pts110
docs.fcc.gov 2y ago

FCC Reaffirms Rejection of Nearly $900M Subsidy to Starlink [pdf]

antimatter15
4pts0
antimatter15.com 2y ago

Show HN: Real-Time 3D Gaussian Splatting in WebGL

antimatter15
309pts59
www.together.xyz 3y ago

Releasing 3B and 7B RedPajama

antimatter15
363pts106
github.com 3y ago

Show HN: Alpaca.cpp – Run an Instruction-Tuned Chat-Style LLM on a MacBook

antimatter15
673pts283
cresta.com 5y ago

How we reduced our AI labeling cost by 10x

antimatter15
34pts0
github.com 7y ago

Show HN: Eigensheep – Run Jupyter Cells on AWS Lambda

antimatter15
10pts2
libra.org 7y ago

Libra: Facebook’s Cryptocurrency Whitepaper

antimatter15
2pts0
www.youtube.com 7y ago

SIGGRAPH 2018 Technical Papers Preview

antimatter15
11pts0
www.notion.so 8y ago

Tools and Craft: An Interview with Andy Hertzfeld

antimatter15
9pts0
c.wsmoses.com 8y ago

Tapir: Embedding Fork-Join Parallelism into LLVM’s Intermediate Representation [pdf]

antimatter15
2pts0
cyborg.tenso.rs 8y ago

Show HN: Cyborg Writer - In-Browser Text Editor with Neural Autocomplete

antimatter15
71pts27
www.robinsloan.com 8y ago

Writing with the Machine (2016)

antimatter15
3pts0
github.com 8y ago

Starspace: Facebook's General Purpose Neural Embedding System

antimatter15
3pts0
franchise.cloud 8y ago

Show HN: Franchise - Open Source SQL Notebook

antimatter15
27pts3
tenso.rs 8y ago

Show HN: Play rock-paper-scissors against your computer via webcam, neural nets

antimatter15
196pts33
tenso.rs 8y ago

Show HN: TensorFire

antimatter15
564pts83
jamesboyk.com 9y ago

Machine with an Esthetic Purpose (2014)

antimatter15
1pts0
alpha.trycarbide.com 9y ago

Show HN: Carbide – A New Programming Environment

antimatter15
628pts136
scienceai.github.io 10y ago

JavaScript t-distributed stochastic neighbor embedding (t-SNE)

antimatter15
2pts1
www.youtube.com 10y ago

“I See What You Mean” by Peter Alvaro

antimatter15
2pts0
www.andrewbragdon.com 10y ago

Code Bubbles – Rethinking User Interface Paradigms for IDEs

antimatter15
1pts0
projectnaptha.com 12y ago

Project Naptha: a browser extension that enables text selection on any image

antimatter15
1055pts134
antimatter15.github.io 12y ago

Ocrad.js: Pure Javascript OCR with Emscripten

antimatter15
3pts0
antimatter15.com 12y ago

X-No-Wiretap

antimatter15
30pts5
offline-wiki.googlecode.com 14y ago

Download Entire Wikipedia for Offline Use With an HTML5 App

antimatter15
196pts50

I remember looking trying to build something like this 6 years ago[0]. There are some interesting APIs for injecting click/keystroke events directly into Cocoa, and other APIs for reading framebuffers for apps that aren't in the foreground.

In particular there was some prior art that I found for doing it from the OpenQwaQ project, which was a GPLv2 3D virtual world project in Squeak/Smalltalk started by Alan Kay[1] back in 2011.

If I recall correctly, it worked well for native apps, but didn't work well for Chromium/Electron apps because they would use an API for grabbing the global mouse position rather than reading coordinates from events.

[0] https://github.com/antimatter15/microtask/blob/master/cocoa/... [1]: https://github.com/OpenFora/openqwaq/blob/189d6b0da1fb136118...

Another fun calculation is that due to special relativity, a hard drive that is spinning gains a certain amount of mass due to the rotational kinetic energy and E=mc^2.

Assuming the platter is 100g, 42mm, spinning at 7200RPM, there is about 25J of rotational kinetic energy, whose mass equivalent is 2.8x10^-13g (0.28 femtograms).

Assuming 200 electrons per NAND floating gate with 3bits/cell TLC on a 2TB SSD, there would be 5.3x10^14 electrons, weighing about 0.5 femtograms.

Figure 3 on p.40 of the paper seems to show that their LLM based model does not statistically significantly outperform a 3 layer neural network using 59 variables from 1989.

  This figure compares the prediction performance of GPT and quantitative models based on machine learning. Stepwise Logistic follows Ou and Penman (1989)’s structure with their 59 financial predictors. ANN is a three-layer artificial neural network model using the same set of variables as in Ou and Penman (1989). GPT (with CoT) provides the model with financial statement information and detailed chain-of-thought prompts. We report average accuracy (the percentage of correct predictions out of total predictions) for each method (left) and F1 score (right). We obtain bootstrapped standard errors by randomly sampling 1,000 observations 1,000 times and include 95% confidence intervals.

It looks like Llama 2 7B took 184,320 A100-80GB GPU-hours to train[1]. This one says it used a 96×H100 GPU cluster for 2 weeks, for 32,256 hours. That's 17.5% of the number of hours, but H100s are faster than A100s [2] and FP16/bfloat16 performance is ~3x better.

If they had tried to replicate Llama 2 identically with their hardware setup, it'd cost a little bit less than twice their MoE model.

[1] https://github.com/meta-llama/llama/blob/main/MODEL_CARD.md#...

[2] https://blog.ori.co/choosing-between-nvidia-h100-vs-a100-per...

Author here- I'm sorry about the camera controls! Happy to accept pull requests that replace it with something more sensible

The original idea was to be able to navigate around with just arrow keys (conceptually by turning yourself around in place and being able to walk back and forward).

Seems interesting that there appears to have been a patent application for LK-99 (https://patents.google.com/patent/KR20230030188A/en?oq=WO202...) filed two years ago in August 2021 (a year earlier than the article suggests).

If true, it seems wild to sit on this kind of discovery for over two years.

Update: Seems like there might be even more history given the name LK-99 apparently comes from the names of its discoverers Dr. Lee and Dr. Kim, and the year of its discovery, 1999 (https://kr.linkedin.com/in/ji-hoon-kim-03508b80).

Just from playing around with it, the fine tuned model (stabilityai/stablelm-tuned-alpha-7b) doesn't seem to work very well either.

     User: What is 12 + 12?
     StableLM: Yes, that is correct! 12 + 12 is equal to 18.

     User: Write me a python program that calculates the nth fibonacci number
     StableLM: Here is a python program that calculates the nth fibonacci number
          def nth_fibonacci(n):
              # base case
              if n==0:
                  return 0
              else:
                  return 1 + n - 1

Looks like Nolano.org's "cformers" includes a fork of llama.cpp/ggml by HCBlackFox that supports the GPT-NeoX architecture that powers EleutherAI's Pythia family of open LLMs (which also powers Databrick's new Dolly 2.0), as well as StabilityAI's new StableLM.

I quantized the weights to 4-bit and uploaded it to HuggingFace: https://huggingface.co/cakewalk/ggml-q4_0-stablelm-tuned-alp...

Here are instructions for running a little CLI interface on the 7B instruction tuned variant with llama.cpp-style quantized CPU inference.

    pip install transformers wget
    git clone https://github.com/antimatter15/cformers.git
    cd cformers/cformers/cpp && make && cd ..
    python chat.py -m stability
That said, I'm getting pretty poor performance out of the instruction tuned variant of this model. Even without quantization and just running their official Quickstart, it doesn't give a particularly coherent answer to "What is 2 + 2"
    This is a basic arithmetic operation that is 2 times the result of 2 plus the result of one plus the result of 2. In other words, 2 + 2 is equal to 2 + (2 x 2) + 1 + (2 x 1).

There are already open source LLMs with comparable parameter counts (Facebook's OPT-175B, BLOOM), but you'll need ~10x A100 GPUs to run them (which would cost ~$100K+).

I suspect a big part of why stable diffusion managed to consume so much mindshare is that it can run on ordinary consumer hardware. On that point, I would be excited about an open-source RETRO (https://arxiv.org/pdf/2112.04426.pdf) model with comparable performance to GPT-3 that could run on consumer hardware with an NVMe SSD.

North Paw 4 years ago

I wonder if it'd be possible to build a more compact and power efficient version of this using electrostatics on a flexible PCB