HN user

ibuildthings

75 karma
Posts11
Comments38
View on HN

Talk on optimizing matrix multiplication with Triton kernels, focusing on low-bit processing and efficient quantization for high-performance AI models.

Aana SDK is an open-source toolkit for building cutting-edge multimodal AI applications: https://github.com/mobiusml/aana_sdk

It addresses key challenges in multimodal AI development:

- Managing diverse inputs - Scaling Generative AI apps - Ensuring extensibility

Built on Ray for seamless scaling, Aana offers a unified framework for multiple data types, easy integration with popular ML frameworks, and a modular architecture.

Aanaphi-2 3B 2 years ago

Currently leading in LLM benchmarks among the 3B categories of models

We are releasing new 2-bit Mixtral models. These ones use a mixed HQQ 4-bit/2-bit configuration, resulting in a significantly improved model (ppl 4.69 vs. 5.90) with a negligible 0.20 GB VRAM increase.

Base: https://huggingface.co/mobiuslabsgmbh/Mixtral-8x7B-v0.1-hf-a...

Instruct: https://huggingface.co/mobiuslabsgmbh/Mixtral-8x7B-Instruct-...

Shout-out to Artem Eliseev and Denis Mazur for suggesting this idea ( https://github.com/mobiusml/hqq/issues/2 )

Sharing our work on model quantization.

- Blog: https://mobiusml.github.io/hqq_blog/ - Code: https://github.com/mobiusml/hqq - Models: https://huggingface.co/mobiuslabsgmbh/

No data calibration needed, extremely fast , works on both language and vision models!

* Why does it matter? Quantization significantly reduces GPU memory requirements but degrades the quality of the models. Having faster and more accurate quantization methods is extremely valuable for the ML community.

* Approach: Sparsity-based error formulation between the original weights and their dequantized version. We used a Half-Quadratic solver to derive a closed-form solution that is 100x faster than backprop via Pytorch's Autograd.

* Quantization speed: ~ 1 minute for Llama2-13B ~ 4 minutes for LLama2-70B (over 50x faster than GPTQ)

* Findings: - Larger models quantized to 3/2-bit outperform smaller full-precision models with similar or lower memory requirements. - Successful 2-bit quantization requires a lower group-size (e.g., 32 or 16) and compression of both the zero-point and the scaling factor for lower memory usage.

While we acknowledge our view might be slightly biased, we genuinely believe that our work will significantly benefit the open-source software (OSS) machine learning community. Code and model are in Apache permissive license.

Github repo should be visible now.

It is not distilling the model, it is reducing the model weights on the fly and uses LoRA for training/fine-tuning. After the training phase, we explain how to merge the LoRA weights with the pruned weights to achieve faster inference speed

I'm sharing a blog post https://mobiusml.github.io/low-rank-llama2/ on our approach to pruning the Llama2 model by leveraging low-rank structures.

In a nutshell, we've managed to reduce the model's parameter count by up to 50%, double the training speed, and increase inference speed by 1.25 times.

For those interested in the technical details or looking to replicate our results, the code is openly available for community use and contributions

While agreeing to the general principle, incentive structures are wired quiet differently in academia vs end consumer oriented gig/service industries.

Publications ( number, when, where, citations) is the primary currency/value in which one is judged within the peers in academia, and reputation outside the immediate academic community has a much lower weight. Whereas for online market places solid revune is the first priority and then comes reputation ( which is a means for the higher reveune). In academia it is the reverse, with reputation ( in a small clique) being the primary motivator, and funding being the means to gather it.

First principles of doing a PhD and taking up an industrial jobs are quite different, which this article sidesteps. I am talking from the perspective of someone who did a PhD, postdoc and migrated to be a founder/CEO.

A PhD system trains you to think about unsolved problems in an given domain deeply with a larger time runway. The end goal is not a tangible product that reaches millions of people, but rather a set of ideas that can take a crack at the unsolved problems in your field in a novel way. A good work should inspire others in the field, and eventually a larger audience to pick them up and expand and build on top of it. To give a small example, a majority of the fundamentals of machine learning was charted out by many, many PhD works over the last 40 years. Implementing a linear classifier is 2 lines of code in 2018, but many Bothans died to bring us this information :-) .

The goals of industry are more immediate. Expect for a privileged few research labs in industry, your work is expected to be monetized, and rightly so. The goal is for you, if you run the business, else your management team to first figure out a problem of high relevance and monetary value. Build products/solutions for that problem, that can be used by someone who is less versed/ambivalent of your technical solutions. Efficacy of solving that particular problem often defines the merit of your contribution.

The fundamental of choosing the PhD or industry should be taking stock of what kind of contribution you want to make as an individual. If it is a few set of ideas to science, which on a later date might become something fundamental in our understanding of the world, then PhD is a good path. If it is a set of contributions towards a product/solution that eases the pain of many users then go into the industry first.

Author here. Just to clarify, motive of this work is to ease curation, with the massive amount of content being created; but by no means an attempt at creativity or originality.

The work is not at all contradictory to Adorno, especially in the sense that it is explicitly trying to as non-reductionist as possible, and assuming notion of aesthetics is a dynamic entity .

There is a finite pattern in the dataset; more interesting, it has its interesting share of subtleties ( for example, as opposed to a image classification problems), and the technological question is whether we can capture these.

But there is another interesting data question. For our work, we curated our training set with the help of expert curators. But the dataset itself is a metamorphising entity; i.e. it is subject to revision ( it is a continuous process for us at the moment), but more interestingly it is a chance for open debate between our curators. In some sense, technology allow to codify and challenge our notion of aesthetics ( especially with the evolution in our training sets) at a given point of time.

The author here. I used the term "understanding", not as in machines understanding the images, but more as scientific attempt in understanding aesthetics. ( <snippet from the text>"empowering me to develop systems for understanding images from a computational and scientific perspective"</snippet ends> ).

I agree with you completely about "bad industry-focused research". It serves no end. My question is that, is it just a reflection of mediocracy and gaming/dishonesty being everywhere, including academica ?

In this case, please trying to take shortcuts towards their goal of improving the amount of publications and grant approvals. There is reward in the system for this kind of behaviour. You being a reviewer/jury position unfortunately do not have the luxury of a filter.

Is there a way to early catch this , by looking at past trends ?

Slight detour is that this is one of my rationale for spending time reviewing papers for journals/conferences. In average, only 10 to 20% of papers I review really stands out or appeal to me, which is correlated to acceptance rate of a top journal/conference.

Pushing away from color lines is very easy. For negative values of lambda in Eq.1 of the paper, (i.e. reverse the cost for color lines ), the optimization tries to push the puzzle shapes away from the color lines. We had tried a few in this configuration, but personally my co-authors and I liked the puzzles generated by the scheme of adhering to the color lines.

Hardness/fun factor in puzzles is a matter of personal taste. Hence the ability to personalize is very interesting. Not everyone like to solve 10000+ piece puzzles , nor color line adherence, but people invest time and effort in solving them.

One of the authors here. Since it is a optimization, the difficulty can be controlled as a parameter ( the lambda parameter in Eq.1 in the paper ).

But you are right, some of the puzzles can be super-hard ( for example, the Seurat puzzle ) that we used to joke between ourself to name our paper "taking the fun out of puzzles".

Personally, what was fascinating for me is the shape of the puzzle curve it produced. Most of the common puzzles are grid based (i.e. four neighbours - up , down, left, down ). But in this scheme, there can be strange neighborhood pieces, with even stranger shapes.

The site has an interesting history. The former Stadtschloss suffered serious destruction during WWII ( https://en.wikipedia.org/wiki/City_Palace,_Berlin ) and the Communist East tore it completely down to build the Palace of the Republic https://en.wikipedia.org/wiki/Palace_of_the_Republic,_Berlin and acted as the hub of DDR government. Once the wall fell, and DDR disintegrated, in 2008 DDR's Palace of Republic was almost completed demolished, and work on new Stadtschloss which very much resembles the original commenced.

I do empathize with the original article a lot. I used to have/still have a strong fear of failing, especially in intellectual tasks. According to my own introspection this is primary angst that caused/causes me to procrastinate. There are two major references I often go back when I find myself paralyzed.

One being an advise I got from one of my PhD advisors: All creative tasks might appear that it requires enormous amount of courage and effort. But usually it is more like a kitchen sink heaped with a lot of unwashed dishes. Chances are that once you wash one dish, you will end up cleaning the full lot; and you often get a strange form of pleasure while you are performing the task.

The other one is this essay http://www-rohan.sdsu.edu/~psargent/Mills_Intell_Craft.pdf on intellectual craftsmanship by Wright Mills. I do now a days actively collect memories of pure immersion and pleasure I experienced while my craft got exposed and exploited to its potential. The thought of me improving as a craftsman, coupled with these memories is a powerful self motivator to me. The shit feeling I gets when I waste my time is another reference. One of the potent lessons was also that craft can be improved only by dedicating time ( which is pleasurable); and by disassociating the end result and fears. The toughest part is to replay this logic while I find myself slipping into vortex of non productivity, but that is something I can work on and probably in my control.

Falsification and Incompleteness are two different things. Since we reason about physicals system using the language of mathematics/logic; it has be based on certain axiom which cannot be proved or disproved ( Godel's incompleteness theorem ). Though this renders certain statements inside physical theorem non-provable ; it certainly does translate to every claim made by a proposed theory. Further many aspects of physicals systems can be disproven experimentally. ( It is still in active debate if Mathematics should treated as science per se : http://en.wikipedia.org/wiki/Mathematics#Mathematics_as_scie... )

While the pen falling from a desk do point out to the incompleteness ( non-Godel sense) of the standard model, which is widely accepted ( http://home.web.cern.ch/about/physics/standard-model : last paragraph ), it does not falsify it. Science is full of open holes, and no one knows ( my bet is against) that it will be completely patched up; but it is the best form of reasoning we have in understanding things, and its ongoing goal is to seek explanations that with the least amount of uncertainty possible.

As i see, the current science is more rigorous because people are producing lot of crap

Good, grief. No!!! It is a way of managing uncertainty and saying something with a precision that is available at a given point of time.

How many discoveries are being overridden by new discoveries coming from future ?

This is beauty/and USP of science. Every scientific proof is always open for scrutiny and revision in light of new data or discovery ( tenants of falsifiability kick in here). That is, it tries hard NOT to be dogmatic by being provisional. For example, science says that we are confident Higgs Boson exists "accounting for one-in-a-million chance on the contrary" ( 5-sigma).

Let me flip your argument on the converse; success rate at which we could make ground breaking theories [ like evolution, theory of relativity , uncertainty principle ] ( which is standing the test of time for extended period of time) using the scientific method is sheer staggering and amazing. The methodology has accelerated our progress and understanding by leaps and bounds which no alternate system has managed to do so, so far!

The amount of data accessible to the people in the past is a lot more when compared to current.

I lost you completely here. Can you please elaborate and the rest of the paragraph. ( My belief: If you take 20 random guesses; one of them turned out to be true; it is more likely to be a coincidence than a mystical insight. If on the contrary, the Monte Carlo filter I routinely simulate might just be the most insightfully entity I have encountered ).

The division between religion/science is very small

Epistemologically they are apples and oranges! Falsifiability is not applicable to religion nor is it is provisional and routinely advocates absolute (and imho dogmatic) reasoning!

There is no such thing as modern science.

The way we approach and do science has evolved drastically ( http://en.wikipedia.org/wiki/History_of_scientific_method ). For example empirical falsifiability which is one of the primary tenant of modern science is less than 100 years old, but forms an essential part on how we do science now a days.

Parent comment's point being; while we may be trying to understand the same principle/phenomena, not only the data available to thinkers that time was very sparse compared to the present; but also the level of rigour applied was of significantly lower standards. While there might be scattered scientific truth in the vedas ( or any other religious document) ; it is insolent to believe that it is good reference manual for scientific knowledge.

Coming from an academic world, the primary currency in my network is the number of publications/citations. Faculty/postdocs/Phd's salary in most universities are not worth bragging about; but it is not uncommon for people to put in regular 60+ hours week to get that extra publication so that you can stay ahead of the curve ( or at-least catch up )to your peers. From my perspective ( of course the domain bias exist) human ego is as powerful a motivator ( and we invariably have a desire to be ahead of people in our immediate network! ).

The fact that Switzerland is a really rich country, with the wealth distributed not so unfairly, makes this experiment even more interesting ( compared to the communist style takeover which happened in past in then poor nations including Russia or China).

To summarize, one of the things that makes capitalism work is competition; and money necessarily is not the only thing ( and might not be the primary thing) that we compete for!

While I agree with the main sentiment of this article, I have a nitpick on the main caption image. For me there is a significant loss of information, when an app summarizes something as non-linear and hard-to-predict as weather into boolean statements. So I am not sure if the umbrella app is solving weather reporting problem correctly.

A slightly related project is this one: http://www.hole-in-the-wall.com/

I remember getting my first PC at the age of 10, and a feeling of empowerment it brought me. Be it writing my first program that did sometime substantial, or playing Prince of Persia ( and clearing those levels ); what was substantial was the feeling that I am in control of this machine, and I can add/mould things as I want it! This sense of empowerment ( and later the ability to make use of the technology ) was possible only due to the latent fun quotient associated whilst using it (and it is best unspoiled by lack of supervision/or other's deciding how I should use the technology ).

This is very pertinent in an Indian context, where most act for "empowerment of poor" is coupled with the assumption that "poor are incapable of making their own decision".

On the contrary, I think these are symptoms of machine learning coming of age [or getting industrialized]. I am not sure "short sighted commercial ends" are the main motivator for likes of Prof. Hinton to shift, but rather the availability of vast amount of data and opportunity to understand them in a scale which was not previously possible.

To bring out the real magic out of techniques like deep learning ( http://en.wikipedia.org/wiki/Deep_learning ), availability of large training sets and the infrastructure required to crunch them are a pre-requisite. Once you have that, it is turning out to be a different ball game all together http://deeplearning.net/2012/12/13/googles-large-scale-deep-.... It turns out that groups like google research are the ones at present which have access to such dataset and infrastructure.

I also predict the reverse shift to happen within few years, once the interesting fundamental research problems has been tackled such people might move back to universities. If that happens, that is indeed a healthy process of academica and industry supplementing each other.