HN user

BUFU

646 karma
Posts42
Comments14
View on HN
github.com 15d ago

Qualcomm acquires Nexa AI, open-sources GenAI runtime for Hexagon NPUs

BUFU
5pts1
nexa.ai 9mo ago

We Ran GPT‑OSS 20B Local on a Phone

BUFU
1pts0
huggingface.co 9mo ago

Qwen3-VL-30B-A3B-Instruct and Thinking

BUFU
6pts0
nexa.ai 10mo ago

New Engine to Run SOTA AI Models on Qualcomm NPU Across Phone, PC, Cars, and IoT

BUFU
1pts0
huggingface.co 11mo ago

First vision language model built off Open AI GPT-OSS

BUFU
3pts0
huggingface.co 11mo ago

First Multimodal AI Model Designed for NPUs

BUFU
1pts0
nexa.ai 11mo ago

Nexa AI Blogs

BUFU
1pts0
www.liquid.ai 11mo ago

LFM2-VL: Efficient Vision-Language Models

BUFU
3pts0
www.youtube.com 11mo ago

Stanford CS336 Language Modeling from Scratch

BUFU
19pts0
ollama.com 11mo ago

Ollama's new app

BUFU
560pts284
claude.ai 1y ago

You can now connect a directory of apps and tools to Claude with one click

BUFU
2pts1
news.ycombinator.com 1y ago

Ask HN: What tools have you tried to run AI locally on mobile?

BUFU
2pts0
github.com 1y ago

A C++ library to efficiently run Gemma-3N across various platform

BUFU
5pts0
techcrunch.com 1y ago

The Trump-Musk feud has been great for X, which jumped up the App Store charts

BUFU
6pts1
openai.com 1y ago

How we’re responding to The NYT’s data demands in order to protect user privacy

BUFU
284pts324
twitter.com 1y ago

ChatGPT Deep Research connects cloud apps

BUFU
1pts0
yummy-fir-7a4.notion.site 1y ago

Local AI generates highly realistic dialogue from a transcript

BUFU
3pts2
www.theverge.com 1y ago

Shroud's Spectre Divide and its developer are shutting down

BUFU
1pts0
www.theverge.com 1y ago

What Went Wrong with Skype?

BUFU
4pts1
www.anthropic.com 1y ago

Anthropic's Recommendations to OSTP for the U.S. AI Action Plan

BUFU
3pts0
nexa.ai 1y ago

Quantized DeepSeek R1 Distill Models with Original Model Accuracy

BUFU
2pts0
neuralmagic.com 1y ago

Multimodal Model Quantization Support Through LLM Compressor by Neural Magic

BUFU
1pts0
huggingface.co 1y ago

DeepSeek-R1-Distill-Qwen-1.5B Surpasses GPT-4o in certain benchmarks

BUFU
39pts17
nexa.ai 1y ago

NexaQuant: Llama.cpp-Compatible Model Compression with 100%+ Accuracy Recovery

BUFU
3pts1
arxiv.org 1y ago

Meta's new Video Understanding Multimodal Model used Qwen model for training

BUFU
7pts1
github.com 1y ago

Llama.cpp Now Supports Qwen2-VL (Vision Language Model)

BUFU
155pts50
nexa.ai 1y ago

OmniAudio-2.6B: Fastest Audio Language Model for Edge Deployment

BUFU
2pts1
moondream.ai 1y ago

Moondream 0.5B: The Smallest Vision-Language Model

BUFU
14pts3
arxiv.org 1y ago

ShowUI: One Vision-Language-Action Model for GUI Visual Agent

BUFU
2pts0
neuralmagic.com 1y ago

What happens if we remove 50 percent of Llama?

BUFU
231pts132
Qwen3-VL 10 months ago

The open source models are no longer catching up. They are leading now.

On a 2024 Mac Mini M4 Pro, Qwen2-Audio-7B-Instruct running on Transformers achieves an average decoding speed of 6.38 tokens/second, while OmniAudio-2.6B through Nexa SDK reaches 35.23 tokens/second in FP16 GGUF version and 66 tokens/second in Q4_K_M quantized GGUF version - delivering 5.5x to 10.3x faster performance on consumer hardware.

Blogs for more details: https://nexa.ai/blogs/OmniAudio-2.6B

HuggingFace Repo: https://huggingface.co/NexaAIDev/OmniAudio-2.6B

Run locally: https://huggingface.co/NexaAIDev/OmniAudio-2.6B#how-to-use-o...

Interactive Demo: https://huggingface.co/spaces/NexaAIDev/omni-audio-demo