Feature demo videos: https://magenta.withgoogle.com/mrt2
HN user
selvan
Founder of CheerArena (https://www.cheerarena.com) - TV grade Live channels on Youtube/Instagram/Facebook/Twitch
http://www.github.com/selvan
Build and play AI musical instruments on your laptop!. It is a live, interactive model that you can control with MIDI and audio, in addition to text.
Creating an AI native solution to manage workflows of my live streaming business (https://www.cheerarena.com)
Most workflow softwares are complex to extend & customize. Building an AI native, structured workflow orchestrator from scratch for agentic era.
As a starting point, have designed and implemented an AI native data store to store semantic linked structured input & output data of workflow steps/tasks. These structured input/output act as spec and guard rails for the workflow tasks.
Thanks. Fixed it.
From the blog " Gemini CLI spawns a new process within a pseudo-terminal in the background, leveraging the node-pty library...So how does this virtual terminal running in the background show up on your screen? Think of it like a video stream. Our new serializer takes a snapshot of the pseudo terminal at every moment—capturing every piece of text, every color, and even the cursor's position. These snapshots are then streamed to you, allowing you to see and interact with the terminal application in real-time. It's not just a stream of text; it's a live feed."
Terminal serializer code: https://github.com/google-gemini/gemini-cli/blob/main/packag...
Uses @xterm/headless npm package.
An MCP server exposes tools that a model can call during a conversation and returns results according to the tool contracts. Those results can include extra metadata—such as inline HTML—that the Apps SDK uses to render rich UI components (widgets) alongside assistant messages.
More: https://github.com/openai/openai-apps-sdk-examples?tab=readm...
May be personalization for narration ?. Different narration style, based on their own interest.
edit: Their demo video shows they allow learners to set different narration style based on their interest.
May be, we are couple of years away from experiencing patent free video codecs based on deep learning.
DCVC-RT (https://github.com/microsoft/DCVC) - A deep learning based video codec claims to deliver 21% more compression than h266.
One of the compelling edge AI usecases is to create deep learning based audio/video codecs on consumer hardwares.
One of the large/enterprise AI usecases is to create a coding model that generates deep learning based audio/video codecs for consumer hardwares.
Cursor - co-pilot/AI pair programming usecases.
Claude Code - Agentic/Autonomous coding usecases.
Both have their own place in programming, though there are overlaps.
Ship AI Agents as a web page :-)
CheerArena - Your Own TV Grade Live Channel on Youtube
Have created a real-time media mixing mobile app that helps to setup TV grade Live channel on Youtube/Facebook/Twitch/Instagram.
Our product scales from individual to institutions, camera in mobiles to network of cameras, indoor to outdoor sports and events.
Details: https://www.cheerarena.com/
Realtime mixing studio - https://play.google.com/store/apps/details?id=com.cheerarena...
Total PRs between Codex vs Cursor is 208K vs 705, this is an enormous difference in absolute PRs. Since cursor is very popular, how does their PRs is not even 1% of codex PRs?.
For simpler games, libraries such as raylib or lightweight opensource game engines such as Amulet (https://www.amulet.xyz/) / Love 2D are good fit.
Curious, what would be the motivation of the sellers to trade high inflation currency?.
It make sense for buyers as they want to move to stabe currency. But how about sellers?. What are they gonna do with the high inflation currency ?.
One motivation could be of very high margin due to high risk involved.
Fraud detection is code.
Not replacing hospitals/doctors, but replacing insurance companies.
Get started documentation on Multimodal Live API : https://ai.google.dev/api/multimodal-live
From the PDF - "One thing that should be learned from the bitter lesson is the great power of general purpose methods, of methods that continue to scale with increased computation even as the available computation becomes very great. The two methods that seem to scale arbitrarily in this way are "search" and "learning".
The second general point to be learned from the bitter lesson is that the actual contents of minds are tremendously, irredeemably complex; we should stop trying to find simple ways to think about the contents of minds, such as simple ways to think about space, objects, multiple agents, or symmetries. All these are part of the arbitrary, intrinsically-complex, outside world. They are not what should be built in, as their complexity is endless; instead we should build in only the meta-methods that can find and capture this arbitrary complexity. Essential to these methods is that they can find good approximations, but the search for them should be by our methods, not by us. We want AI agents that can discover like we can, not which contain what we have discovered. Building in our discoveries only makes it harder to see how the discovering process can be done."
Working on - "real-time conversations in rich video streaming". Have created rich video composition, mixing, streaming studio (http://www.thecheerlabs.com), working on to bring real-time conversations that can be mixed in real-time for streaming/recording.
Have used https://github.com/redotvideo
You may wanna move this post to "Ask HN:"
VS Code Editor which is based on Electron, is really fast, even with large codebase & many open tabs. Their monaco engine (https://microsoft.github.io/monaco-editor/) uses custom, virtual code processor that is optimized for surgically updating underlying DOM. It also uses WebGL + canvas rendering to show minimap of the file.
Similar approach (custom virtual processor) is leveraged by Google docs/sheets.
Canvas rendering may be the last resort when nothing worked.
Chat, Audio/Video Conferencing apps are other examples.
Nice, (code like) Refactoring meets speech-to-text
Microsoft and AWS would have a partnership with AMD/Intel for their GPUs, if those are capable and widely used as Nvidia's.
Microsoft has partnetship with OpenAI and also with Mistral.
Present convenience may not hold true in future. Nvidia knows that well.
Ad generation usecases are getting interesting with Video generation + Controlnet + Finetuning
https://nammayatri.in/open/ - Raid hailing service that uses beckn
https://ondc.org/ - P2P commerce network that uses beckn
Sounds like failure of financial due diligence. A proper due diligence could have uncovered the concealments.
Wouldn't the details audited by accounting firms?
super cool & informative !!
The documentation shows ffmpeg is used only for encoding/decoding, composition/animation/effects are driven by openCV/numpy/scipy/PIL
See: https://zulko.github.io/moviepy/getting_started/quick_presen...
The Gallery section has more advanced demos on vector/3D animations & audio mixing: https://zulko.github.io/moviepy/gallery.html
IMSTRONG | Bangalore, India | Full time | www.imstrong.co
ImStrong is making people around the world reimagine the way to stay fit. We're reinventing how anyone can exercise at home by delivering a live video-streaming fitness experience. ImStrong makes it easy and convenient for busy people and homemakers to access awesome trainer-led fitness classes from the comfort of their home at their chosen time.
We are looking to agument our team with couple of backend engineers. Our backend stack is built with Node.js and leverages various AWS service (EC2, RDS (MySQL), SES etc)
Polyglot programming experince with other backend technologies such as Phoenix(Elixir), RoR/Django, Golang is an added advantage.
Exposure to Event driven, Message passing, Communicating sequential processes, multi-threaded programming environments is a plus.
If interested, send your github profile to selvan@imstrong.co