Someone from GitHub please give us the tea on what the hell is really going on over there.
HN user
dogman123
yup
neat. i'm pretty novice in the guts of this kind of stuff, but how does this work under the hood for blocking operators where they "cannot output a single row until the last row of their input has been seen"?
i think this is where spark shuffling comes in? but how does it work here.
https://duckdb.org/docs/stable/guides/performance/how_to_tun...
pretty awesome that the new yc website touts gary tan's work at palantir as a positive
"he was an early designer and engineering manager at Palantir (NYSE:PLTR), where he designed the company logo"
Is there a way to have the model inside of codex to make use of chunkhound instead of its “built in” search/explore functionality with rg? Whenever I spin up a new agent using xhigh thinking it spins its wheels for a while to get up to speed — wondering if chunkhound can make this process faster.
im building a backtesting framework that uses polars as the underlying engine instead of the more traditional pandas. i've learned a ton.
hell yea brother
One thing that I never really see mentioned in these types of articles is that a lot of DuckDB’s functionality does not work if you need to spill to disk. iirc, percentiles/quartiles (among other aggregate functions) caused DuckDB to crash out when it spilled to disk.
I’m often surprised how little people talk about the iOS Orion browser on here and it’s ability to let you use both Firefox and chrome extensions. I’ve been using it for a while now and it’s been great. It’s a little bit buggy sometimes, but nothing that would make me switch.
My dad was diagnosed with multiple myeloma 2 years ago. His bone marrow transplant failed (frequent first line of defense) and he just finished CAR-T therapy a couple months ago. The initial side effects from the treatment were _bad_, but everything is looking good right now. CAR-T is really mindbogglingly insane cyberpunk stuff.
can someone ELI5 how these proof-of-work captchas work under the hood to detect whether i'm a bot or not?
This could be incredibly useful for me. Currently struggling to complete jobs with massive amounts of shuffle with Spark on EMR (large joins yielding 150+ billion rows). We use Glue currently, but it has become cost prohibitive.
I think I'd need gemini 2.5 built into the trial to try this out. It's crazy how bad claude has become.
I thought the era of gratuitous cursing was behind us. Oh well.
Got it, thx!
Curious on this: "Atuin's sync keeps my history on all of them"
I just checked on their GitHub and it says "Additionally, it provides optional and fully encrypted synchronisation of your history between machines, via an Atuin server."
So you trust all of your shell commands to be stored on a server that you don't control?
Maybe I'm missing something here.
The Barnes Foundation is one of my favorite museums in the country. There isn’t anything else quite like it.
This seems to violate the spirit of Show HN. Assuming it’s a marketing gimmick for Feeld or something like that.
Who’s the “I”?
“Designed by Cozy Ventures” … “We're a company that creates advanced digital solutions for early-stage startups.”
It allows Firefox and chrome extensions!
Orion on iOS has been life changing for me.
fwiw, i had it do something _far_ more complex that i am currently dealing with at work and it performed perfectly in my few test cases. i see very heavy use of this tool in my future. just figured i'd give a shot about the quickstart not functioning as planned :)
i tried the reddit quickstart example in the repo and it seemed to be incapable of completing the task.
sqlmesh execution engine + cloud resource provisioning
provision spark on emr or duckdb on beefy ec2 -> run sqlmesh -> wipe resources.
i'm still in MVP phase of revamping my company's current data platform, so maybe there are better alternatives -- which i'd love to hear about.
that's great to hear. it mirrors my observational experience from being in their slack channel. i'm aware of the technical risks of being an early adopter of a product like this, but i must say part of me is excited to be on board early to help to shape it from a user perspective. i'm still not totally bought in yet (still in mvp phase) but our use case as we scale almost requires multi-engine execution (athena, spark on EMR, duckdb) and it doesn't seem like anyone is doing it better.
i'm working on a project to do this with iceberg and sqlmesh executed via airflow at my job. sqlmesh seems really promising. i investigated multi-engine executions in dbt and it seems like you need to pay a lot of $$$ for it (multi-engine execution requires multiple dbt projects) and is not included in dbt core.
RIP Partiful
Did anyone else get the seemingly _actual_ NordVPN that clicked through to NordVPN? Wondering if they sponsored this.
Hell yea, lick that boot !