Can I ask how your orgs measured productivity before and after remote work/RTO? I agree that employees would have a much easier time accepting RTO if they were given hard evidence it makes the business more productive, but it seems like almost all companies don’t do this and simply justify RTO “because we know best” which frustrates ICs.
HN user
karagenit
http://caleb.software
Yep, and following that with
and Chinese models are poised to take the lead.
makes it sound like the second part is a continuation of the first quote from the same source, but actually the second link is just some random person’s substack post from almost a year ago.
Tangent, but I’m really curious what country you’re from that uses the endonym for Göteborg but then also spells the capital of Denmark like Kopenhagen?
You will encounter business decisions you think are terrible but you still have to sell to your team. You cannot vent your frustration to the people you lead.
I actually disagree with this point a lot, as an IC. My manager shares his honest opinions with us, and I respect him more for it. It seems like the rest of the team feels the same.
I’ve had managers try to sell <obviously bad thing> as something good for the team, and it sucks. It feels like being gaslit. I think honest, open communication is a much better way to run a team. We’re all adults and professionals too; we can handle the truth.
It seems like based on e.g. [1] the article originally made some stronger claims about “no difference in bugs” that have been corrected. I agree that now it seems fine, but those edits might be why it feels like some commenters read a different article than you.
Curious how you’re handling prompt caching, as I understand it most LLM providers essentially inject tool definitions in the system prompt, so changing tools dynamically breaks the cache. This has been a big annoyance for me in a separate project; I currently just implemented my own tool-ish system that defines schemas in user messages and instructs the LLM to return matching JSON, but it’s less reliable than using the native tool calling + structured outputs available in the API.
True, but the article also says:
That's it. No rate limiting. No account lockout.
To me, if he confirmed that there’s no rate limiting on the auth API, this implies a scripted approach checking at least tens (if not more) of accounts in rapid succession.
I would highly recommend giving this excelled LessWrong post a read: https://www.lesswrong.com/posts/rarcxjGp47dcHftCP/your-llm-a...
It isn't a perfect fit, since the article talks a lot about the scientific method which doesn't apply super well to philosophy+math, but I think there are some strong parallels here.
Looks cool! Does it support prompt caching? And do you have any data showing how your latency compares to going directly to the model providers? I’m thinking about trying it out but those are my two big reservations.
What if the number of game critics just hasn’t increased, and since they can only play/review a fixed number of games each year due to time constraints, the number that they acclaim each year hasn’t grown? Not saying this is necessarily the case, just suggesting the possibility.
Yeah, I'd probably go that route if I wanted to scrape more than the handful of pages I needed for this project. I wonder if it would work on Zillow or not. Even my simple workflow of "click the next button, save the request in devtools, repeat" was suspicious enough to trigger a captcha-type "are you a bot?" challenge. Maybe it was just too many requests quickly like you mentioned, or maybe they're doing something more advanced like mouse movement tracking.
Do you have a citation for this? The only relevant study I saw on the LISTEN website was a preprint of a study showing data on self-reported post-vaccine symptoms, but didn’t really talk about causes or gene edits (Krumholz et al. 2023).
Yep, been waiting for the same thing. Maybe at some point it’ll be possible to use a large multilingual model to translate the dataset into one programming language, then train a new smaller model on just that language?
Hah at least it’s two/three letters, my personal site is at .software and most people get really confused by an eight letter TLD.
In terms of total petroleum products (including crude, gasoline, and diesel) the US has become a net exporter in the last few years.
In 2020, the United States became a net exporter of petroleum for the first time since at least 1949. In 2022, total petroleum exports were about 9.52 million barrels per day (b/d) and total petroleum imports were about 8.33 million b/d, making the United States an annual net total petroleum exporter for the third year in a row.
https://www.eia.gov/energyexplained/oil-and-petroleum-produc...
For the foregoing reasons, we hold that district courts must apply the traditional four factors articulated in Winter when considering the Board’s requests for a preliminary injunction under §10(j). We therefore vacate the judgment of the Court of Appeals and remand the case for further proceedings consistent with this opinion.
Correct me if I’m wrong, but it sounds like this is just about a temporary injunction to give the employees their jobs back while the actual labor case gets decided? The lower court used the two part McKinney test to grant the injunction, and the Supreme Court said no, you have to use the four part Winter test like other district courts, go back and decide again.
I have zero medical expertise, but were the two blood tests and an xray really necessary for what seems like a simple case of dehydration?
Anecdotally, I recently got a simple blood test to keep an eye on something that’s been borderline in the past. Unbeknownst to me, my doctor ordered nearly 30 things to be checked in the blood panel, at $50-$100 a piece. Luckily my insurance negotiated a much lower price and covered most of it, but I can’t imagine paying $2,000+ just to check e.g. your cholesterol level.
I think the problem is alignment of incentives: for a doctor, over-testing has little to no negative consequences, but under-testing could lead to guilt due to a patient dying, loss of reputation because something was missed, lawsuits, etc.
As a society, is this how we should be using our limited medical resources? What if we made fewer xray machines and instead spent more on something like cancer screening? How many net lives could we save?
Arguably because it’s a two sided marketplace (matching drivers to riders) the network effect is very strong, so to have any success you have to dump a ton of money into marketing to gain enough users to make the app viable.
According to the EIA domestic crude production is at an all time high: https://www.eia.gov/todayinenergy/detail.php?id=61523
And it’s equally annoying to see this comment about that comment about the paper every time. Recursion!
Very neat. Would it be more efficient to start with every known pangram (~35k) and calculate the score from there instead of scoring every possible set of 7 letters (~8B) and then filtering for pangrams?
I don’t know, it was a pretty simple demo (just audio recording, transcription, and summarizing) so would it really be worth faking? The only things the article mentions are the typo (which actually makes me think it’s more likely to be real than faked) and the audio mismatch (which he explains on twitter was simply because he clicked an earlier recording accidentally - makes sense, I would do a dry run before actually recording the demo too).
Yeah, I've been bitten by those quotes in the past too. I noticed recently that VSCode (probably other IDEs too) highlight these characters pretty clearly to help avoid these issues.
What’s up with North Dakota? There’s a spot in the middle of nowhere that has pollution levels on par with major metro areas.
Definitely some interesting results, though I wish they went into a little more detail on their methodology. For example, is lower case better because it’s harder to read Titles With Every Word Capitalized? Or is it because PR titles in ALL CAPS tend to get ignored for being annoying which biases the results against capital letters?
Also I don’t buy their time saving numbers at the end. According to their own numbers the average dev spends 1.34 * 590 = 791 days per year waiting on review?? At least at my company I’ve never had to sit and wait around for a review, even if thr PR is blocking some other work there’s always been unrelated stories to do instead.
O2 actually has a stronger narcotic effect than N2, so probably it would have the opposite effect if any.
I don’t think that’s right. Sometimes the cartilage will break which can sound similar, but broken ribs is a serious condition that can lead to internal bleeding and such.
Broken ribs are present in 3% of those who survive to hospital discharge, and 15% of those who die in the hospital
https://en.m.wikipedia.org/wiki/Cardiopulmonary_resuscitatio...
I totaled up the results from only the "crosschecked" CSV files, here's what I saw:
APC: 5928825
LP: 4731127
PDP: 4555334
NNPP: 1019045
I tried to manually verify about a dozen rows myself, half were so blurry/low res they were illegible but the ones that were legible were all correct.And for the "unsure" CSVs:
APC: 1308067
LP: 578482
PDP: 736183
NNPP: 513245
Also checked about a dozen, and all but one of them were wildly inaccurate so I wouldn't trust these much.Huh? The tax credit for small vehicles is the same for both at $7500. The $40k credit is only for large trucks, capped at 30% of the vehicle’s value, and with all of the same restrictions.