I'm with you. I haven't materially been more satisfied with the code or reasoning with 4.8 than I was with 4.7. But I'm also not vibe coding, I'm reviewing all of the output. Maybe 4.8 has been making fewer mistakes that I otherwise would have corrected on, but I was perfectly happy going through a few iterations with 4.7 to get it over the finish line. This trend just has me startled and I'm now realizing that my workflow will need to shift to open-weight models very soon. They're cranking the costs and there's no way I can get my employer to cover what's apparently become $2k a day in token use.
HN user
pvankessel
Senior director of data science and AI at WorkMoney, a non-profit dedicated to helping working and middle-class Americans. Former senior researcher at Pew Research Center.
Anecdotally this tracks what I've felt over the past month, though I haven't rigorously quantified it. I've just been burning through my quota considerably more quickly than I was a week ago. Hitting limits I didn't before. I hit my weekly max yesterday, it resets tomorrow so I asked my admin to add $50 to my overage limit so I could bridge the gap. Burned that in an hour, I was astounded. Two weeks ago that much bought me two days of work. I asked for another $10 so I could simply have Claude dump handoff notes that could be picked up by Codex, and got through 4 of the 10 agents I'd had running before running out. When it resets tomorrow I'm setting the default back to 4.7, my strong suspicion is 4.8 was designed to lighten the load by burning more quota (not necessarily tokens) on the backend. I don't understand the mechanism but they're clearly putting the squeeze on power users. Curious what others have experienced.
This view of the world puts everything on the individual. It might be worth reading up on structuralism to balance that perspective out a bit. I'm somewhere in the middle of the two extremes myself, but surely one must acknowledge that there are larger systems at play that can constrain an individual's ability to "optimize".
The automation one is so true! When I first deployed a huge job to MTurk, with so much money on the line I wanted to be careful, and I wrote some heuristics to auto-ban Turkers who worked their way through the HITs suspiciously quickly (2 standard deviations above the norm, iirc) - and damn did I wake up to a BUNCH of angry (but kind) emails. Turns out, there was a popular hotkey programming tool that Turk Masters made use of to work through the more prized HITs more efficiently, and on one of their forums someone shared a script for ours. I checked their work and it was quality, they were just hyper-optimizing. It was reassuring to see how much they cared about doing a good job.
I used MTurk heavily in its hey-day for data annotation - it was an invaluable tool for collecting training data for large-scale research projects, I honestly have to credit it with enabling most of my early career triumphs. We labeled and classified hundreds of thousands of tweets, Facebook posts, news articles, YouTube videos - you name it. Sure, there were bad actors who gave us fake data, but with the right qualifications and timing checks, and if you assigned multiple Turkers (3-5) to each task, you could get very reliable results with high inter-rater reliability that matched that of experts. Wisdom of the crowd, or the law of averages, I suppose. Paying a living wage also helped - the community always got extremely excited when our HITs dropped and was very engaged, I loved getting thank yous and insightful clarifying questions in our inbox. For most of this kind of work, I now use AI and get comparable results, but back in the day, MTurk was pure magic if you knew how to use it to its full potential. Truthfully I really miss it - hitting a button to launch 50k HITs and seeing the results slowly pour in overnight (and frantically spot-checking it to make sure you weren't setting $20k on fire) was about as much of a rush as you can get in the social science research world.
I got this one a few months ago and have been running it in my basement directly under my living room, separated only by the floor and a bit of insulation. Can't hear it at all. It's been working well and it's a fun low-investment hobby. I live on a glacial moraine so there are lots of unique rocks in my backyard, and my son enjoys digging for them. https://a.co/d/4HSnVVX
Hard disagree. I gave up a top-10 engineering scholarship and switched to liberal arts largely because my entire curriculum was predetermined in the former. Five courses in calculus and two slots for electives in your entire undergraduate schedule - that doesn't teach you how to think. Political philosophy, symbolic logic, comparative history, econometrics - having the freedom to explore and dabble and push yourself into new ideas instead of being fast-tracked into a pipeline, that's how you learn how to learn. And the "difficulty" is entirely what you make of it. Sure, if you show up to college and want to major in anthropology and put no effort in, you get nothing out. But I saw very quickly that with absolute unfailing effort applied to my engineering degree, I was still going to get exactly one and only one thing out of it. The liberal arts gave me a cornucopia of possibility. I've gone on a human trafficking sting op with the FBI, I've presented my research at the White House, I've been cited by the Pope - that's all wild shit that an engineering degree never would have enabled. Breadth of learning and soft skills matter. I'd be a shell of a person today if not for my liberal arts education. I owe everything to it, and the constant condescension towards non-STEM education in tech would frustrate me more if I didn't run laps around my peers.
Well that's kind of my point - liberal arts and humanities set you up with a very versatile baseline. With a proper education in those disciplines you learn how to think, and that's applicable to a wide range of fields. The woman I dated in grad school at UChicago studied war history and wound up being an analyst for a prominent wine auctioneering firm as a key researcher. My master's thesis was on the meaning of life, and now I'm running data science at a non-profit. So many of my fellow liberal arts grads have gone on to do incredible things entirely unrelated to their chosen subject of study.
Oh I agree with you on that wholeheartedly. I think our society would be substantially healthier if we required civics, philosophy, economics, etc in high school. But if we're already struggling to have evolution taught in schools and we have state boards of education removing references to the slave trade and founding fathers from history curriculum (https://www.theguardian.com/world/2010/may/16/texas-schools-...), expanding liberal arts in public education is a non-starter. Hell, half the country would love to see it wiped from post-secondary education. Best I figure we can do at this point is defend the idea itself to the extent we can - for instance, in Hacker News threads where the liberal arts are being dismissed as an unnecessary lesser-than academic pursuit.
Except many STEM graduates are having a harder time finding jobs right now than liberal arts and humanities majors: https://www.newyorkfed.org/research/college-labor-market#--:....
For what it's worth, I have enjoyed a very successful career in data science and software engineering after taking some AP STEM courses in high school, followed by three liberal arts degrees. Many of the best engineers I've known have had similar backgrounds. A good liberal arts education teaches one how to think and learn independently. It's not a substitute for a highly-specialized education in, say, molecular biology, but it provides a really solid foundation to easily pick up more logic-derived technical skills like software development. It's also essential for an informed citizenry and functional democracy.
Curious about this, is there actually a canonical explanation in the trilogy somewhere?
Heh, that's actually pretty compelling - it sounds like a darker twist on the plot of a Stargate episode I recently rewatched: https://en.m.wikipedia.org/wiki/Revisions_(Stargate_SG-1)
Oh this sounds like exactly like what I've been looking for, can't wait to give these a try - many thanks
Are there any models out there for cleaning up an image, not just upscaling? I have a bunch of old photos taken on early low-res point-and-shoots that have JPEG artifacts etc and this seems like something a modern model could easily be fine-tuned to resolve, but every few months I look around and have yet to find anything
Here's a study I did a few years ago that broke down popular YT videos into more granular categories, might be of interest: https://www.pewresearch.org/internet/2019/07/25/childrens-co...
They actually held out for a couple of years after Facebook and didn't start forcing audits and cutting quotas until 2019/2020
This is such a clever way of sampling, kudos to the authors. Back when I was at Pew we tried to map YouTube using random walks through the API's "related videos" endpoint and it seemed like we hit a saturation point after a year, but the magnitude described here suggests there's a quite a long tail that flies under the radar. Google started locking down the API almost immediately after we published our study, I'm glad to see folks still pursuing research with good old-fashioned scraping. Our analysis was at the channel level and focused only on popular ones but it's interesting how some of the figures on TubeStats are pretty close to what we found (e.g. language distribution): https://www.pewresearch.org/internet/2019/07/25/a-week-in-th...
First thing I thought of, 7 years old and more prophetic by the day: https://m.youtube.com/watch?v=7Pq-S557XQU
It's only a matter of time. Five years, twenty - we need to prepare for mass unemployment.
Lol, not yet it seems
I mean, that's definitely a huge factor, but there IS a limit to how much Amazon can pay their hundreds of thousands of workers and still remain competitive - and I have NO idea where that threshold is, which is why I'm fascinated to see how this shakes out. Either Amazon workers start getting treated better which is great, or the company collapses and turns into an unprecedented case study of what not to do. Either way, I'm grabbing some popcorn!
Oh don't get me wrong, much of it is definitely the result of their policies and could have been avoided if they hadn't treated their labor force like a discardable, consumable resource for years. I just thought it was interesting to see something pop up in the news that echoes what I got a glimpse of a few months ago - that Amazon is starting to recognize that they have a problem on their hands (finally). Just thinking about it in dispassionate scientific terms, it's a fascinating and unique problem - it might be too late for them to pivot and shift back towards "sustainable" practices, and if they fail I'm certainly going to enjoy the schadenfreude, but I'm really curious to see how they attempt to deal with all this.
Not surprised to see this pop up. I interviewed for a position with Amazon a few months back to lead up a new research program to determine how they can better recruit and retain hourly workers. I had no intention of taking the job, but was curious so I took the interview. What stood out was just the sheer scale at which they're operating - they're literally up against the constraints of domestic labor supply. I have plenty of strong opinions about how they treat their workers and have no desire to work for such a company, but I was surprised to find that I did sympathize with them to an extent - it's not just about offering better pay and bathroom breaks, they're also on the verge of exhausting the viable labor market. I wish whoever took the job the best of luck - I hope that they're taking the research effort seriously and it's not just performance art.
I love Vaillant's work. His and many other studies definitely point to the importance of relationships as determinants of happiness. We found evidence of the same a few years ago - people who mention their friends and their spouse/partner rate their lives more highly: https://www.pewresearch.org/fact-tank/2018/11/20/americans-w...
What I'm really interested in, though, is not just what the formula for happiness is - it's whether people are actually tapping into the key components of that formula. That is, knowing that strong and healthy relationships bring meaning and satisfaction to life, to what extent do people have and prioritize those relationships?
The fact that our latest research shows family and friends ranking highly around the world suggests that many people are doing just that, but there are other places where there might be room for improvement. For example, compared to four years ago, Americans are half as likely to mention their spouse or romantic partner when describing where they find meaning in life: https://www.pewresearch.org/fact-tank/2021/11/18/where-ameri...
We started off by trying LDA and NMF, but the topics were too messy so we wound up switching to CorEx (https://github.com/gregversteeg/corex_topic), which is a semi-supervised algo that lets you "nudge" the model in the right direction using anchor terms. By the time our topics started looking coherent, it turned out that a regex with the anchor terms we'd picked outperformed the model itself. This case study was on a relatively small sample of relatively short documents (~4k survey open-ends) but for what it's worth, we also tried to use topic models to classify congressional Facebook posts (much larger corpus and longer documents) and the results were the same.
Overfitting is certainly part of the problem - in one of my earlier posts I talk about "conceptually spurious words," which are essentially the product of overfitting - but the more difficult problem is polysemy. I'm sure there are ways to mitigate that - expanding the feature space with POS tagging, etc. - but ultimately I think the solution is to simply avoid using a dimensionality reduction method for text classification. Supervised models are clearly the way to go - even if those "models" are just keyword dictionaries curated based on domain knowledge.
Wow, I had no idea. 1m was enough to work with, but... wow. Thanks for the reference.
To clarify - the current API has a limit of 10,000 "query points" per day for new API keys (most endpoints cost 1-5 points per query). It used to be 1 million; they've since throttled everyone down and started forcing audits. 10k is still something, but it certainly doesn't allow large scale research.
Would love to see what you come up with, will stay tuned!
As for the API restrictions, they aren't advertising it but about a year ago they started warning users about forthcoming extensive audits to maintain access, and about six months ago they started reducing access for API keys if you stopped maxing them out for a day or more. Our last API key got shut down for good a couple of weeks ago. We're going to try to fill out the form and get our access reinstated, but I'm not sure how willing they'll be to allow access for research. The form seems intended for client-facing apps. But who knows - Facebook/CrowdTangle/Twitter have been very supportive of legitimate research initiatives, I'm hoping YouTube follows that trend!
This is really great - and kudos for providing methodological details and code! I love seeing this kind of large-scale descriptive research, it's a real bummer that YouTube is starting to close up access to their API. We did a similar kind of analysis looking at videos posted by popular channels last year, including some analysis of keywords that boosted views - figure you might find it interesting (and I'd love to see if our findings hold up with your dataset!) https://www.pewresearch.org/internet/2019/07/25/a-week-in-th...
Full report is here, in case it interests anyone: https://www.pewinternet.org/2019/07/25/a-week-in-the-life-of...
That's actually why we had to come up with our own topic typology in this report - some of the category tags that YouTube provides were too broad to be useful ("TV Shows") and others were way too specific ("Music of Latin America")