Is there some sort of a leaderboard for this test? Like if you'd give each of Opus 4.8 and GPT 5.5 a score out of 100, what would the scores be?
HN user
ammar_x
My website: ammar-alyousfi.com
Linkedin: linkedin.com/in/ammar-alyousfi
Data scientist.
Absolutely! We need new and better benchmarks like this.
I have a question: why not use the maximum available reasoning on each LLM? For example, I see that Opus 4.7 at `max` reasoning but Sonnet 4.6 at `high`. Wouldn't it be a fairer comparison if all were at max?
I usually do this for complex features:
- Opus 4.7 writes the code - I make GPT-5.5 in Codex to review it (given context) - I provide the review back to Opus and ask it to verify the review findings - Make Opus plan the fixes then execute them - Ask GPT-5.5 to review the fixes and check if they solve the problems
Cool, but body font size is too small for comfortable reading!
I've used Brave Search and found it better than Google's in some cases
This looks great for quick audio operations without the need to use heavy apps.
One question: I tried the "Fade In" effect; is there a way to control its timing (i.e. the part of the clip where the effect is applied) ?
You can use V4 Pro with Claude Code [1].
I tried it and it's impressive.
[1] https://api-docs.deepseek.com/quick_start/agent_integrations...
My "trick" was to divide things into batches (which can be big with LLMs with larger context sizes) and classify the items in each batch, then take the resulting categories from each batch and feed them into an LLM to group semantically similar categories into groups with a representative category for each group. The representative category can be chosen from the group or created by the LLM. This is an over-simplification of the process but that's the gist of it.
Language support is not mentioned in the repo. But from the paper, it offers extensive multilingual support (nearly 100 languages) which is good, but I need to test it to see how it compares to Gemini and Mistral OCR.
Claude Skills seem to be the option that offers highest flexibility to add more capabilities at most simplicity. Better than MCP in my opinion. Hope it becomes a standard and get adopted by OpenAI and the rest of labs.
Good question! I selected the edition with the smallest Goodreads ID¹ that has the publication date and cover photo available. If all editions don't have publication date nor cover photo, then we get the one with the smallest ID.
And you're right, in a few cases, this resulted in getting less widely read editions for some books.
1: Assuming smaller ID means earlier addition to Goodreads' database.
Hi Jeremy, congratulations for the launch.
How does this compare to Dash?
I've used Dash for many applications, so I'm wondering what are the advantages of FastHTML?
Been looking for such a website to show weather for the whole year like this. Thanks for sharing.
I have Raycast extensions for GPT and Claude models. Whenever I have a question, the most powerful LLMs in the world are two key strokes away.
This way is easier than going to the browser then ChatGPT tab for example then creating a new chat.
I found myself using LLMs more and getting more out of them because of this frictionless interaction. They've become more of actual "helpful assistants."
The article compares GPT-4o to Sonnet from Anthropic. I'm wondering how Opus would perform at this test?
How does it compare to Plotly Dash?
Can you explain more? Like which tool do you use for this wiki page? Or is it an internal tool? And do you use it to write meeting notes and then discuss on the same page?
How is this different or better than Dash or Streamlit?
Like many people here have noticed, it's definitely less quality now than before. It's annoying to be honest to reduce the quality significantly without a notice while we are paying the same amount. I'm willing to pay $40 for the original GPT-4, though.
GPT-4 was remarkable 2 months ago. It could handle complex coding tasks and the reasoning was great. Now, it feels like GPT-3. It's sad. I had many things in mind that could have been done with the original GPT-4.
I hope we see real competitors to GPT-4 in terms of coding abilities and reasoning. Absence of real competitors made it a reasonable option for "Open"AI to lobotomize GPT-4 without a notice.
I couldn't find it on the app store, can you post the link?
Well, we have less than 2 TB of data, and although we are running MySQL on a large instance with ~120 GB of RAM, it's extremely slow when dealing with big tables (like a 25 GB table) and that's why we need "big data" tools like BigQuery.
What is Everyprompt? I searched but found nothing relevant.
Made with Jekyll and hosted on Github + Netlify. All for free although it received tens of thousands of visits (especially on the YouTube analysis post) without a problem.
I've burned out twice
When you say "burn out" can you describe what it feels like? Is it like depression?
Setting up a home gym to lift weights and trying to eat healthier...
Does a treadmill work instead of weight lifting? Curious to know based on your experience.
This is an inspiring story, thanks for sharing.
Exactly! These points drive me crazy!
I think YouTube is filled with interesting content but they always choose to recommend the same videos for weeks!
However, their "New to You" section is much better and I think it should be on the home page.
I've used Miro before for mind maps but after your suggestion, I think it's great as a whiteboard, thanks.
Gather looks fun but I'm afraid it will make you more prone to interruptions and therefore have less focus time. For me, I think it would be stressful to be in a situation like that.
collaboratively manipulating shared representations
Thanks, I would think this is key to enhance meeting quality.
I got your point. Yes, telephone calls would eliminate distractions but it lacks a lot of necessary features. Sharing screen is a must in our company. Also no whiteboard alternative. And we have many non-native English speakers so voice only might not be enough to always understand what they're saying.
I think a telephone call might be good for a few cases but in our case (technology/marketing company,) I don't think it's a good way to communicate remotely.