HN user

gpapilion

1,086 karma
Posts37
Comments369
View on HN
blog.hypergeometric.com 13y ago

Fork Less in Bash and See Performance Wins

gpapilion
1pts0
blog.hypergeometric.com 13y ago

Getting Unique Counts From a Log File

gpapilion
9pts8
blog.hypergeometric.com 13y ago

Just Enough Ops of Devs

gpapilion
2pts0
blog.hypergeometric.com 13y ago

Thanks Mr. Jobs, But it Seems I Can Use a Linux Laptop Now

gpapilion
170pts165
blog.hypergeometric.com 13y ago

The Pitfalls of Web Caches

gpapilion
1pts0
blog.hypergeometric.com 13y ago

Infrastructure – The Challenge of Small Ops – Part 3

gpapilion
1pts0
blog.hypergeometric.com 13y ago

Solr Upgrade Surprise and Using Kill To Debug It

gpapilion
1pts0
blog.hypergeometric.com 14y ago

What I Wish Some Had Told Me About Writing Cron Jobs

gpapilion
3pts3
blog.hypergeometric.com 14y ago

6 Phone Screen Questions for an Ops Candidate

gpapilion
1pts0
blog.hypergeometric.com 14y ago

Rally Cars and Redunancy: Understand Your Failure Boundaries

gpapilion
1pts0
blog.hypergeometric.com 14y ago

Two Helpful Data Concepts

gpapilion
1pts0
blog.hypergeometric.com 14y ago

Keep it Simple Sysadmin

gpapilion
2pts0
blog.hypergeometric.com 14y ago

Monitoring – The Challenge of Small Ops – Part 2

gpapilion
1pts0
blog.hypergeometric.com 14y ago

Three First Pass Security Steps

gpapilion
2pts0
blog.hypergeometric.com 14y ago

Monitoring Your Customers with Selenium and Nagios

gpapilion
1pts0
blog.hypergeometric.com 14y ago

The Challenge of Small Ops (Part 1)

gpapilion
1pts0
blog.hypergeometric.com 14y ago

Good Nagios Parenting, Avoids a Noisey Pager

gpapilion
1pts0
blog.hypergeometric.com 14y ago

Redunancy Planning, more work than adding one of everything

gpapilion
1pts0
blog.hypergeometric.com 14y ago

Three Monitoring Tenants

gpapilion
1pts0
blog.hypergeometric.com 14y ago

Groovy, A Reasonable JVM Language for DevOps

gpapilion
1pts0
venturebeat.com 14y ago

You don’t need a Mayan calendar to predict this potential disaster

gpapilion
3pts0
blog.hypergeometric.com 14y ago

A few things you should know about EC2

gpapilion
1pts0
blog.hypergeometric.com 14y ago

Configuration Management Tools Still Fall Short

gpapilion
2pts0
blog.hypergeometric.com 14y ago

SSH Do’s and Don’ts

gpapilion
107pts61
blog.hypergeometric.com 14y ago

Techincal Debt Better Than Not Doing It

gpapilion
3pts0
blog.hypergeometric.com 14y ago

User Acceptance Testing for Successful Failovers

gpapilion
1pts0
blog.hypergeometric.com 14y ago

Solr Query Change Beats JVM Tuning

gpapilion
2pts0
blog.hypergeometric.com 14y ago

Language Importance for DevOps Engineers

gpapilion
2pts7
blog.hypergeometric.com 14y ago

Stupid Bash Expansion Trick

gpapilion
1pts0
blog.hypergeometric.com 14y ago

Dealing with Outages

gpapilion
1pts0
SpaceX S-1 2 months ago

Yes and no. They paid what everyone else pays for those gpus. NVIDIA make the profit and leaves crumbs for the rest. For other components they paid less, but since the gpus are the majority of the cost…

So recently I moved from a Anthropic model to a qwen 3.5 model running on my Mac to summarize ticket activity over 7 days. I used to do this manually with a colleague and it would take us a couple hours to go through. Opus took 58 seconds, and Qwen took 2.5 minutes. The quality of the qwen output was comparable, but the there was a 2.5x difference in time.

All that said I actually don’t think that matters much. I think we are dragging attention economy concepts in to ai responses, and it doesn’t matter. Both options saved me hours per week, and the difference between 3 and 1 minute may not be worth the additional cost.

Also there are times when the model output is much better with anthropic, but it’s not all the time. I think it becomes a question should we be using the best model for all questions?

Cerebras S-1 3 months ago

The initial cost of serving is very high, and while super performant not great for scaling up.

In practice they are also not very flexible when compared to gpus.

It’s a very different company post the PwC purchase. They have around 1/3 of the revenue from consulting which tends to push the valuation down due to its relative low margin when compared to software. This also inflates the number of employees.

Just to level set here. I think its important to realize this is really focused on allowing things like search to operate on encrypted data. This technique allows you to perform an operation on the data without decrypting it. Think a row in a database with email, first, last, and mailing address. You want to search by email to retrieve the other data, but don't want that data unencrypted since it is PII.

In general, this solution would be expensive and targeted at data lakes, or areas where you want to run computation but not necessarily expose the data.

With regard to DRM, one key thing to remember is that it has to be cheap, and widely deployable. Part of the reason dvds were easily broken is that the algorithm chosen was inexpensive both computationally, so you can install it on as many clients as possible.

The large api/token providers, and large consumers are all investing in their own hardware. So, they are in an interesting position where the market is growing, and NVIDIA is taking the lion's share of enterprise, but is shrinking at the hyperscaler side (google is a good example as they shift more and more compute to TPU). So, they have a shrinking market share, but its not super visible.

Realistically groq is a great solution but has near impossible requirements for deployment. Just look at how many adapters you need to meet the memory requirements of a small llm. SRAM is fast but small.

I would guess their interconnect technology is what NVIDIA wants. You need something like 75 adapters for an 8b parameter model they had some really interesting tech to make the accelerator to accelerator communication work and scale. They were able to do that well before nvl 72 and they scale to hundreds of adapters since large models require more adapters still.

We will know in a few months.

I would think this is for rental fleets or bike share. The weight and design would seem to make sense for that. Though the single speed seems like and odd choice for that.

I think that the private carriers are more likely to be helped by this, since they will manage the paperwork.

It’s more likely a set of products that were shipping directly from factories disappears from the market. For example, the direct from factory Halloween costume.

It could end up being a step backwards in living standards and access to daily luxuries.

Gradual damage is consistent with over heating. I've seen racks of servers do the same thing.

Overall, there is a continued challenge with CPU temperatures that requires much tighter tolerances both in the thermal solution. The torque specs need to be followed and verified that they were met correctly in manufacturing.

More generally beats better. That’s the continual lesson from data intensive workloads. More compute, more data, more bandwidth.

The part that I’ve been scratching my head at is whether we see a retreat from aspects of this due to the high costs associated with it. For cpu based workloads this was a workable solution, since the price has been reducing. gpus have generally scaled pricing as a constant of available flops, and the current hardware approach equates to pouring in power to achieve better results.

Scope, it’s all about scope of your team. Em to director requires opportunity as well as performance.

For you that means focusing on a growing area of the company, and finding new areas to grow your team in. You also need to have a team of managers, who are growing their scope as well.

Apple M3 Ultra 1 year ago

I think this will eventually morph into apples server fleet. This in conjunction with the ai server factory they are opening makes a lot of sense.

I don’t know this is significantly different than modern engines. They require special tools and software too.

The bigger issue I think is most of the cars are teslas, which didn’t behave like a normal automaker for better or worse. For example the work done during the pandemic to avoid supply chain crunches may result in a maintenance headache a few years from now.

The headline discussion on the podcast covers whether chatgpt is actually successful. They point to relatively few use cases emerging, and the continual or press around agi. They cover how there is now pressure to build an ads into the platform to build revenue.

They weren’t interested in creating an open solution. Both intel and AMD have been somewhat short sighted and looked to recreate their own cuda, and the mistrust of each other has prevented them from a solution for both of them.

Consumers want faster processing the instructions are just the method to get there. And they aren’t the best since the area dedicated to the instruction could be used for something else.

It is insane especially if you think emulation is performant enough to allow for a switch.

Mishandling aside, the issue I've seen is there really isn't consumer demand for this. Prior to AMD having AVX512, most of the comments were around wasting the silicon on SIMD, rather than improving other aspects of the CPU. I'm pretty sure there was good reason to think it was largely a dark area of the chip.

From what I've seen, but haven't heard discussed much, the naive implementation vs AVX512 is a huge gain, but AVX2 vs AVX512 was not very impressive for the application I was looking at. The complexity this code added, and the cases where we needed it to run on AMD (for other reasons), basically made taking advantage of the feature undesirable for a single digit gain.

Things like VNNI or AMX are better wins, but they are only needed in very specific cases. VNNI in particular looked to be a 30% improvement in a BERT workload.