HN user

mikiem

742 karma

Founder and CEO of M5Hosting.com mike at m5hosting dot com

Posts2
Comments99
View on HN

Initially, it seemed like a DoS to us too, but it was not. This was confirmed by upstream provider metrics. No major traffic spikes. It was a combination of non-malicious things. More info later, some of us need sleep.

I'm a little surprised at the response here.

I feel like there is an element of "Body Doubling" here... a strategy used by those with ADD/ADHD. I recently looked in to this when curious about my own observation that I work longer and with better focus when working in close proximity of someone else.

HN is up again 4 years ago

This morning, I googled for issues with the firmware and the model of SSD, I got nothing. But now I am searching for "40000 hours SSD" and a million relevant results. Of course, why would I search for 40000 hours.

This thread is making me feel a lot less crazy.

HN is up again 4 years ago

They were in two mirrors, each mirror in a different server. Each server in different racks in the same row. The servers were on different power circuits from different panels.

HN is up again 4 years ago

These were made by SanDisk (SanDisk Optimus Lightning II) and the number of hours is between 39,984 and 40,032... I can't be precise because they are dead and I am going off of when the hardware configurations were entered in to our database (could have been before they were powered on) or when we handed them over to HN, and when the disks failed.

Unbelievable. Thank you for sharing your experience!

HN is up again 4 years ago

You are never going to guess how long the HN SSDs were in the servers... never ever... OK... I'll tell you: 4.5years. I am not even kidding.

It was part of a mirror of identical SSDs on an LSI MegaRAID RAID card. We see occasional "spectacular" drive failures that take the machine down with a single disk failure. Usually it's just a reboot to come back up, and a disk replacement, then some hours of time to rebuild the array and get back to situation nominal.

HN was down 5 years ago

Thank you for sharing your positive experience! We can power cycle power outlets remotely and can connect a console (ip kvm)... and we are staffed 24x7.... in case you need another server. Thanks again!

HN was down 5 years ago

Oh hi! Thank you for the kind words. I cant tell who you are by your name here, but if you've been with us since 2011, we have certainly spoken. Are you using our second San Diego data center for your failover location? If you and I aren't already talking directly, ask to speak with Mike in your ticket.

HN was down 5 years ago

Unrelated issues, but I did hear from our other clients that O365 was having issues at the same time as our network outage affected HN and many others.

HN was down 5 years ago

Founder and CEO of M5 Hosting here. We did have a network outage today that affected Hacker News. As with any outage, we will do an RCA and we will learn and improve as a result.

I'm a big fan of HN and YC in general, we host of other YC alum, and I have taken a few things through YC Startup School. During this incident, I spoke to YC personally when they called this morning.

Cogent and Cox are also having problems, but we are seeing a lot more successful traffic on Cogent than CenturyLink. It appears that CL is also not withdrawing stale routes. It seems CLs issues are causing issues on/with everything connected to it.

M5 Hosting here, where this site is hosted. We just shut down 2 sessions with Level3/CenturyLink because the sessions were flapping and we were not getting complete full route table from either session. There are definitely other issues going on on the Internet right now.

Actually, we just bought bandwidth for a roll out at Equinix in Munich. $0.50 for Cogent (when added to several other 1G commit on 10G ports in our account. A single 1G commit on a single 10G port would cost more) and we were quoted $1.43 for Level 3 after rejecting a $1.70 quote. Both Cogent and Level 3 were 2yr terms, and in an "on net" location. We are going with another provider besides Level 3 there, but I used these as examples in the parent. You thought it was not "reality", and I refute that.

As a provider of IaaS Cloud and of dedicated servers and colo, I hear this argument all the time. No one ever seems to include the Network Engineers, monitoring systems, the routers (better have more than 1!), the switches (distribution and access layers), the maintenance, software licenses (where applicable), customer support, cost of IP addresses, Account Payable, ARIN membership, RADB membership, cross-connects, optics, spares and/or support contracts, etc... and finally, you do not use a 1Mbps at 100% for 24hrs per day, so while 1Mbps for a month is ~320GB, in reality, the way most people transfer data, 320GB would look more like 3Mbps at 95th percentile (the way burstable bandwidth is billed)

A basic 1Gbps commit on a 10Gbps port in a data center might cost you from $0.50/Mbps (something like Cogent) to maybe $1.50/Mbps (let's say Level 3), other providers could be $4+/Mbps. By the time you factor in all of the above overhead costs, the true cost of the bandwidth is much much higher on a per Mbps basis.

Don't forget to significantly over-build your stuff, or you might get knocked off-line for anomalies or DoS attacks.

Admittedly, the scale of Google, AWS, Azure makes the cost per Mbps much much lower, but when as others have pointed out, AWS, Google, Azure don't need to charge less than they do.

I am seeing a lot of comments about the need for stable power. It seems logical, but it’s not an absolute need. I managed a data center win San Diego at the time of the California electrical power crisis [1] near the end of the Dot Com days. I have operated more than one at a time since. During those times, which produced rolling blackouts, we just ended up running our generators more often. Power outages in Data Center happen during maintenance and upgrades, not so much during power grid outages. I have still never never had a data center go down during a power grid outage… but have seen many during upgrades, and maintenance.

In the years since the rolling blackouts, I have been colocated in data centers that have made the decision to use special pricing available from the power utility, available when you can have a preemptable load… during peak hours of peak season the power company can tell you to get off the grid with short notice. The data center operators did OK when they did this for most years. They got a lower price on power all year, in exchange for running their generators more. I’d say that the typical year, they ran the generators for several hours per day for 7 to 10 days in the Summer months. The last year they dd this, they ran a much higher number of days than expected… maybe 15 - 18 days from noon to 7pm. I understand they had to stop doing this because of the pollution of the generators or the permits required for such heavy usage... but that info is not first hand. It could have been many other reasons such as neighbors complaining about the noise in the business parks (2MW diesel generators are deafeningly loud), costs related to running the generators for so many hours, etc.

To operate like this, you have to be on your game for maintenance of the generators and checklists and training. You also have to have multiple contracts with fuel delivery trucks, just incase your outage lasts for a while. These data centers were all under 5MW in size each. We never lost critical load during a power supplier outage. I hope I have illustrated that a reliable power supply is not strictly required to run a reliable data center or service that is dependent upon a reliable data center.

[1] https://en.wikipedia.org/wiki/California_electricity_crisis

You can Google the phone number in the letter and also look up the number for the field office they are from, call the office and ask for the agent.

Yes, it may take time. But you don't get many NSLs, so you do it. You do it to protect yourself (liability of disclosing info without a legal order) and your customer/user who is the subject of the letter. Every time.

Interesting. As a service provider (hosting) we have received many "court orders" that are very similar to these NSLs... but they were not NSLs. Now that I see these NSLs, I am not that freaked out by them. I'm not sure of all the hub bub, at least for these particular NSLs. The scope of these is basically limited to identifying the user. These specifically say to not provide content of the account to the FBI. The not-NSL court orders we have received have included verbage to not disclose the request to the subject of the request.

I thought NSLs were supposedly non-contestible, broad and were for communication detail. These don't seem to be any if that.

The requests we have received have been from a variety of organizations (but signed by a magistrate) ranging from local law enforcement to three letter acronyms and one entity that is neither. While the requests don't say why the order is being issued, we usually receive a call from the agent/detective beforehand and dialog ensues in which they explain what's going on.

While many companies will just give the info, we scrutinize the request and ask the agent/detective politely and apologetically that we can help, but only if they acquire a court order. We have caught not-legitimate requests before, so we verify the request is legit before responding. We have never been asked for content of communications. If Google is not doing the same thing... oof. Just as a matter of process I assume they do. I recall in the past some networks having right in their WHOIS info, how/where Law Enforcement can send FAX requests.

It's a common and well known business model. We have Web hosting companies, game servers, mobile apps, cloud-managed gizmos, SaaS platforms, etc... all are common well known business models. Do you know any more about them because of this post or the information I have talked about? Maybe you didn't know they business model exists.

I should add that some sites (especially ones counting on ad revenue, or that don't want "scalpers") may consider bots hitting them as abusive and will either contact us or just block specific IPs. At that point we know what is going on. Our customer may call us because "it stopped working" when they have been blocked or rate limited. I will point out that many sites that bots hit do not care at all and know it is happening.