Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

If you're on a shoestring budget, can't you just afford the day of downtime once every few years? Yeah, it's annoying to be down, but each 9 you add past 99% costs more than the last.


For reference, here is how many days/hours the system is unavailable with different percentages:

  90%      36.5 days ("one nine")
  95%      18.25 days
  98%      7.30 days
  99%      3.65 days ("two nines")
  99.5%	   1.83 days
  99.8%	   17.52 hours
  99.9%    8.76 hours ("three nines")
  99.95%   4.38 hours
  99.99%   52.56 minutes ("four nines")
  99.999%  5.26 minutes ("five nines")
  99.9999% 31.5 seconds ("six nines")
http://en.wikipedia.org/wiki/High_availability#Percentage_ca...


It's even easier if you measure per week instead of per year since a week has ~10,000 minutes.

Then the rule of thumb is 10 minutes is 0.1%, 1 minute is 0.01%, and 6 seconds is 0.001%; or 10 mins = 99.9%, 1 min = 99.99%, and 6 seconds = 99.999% respectively.

(Note that 6 seconds a week * 52 weeks = 5.2 minutes, same as the reference table above.)

Programmers don't think in years, but they can think in "any given week". This rule of thumb puts things in an easy-to-remember perspective.


I spent a year as a High Performance Computing Sys Admin. After careful consideration, we just decided reliability wasn't worth the cost. We bought more hardware instead of a UPS system for our data center. Once or twice a year, we experience a power blip and lose all our compute nodes (infrastructure is on UPS). We would send out an apology email to our users and tell them to resubmit their jobs. Worst case scenario was someone had to restart a 14 day job. Had we bought the UPS system, that 14 day job would take almost twice as long to complete, spending a week or so longer in the queue. We started embracing the idea of acceptable failures and saw it had a great impact on our service.


Do your users run 14day jobs with no checkpointing whatsoever? I'd be afraid of a bug in my code crashing the computation 90% of the way through. The MapReduce setup seemed much more resilient to things like this for example.


each 9 you add past 99% costs more than the last.

You're right about that. In fact, each 9 past 99% costs twice as much as the last. But on the other hand when you have paying customers, one day of downtime they might forgive you. Two days and they'll be annoyed. But on the third day they'll come for you with pitchforks in hand. If you can avoid pitchforks with clever architecture and slightly higher costs, I think it's worth it.


Is it only twice? In my experience, 99.9999% uptime is significantly more than 16x more expensive than 99% uptime.

For comparison, 99.9999% uptime means about 30 seconds of downtime in a year, while 99% uptime is about 3 days of downtime. You can get 99% uptime with a singly-homed, not terribly reliable commodity system. For 99.9999%, you need multiple redundancies in every architectural component with automatic error detection and failover, and have to watch every change to make sure it doesn't introduce the possibility of system instability or cascading failures. Those are qualitatively different approaches to software engineering.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: