Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

There's pretty much two ways to deal with this. Either admit this is a low probability failure scenario and it isn't cost effective to have global redundancies. The outage will be resolved as soon as possible. Or, admit you failed to build a georedundant HA infrastructure and apologize with a tentative plan to build out a redundant infrastructure in a different catastrophe zone.

move the servers? On what planet is server infrastructure movable on a whim?



It is movable. It would take a few days. We have another DC. We wouldn't say that if we knew it wasn't doable.

EDIT: Just because it's feasible doesn't mean we will actually do it; just wanted to clarify that we weren't firing from the hip.


Wow you expect fog creek to be down for that long when it does? Why can't you just scp everything on there to your other DC? or at least just move the harddrives? Seems like it would cost less time.


Kiln alone has >4 TB of data; you want to SCP that with 90 minutes heads-up?

Having power and/or rack space is not the same as having servers, switches, etc. anyway.

Hopefully it will not be down that long. We'll let you know more when we know more.


Ha, I wouldn't worry too much about Kiln. I can push when the servers come back. Heck, considering how often some folks I know of push their changesets I wouldn't be surprised if a lot of them never even notice that anything happened.

FogBugz is more of a problem. Some days the "Resolve" button is my only source of job satisfaction.


You could have rsync'd it with a few days heads up, and freshened that in the last 90min.


>4 TB of data? I didn't realize Kiln's gotten that big already. When this whole thing is sorted out, @gecko @kevingessner - would love to see a "State of the Kiln" post and some stats!


Do you guys do off-site backups?


Yes, we have multiple off-site backups (cloud as well as an offsite storage DC) for all customer data. All data is still safe in NYC -- we've just brought down service to prevent problems in case of an abrupt power failure.


Yes we do. Unfortunately, that's all they are -- data backups, sans infrastructure.


Sure, but that invalidates the complaint of it being infeasible to scp out 4 TB of data before your NYC DC runs out of power. Those 4 TB of data are already out of NYC, safely in some other DC where you have hardware. You just need new servers/VMs in/near that DC to restore the backups to.

I'm not trying to 2nd guess your ops team, but the whole point in having off-site backups is to facilitate a your RTO plan in case you lose your primary DC with no warning. I guess I'd be surprised if you don't have a < 24 hour RTO plan in place. With how quickly you can get VMs and even dedicated server provisioned by many hosting providers (minutes to a couple hours), the idea of physically moving servers off-site into a new racks, with new networking, etc... seems kinda nutty...


I can't imagine that relocating an entire environment for a multitude of applications and services is as simple as scping things over. I'm sure there is a process wherein scp could be a step, but I don't see it being any easier/faster.

Moving the hard drives could be an option, and I believe it is sometimes done, but it assumes there are empty boxes on the other end waiting to receive the hard drives in a similar configuration to how the hard drives came. Also there's separate issues depending how many drives they are dealing with, and what redundancy is involved. If it's very few, then you might as well move the server outright. If it's very many, then there's extra human overhead (and room for error) in keeping the drives together.


Why would moving the servers be hard? You can fit 100 terabytes of storage in a shoe-box these days. I'd be extremely surprised if you couldn't run all of FogCreek off of a single 10U blade enclosure. That would be up to 128 CPU cores; I suspect they need only a small fraction of that.

On the other hand, with a 1000/Mbit uplink that they were allowed to saturate, they'd still only be able to copy out 1 terabyte in 3 hours.

Essential quote (literally from Networking 101): "Never underestimate the bandwidth of a station wagon full of tapes hurtling down the highway."


Depends on your mental model of "servers." People have different views of servers from "that nosiy hot thing under my desk" to multiple cages and dozens of racks with 1G, 10G, 40G fiber, rack switches, core switches, edge switches, terminal servers, database servers, application servers, monitoring servers, corporate boxes, .... All with various weights, accessibility, cable routing, and those damn three servers with stripped mounting screws that have been there for six years.

100 TB of high speed RAID-10 would fit in maybe 30 shoe boxes.

Then, after things are moved, you have to deal with drives that have jiggled loose, components that outright fail to work again, or things that get accidentally broken in transit.

Let's just declare an emergency federal holiday until Nov 2 so everybody can recover without dangerous heroic measures.


In this case my mental model of "servers" is "the computers that run the specific small company under discussion, who has already said that they can move them if they want to".

We seem to be arguing separate points -- I'm saying that it's not unreasonable that a small company could be moved fairly easily. Possibly as easily as unplugging a blade enclosure and throwing it in a station wagon. There are loads of small businesses that can run on an amount of hardware that can be easily transported. (When I used to gig on electric bass, my amp and other rack gear was in a portable 8U rack and that was more "portable" than the 100 pound speaker cabinet.)

I can't tell if your point is that it's unreasonable for all companies, which is wrong, or that it's unreasonable for some companies, which is obvious.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: