Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

It's the best definition. It makes "big data" the name of the problem you have when your data can not be worked in a coherent way, what must be solved by distributed tools.

If you can buy a bigger machine, you can make "big data" bigger, and maybe evade this problem; if you must access it a lot of times, fitting on disk is useless and "big data" just got smaller; etc.



How then does "big data" differ from traditional HPC and mainframe processing? Those fields have been dealing with distributed processing and data storage measured in racks for decades.


I think the simplest answer is that it's often essentially the same thing but approached from a different direction by different people with different marketing terms.

One area which might be a more interesting difference to talk about might be flexibility/stability. A lot of the classic big iron work involved doing the same thing on a large scale for long periods of time whereas it seems like the modern big data crowd might be doing more ad hoc analysis, but I'm not sure that's really enough different to warrant a new term.


I dunno. Is it useful to separate them?




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: