Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I use a telegraf, influxdb, grafana stack. I find the cost in maintainance and initial setup quite low. Telegraf is super easy, just uncomment the things you want it to collect from the config file. Influxdb is just adding the correct users and a database, never touched it since. Grafana can be a time sink if you want to bikeshed your dashboards, but there are a lot of pre-made ones that handle common usecases.

It's really nice to have some metrics when for instance a service goes down. It's super easy to spot a OOM situation or other vertical scaling issues.



Thanks. I completely agree with the bikeshedding. We're playing with Prometheus/Grafana + ELK, and being able to visualise all this data, its hard to work out what is useful and what is just fun to play with.

I'm now wondering where the line is between necessary monitoring and bikeshedding? I could look at introducing distributed tracing, but will it actually add any meaningful value?




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: