Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I also work in this field and what we typically do is a lot of nested cross-validations to get some bounds on a model building process and some idea of how it would perform on repeated unseen data. Data leakage is always on our mind and we do our best at all stages to avoid that. We also train on data from many sites. It can be done and it can be done properly. As you say, it is always best to collect some completely naive test set to back up the model-building process. If you design your pipeline properly, the test set should fall within the bounds you got during cross-validation. It all depends on how much data you have and I think as long as you design your pipelines with that in mind and acknowledge limitations with smaller datasets, then the research is valid and useful.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: