Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Good stuff.

I do a good bit of scraping, and made RubyRetriever[1] to make my life easier but it seems like I'm getting roadblocked on occasion, probably due to some of the things you mention in your article.

Is there any way for a site to verify that only their JS and CSS files are linked? Like preventing injection?

[1]: https://github.com/joenorton/rubyretriever



You could inspect the src attributes of script tags, and the href attributes of link tags with rel="stylesheet", for acceptable domains. I doubt it would cover all cases, but it might be a start.


I got the 100th star on the repo! What do you mean by the verifying part?




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: