I do a good bit of scraping, and made RubyRetriever[1] to make my life easier but it seems like I'm getting roadblocked on occasion, probably due to some of the things you mention in your article.
Is there any way for a site to verify that only their JS and CSS files are linked? Like preventing injection?
You could inspect the src attributes of script tags, and the href attributes of link tags with rel="stylesheet", for acceptable domains. I doubt it would cover all cases, but it might be a start.
I do a good bit of scraping, and made RubyRetriever[1] to make my life easier but it seems like I'm getting roadblocked on occasion, probably due to some of the things you mention in your article.
Is there any way for a site to verify that only their JS and CSS files are linked? Like preventing injection?
[1]: https://github.com/joenorton/rubyretriever