Detecting a full web page file size from within Java is not easy. On the surface, it sounds pretty easy - in fact, in a couple lines you could open an HttpURLConnection, grab the stream, count it and bang, accurate size count.
But the full web page file size includes scripts, styles, images, etc. Ok. So let's parse out all the links and retrieve those with HttpURLConnection. Hmmm, that's a little tricky, because now we have to recursively step through frames and imported stylesheets, but we know HTML so it's possible. Phew, disaster averted.
But uh oh, these pages use a script to preload some images. Now what? Put in Rhino JS interpreter and look for those calls? Yech, now it's just getting nasty.
My solution is in a way nastier and yet more elegant as a whole. Run the page through a renderer, a la WebRenderer, and care less about the technology it uses. Instead, point the renderer at a proxy server on a different thread (finding a simple java proxy server is a whole different post), and count the file size on the proxy. This way, we make sure we trap every little piece of the page for file size counts. Even those little browser icons.
Oh, and just as a note, remember to disable any cache in your renderer and proxy server so it doesn't bypass cached fetches. And obviously this solution would become much more complex in a multi-threaded application. Offhand, if the proxy was lightweight enough, I'd probably just open a different proxy on a different port for every thread. This way you could still track sizes in a consistent way.
Look for a proxy server post coming up where I examine some pros and cons of different free solutions for doing this and certain other tasks I need a proxy for.
Subscribe to:
Post Comments (Atom)
No comments:
Post a Comment