Monday, June 18, 2007
Detecting Web Page File Size
But the full web page file size includes scripts, styles, images, etc. Ok. So let's parse out all the links and retrieve those with HttpURLConnection. Hmmm, that's a little tricky, because now we have to recursively step through frames and imported stylesheets, but we know HTML so it's possible. Phew, disaster averted.
But uh oh, these pages use a script to preload some images. Now what? Put in Rhino JS interpreter and look for those calls? Yech, now it's just getting nasty.
My solution is in a way nastier and yet more elegant as a whole. Run the page through a renderer, a la WebRenderer, and care less about the technology it uses. Instead, point the renderer at a proxy server on a different thread (finding a simple java proxy server is a whole different post), and count the file size on the proxy. This way, we make sure we trap every little piece of the page for file size counts. Even those little browser icons.
Oh, and just as a note, remember to disable any cache in your renderer and proxy server so it doesn't bypass cached fetches. And obviously this solution would become much more complex in a multi-threaded application. Offhand, if the proxy was lightweight enough, I'd probably just open a different proxy on a different port for every thread. This way you could still track sizes in a consistent way.
Look for a proxy server post coming up where I examine some pros and cons of different free solutions for doing this and certain other tasks I need a proxy for.
Friday, June 15, 2007
Using SVG in XML Thesis
The first challenge that I looked into was graphs and charts; hopefully I found a procedure that I'm ok with. I wanted to use SVG, which Prince supports, as a scalable image source. To generate the SVG, there's really two feasible options in my mind:
- Embed the data directly into the source XML file and preprocess with an XSLT file, a la here. This is a really cool option, but you have to have a smooth XSLT translation and you'd lose fine-grained control over a lot of the elements without specifying a lot of presentation information in your data. However, I will play with that a bit to see if I can get something working.
- Use a SVG graph/chart exporter to create the SVG files. This is actually harder to do then you'd think. No MS product does this, and while OpenOffice can theoretically, its a two program task (Calc->Draw->SVG) and I personally couldn't get the text to output. I finally found GraPL, which in my opinion was the best way to get a graph into SVG from a client-side spreadsheet and looking good. The trial version actually lets you import data and export to SVG, so if you are ok with not saving your graph settings you can do it for free.
Like I said, I'll try the first method first, but I'm honestly expecting it not to live up to my ideals. I'll give updates on how this goes, in addition to my explorations with other neat things, like an XML-driven JSP thesis-o-meter.
Choosing a Thesis Editor
- Simplicity - I can very easily get visual feedback on my document as I write it. While I really do appreciate the theoretical idea of separating presentation and content, its kind of annoying for me sometimes as a visual person.
- Collaboration - Versioning and edit tracking in Word is very good. Similar concepts in LaTeX just don't exist, you have to use custom commands or use very basic text diffing, which is not nearly as convenient.
- Portability - Word is very common and even if it wasn't, it's a simple to understand package. LaTeX is a "system" and a "language" and requires a lot more for a new user to understand.
However, the main disadvantage is that I have to spend several hours minimum switching a paper format when I resubmit it to a different conference. And that's with using all the Word helpers, like fully Styling the paper. With a thesis that is going to be much larger, I could definitely see things getting sort of hairy and presentation becoming inconsistent as I lose track of the whole document.
So I started off this morning by finally considering LaTeX. However, while it is very accepted and seems to do a decent job, it would essentially require me learning a new language. So before I ran off to B&N to sit down and start grinding through another presentation language, I looked into some other options and ended up at Prince XML.
Prince XML is a product that is developed by YesLogic in Australia. It takes XML and CSS, combines it, uses a couple of proprietary CSS tags, and generates a print-ready PDF file. The reason why this appeals to me is several-fold:
- It's a more web-specific technology. My research and personal interest is the web area so why I shouldn't I also use it for my thesis?
- I already know the technology. I've been writing CSS and XML for years, so the fundamental technology I already understand. Furthermore, I'm proficient in analyzing text in XML documents so I can more easily write cool tools like a thesis-o-meter.
- It's popular. XML and CSS are very well known, popular technologies. I can leverage that with a better suite of tools for source editing, display, and version tracking.
My only concern is that I won't have as much control over the output as with LaTeX, so I'll either have to spend time writing hacks or actually changing the content to make it easier for Prince. But at least for now I'm going to attempt to do it this way and see how far I get before I run into problems.
Thursday, June 7, 2007
Review: Reproduced and Emergent Genres of Communication on the World Wide Web
The first paper I've chosen to focus on is a web genre classification paper. Honestly, I've read less of these papers than anything else, mainly because there are so few. If you expand the field to general text genre classification, it gets a bit larger, but the core web genre classification field is limited to less than 25 peer-edited unique papers, and I believe I'm being very liberal with my categorization.
Reproduced and Emergent Genres of Communication on the World Wide Web was one of the first web specific genre papers out there. It was first published at the Hawaii International Conference on System Science in 1997, but a more thorough version of the paper appeared in The Information Society journal in 2000. It was written by Kevin Crowston and Marie Williams of Syracuse University (go SUNY!) who are both information scientist, not computer scientists. This is common to the field currently. It doesn't really become a CS field until it actually gets implemented or the machine learning takes place.
The core idea of the paper is that the genres that exist in traditional print and electronic text were changing as they moved onto the Web. New ones were being created and old ones were being changed. They showed this by sort of random sampling 100 URLs and then classifying each one into a genre and finally examining the genre list. Honestly, the mutation of communication patterns across media is really not my thing (a little too abstract for my CS liking), but I loved the paper as a first step toward establishing genre on the Web. Again, my interest in genre is purely from its applications to assisting the information search process.
Most of the papers they have written since then are somewhat of an extension of this work and some of them have more interesting things to talk about in my opinion, but I chose this piece to highlight because its the first paper that really brings genre classification to the Web specifically. Previous papers may have used the Web as an example or even gotten data from the Web, but usually the genres were text-specific. This was the first place we see Web-specific genres like home pages, hot lists (which are generally known as hubs in the HTML community), and server statistics used as a genre type. Frankly, I think some later papers that use non-Web specific genres should have taken a clue from this direction to do their classification studies.
The original paper was interesting because it actually lists the 100 pages they hand-classified and the genre they chose. Sometimes its refreshing to see actual data coding examples so you can get a more concrete grip on what genre really meant to the paper. However, it was kind of a small sample, they didn't cross-code for consistency, and they didn't do any work with organizing the genres, so the analysis feels a bit immature.
Those things were all improved in the later version of this paper. The sampling method was to actually get random URLs from Altavista, which was an improvement over the earlier manner. They also increased the sample size to 1000. This is a much better number when you figure that at the fine-grained granularity they define their genres to have, they had almost as many genres as documents in the 100 sample. They also did a much more thorough analysis of the genre types and the role they play followed by a very large future works section. In particular, I enjoyed the discussion on their sampling methodology, since I've dealt with the problem of creating good sampling methods using limited web databases (see my paper last year). Also, the point on pages as the unit of representation in genre classification is of particular interest to me in my research. The only problems I had with the later version was that it felt bloated in my mind. Included statistics on basic link and domain distribution had very little to do with the paper, nor did the section on web site design. Go write a web survey paper if that was your intention.
So just to summarize the main contribution: transitioned the concept of genre to Web-specific genres, creating the first Web-specific genre palette.
Tuesday, June 5, 2007
On Thumbnail Preview
More worried because (at the risk of sounding like a promoter) the feature is incredibly useful. By mousing over the links I can very quickly weed out spam-like sites that I'm not really interested in and get to the more appropriate ones that I need. Search UI assistance is particularly useful in dealing with broad queries, people who don't know how to form queries, and one word queries. In all of these cases, the results will be quite varied and manual filtering is required. By putting a thumbnail preview on the page, I take out a step for all of those query types. I don't have to page back and forth through results, looking for the one nugget of goodness.
However, I'm reassured because people are lazy and the process isn't perfect. First of all, sometimes the information is not there or is old or couldn't be rendered correctly for the preview. However, any classification system would suffer from similar issues, so I'll not fault them there. The fundamental problem is that the searcher still has to do the work. The positive aspect of this is that in general humans are smarter than machines. On the otherhand, they still have to go through the same process, it's just faster. In essence, the tool isn't really helping anyone find information, it's just making the task quicker.
Other aspects of Ask.com and Google do make my life easier. When they detect that I'm searching for a city and give me weather and maps and all sorts of neat things, that's taking a good guess at what I want to know. I think that genre classification falls more into this type of category. And that is what strengthens my resolve as I progress my research in that area.
Welcome
As way of introduction, I'm a PhD student in Information Retrieval at Binghamton University. I'm currently doing a dissertation on applying higher order HTML document features to several different fields of HTML analysis, currently focusing on genre classification. Genre classification is looking at consistent themes in web page design and information in order to help guide a user to the information they are looking for.
The main point of this blog is to serve as a brainstorming pad for my ideas on using visual features. A large area of research in academia is devoted to analyzing the text of an HTML web page, yet VERY little is devoted to analyzing the layout of the page. This is primarily because text documents were all the rage 50 years ago when most of the research was being done. Furthermore, in several tasks like content classification, its not really as important. However, in other fields like the up and coming genre classification, it can be much more useful.
While I believe starting a blog about academic research is in some ways dangerous (there's this pervasive academic atmosphere of quasi hiding what you're doing until you publish it or at least have the resources to do it better than anyone else), I think it should spur me to be quicker in my implementation of ideas.