Monday, December 17, 2007
Eclipse SWT-XPCOM Renderer
At first I was turned off because it appeared fairly heavyweight in dependencies and confusion. However, the swt library is one jar, the XPCOM libraries are one jar, and then you need a GRE installed (through XULRunner). Then use both the snippets from the SWT site and the XPCOM interface APIs and you can do almost everything you can from within native Mozilla code through Java fairly easily.
This is sort of bad timing for me since I'm currently in the middle of lockdown mode where I'm trying not to go off on coding binges and spend more time writing my dissertation. However, I may have to take a small detour, just for the sake of the library.
Just to go over several advantages of SWT-XPCOM over WebRenderer (sorry guys):
1) It gives me access to file size information through the http cache Mozilla has
2) It shuts down, which WebRenderer doesn't do...they hold open a native library listener until the system exits
3) It's easier to upgrade the mozilla version, you just get a new GRE
4) It doesn't require me to pass interesting page calls through serialized Javascript, I can directly access the native objects through JNI interfaces
5) Their JNI interfaces appear to be better written, some calls in WebRenderer take forever
Disadvantages:
1) Requires a GRE installed
2) Requires some understanding of how Mozilla code works (runtime interface casting)
Sunday, December 2, 2007
Plugin Frameworks
I quickly found JPF, which is beautiful in the way they take the plugin architecture to the extreme. However, in my opinion, architecturally they took it too far when they allowed plugin initialization methods. This enables them to start applications without having a main class really. The entire application is just plugins and then you point a configuration file to the plugin that is the "main" plugin. Again, I appreciate the elegance of the solution, but realistically I think people like seeing the main application code. The one thing that JPF does that Nutch doesn't do is create their own classloaders, so the code spaces for plugins is separate. This is a really neat feature that I'd like to add to WebSeer eventually, but....
I ended right back at the Nutch plugin framework. It's not as powerful but I think it's easier to understand. This application has to be understandable by IR students and having core functionality wrapped in plugins would be hard for the ridiculously poor students I normally deal with to understand. Core interfaces and the code that brings those interfaces together to run something will be in the main library and implementations of those interfaces will be the plugins.
On a side note, anyone know the history of these plugin frameworks? JPF and the Nutch frameworks are suspiciously similar, Nutch is just a dumbed down version. I didn't look into the histories of these projects to see how they overlapped by one was definitely inspired by the other.
Sunday, October 7, 2007
Webseer Architecture
FeatureExtractor - a visitor that collects features on a model
DataSourceIO - turns bytes into models
Transformer - transforms models into other models
I haven't figured out the classification side yet, but I need to abstract the common classification tasks into an interface. Something with more a little more power than just being able to plug in weka classes.
Thursday, September 20, 2007
Nutch Rearchitecture
Sunday, September 2, 2007
Mozilla Flaw
Monday, August 6, 2007
WebRenderer
However, up until this year, I wasn't interested. Why? To be honest, their API and event model was fairly annoying and hard to use. They don't give you very sophisticated DOM element methods, so you're forced to reconstruct JS calls on the document in order to get things like background color on a table cell. Their event model is shaky, though to be fair, it's a hard problem. Determining when a complex document is loaded is not easy I guess. Often I won't get document complete events and sometimes I'll get multiple if there are behind the scenes redirects. I understand that it's hard to avoid this with what HTML can do these days, but I want some of this encapsulated better.
The thing that made me reexamine WebRenderer is that they finally made it lightweight. The advantage of the "pure Java" implementations out there (IceSoft Browser for instance) is that they are lightweight so can be drawn on top of and tossed around without visualizing. (The main disadvantage is that the rendering is often pretty bad.) The main version of WebRenderer doesn't allow this type of rendering. It essentially dedicates a particular part of the screen to allowing Mozilla code to draw the HTML document. You can't touch it after it hands over control really, besides some integrated event handling.
It's nice to be able to highlight portions of the screen and interact with the actual document and the tool that I'm building within WebSeer to do labeling needs that ability. So I'm a strong proponent of WebRenderer now for this reasion. I am concerned that it still has some major flaws from a document analysis perspective. However, if you're just looking for integrated lightweight Java rendering, I'd highly recommend it.
Tuesday, July 31, 2007
Entropy Segmentation
Elements are laid out horizontally on a page in this pattern: XXAXAXAXAXAX. For simplicity, A elements are all the same type of element and X elements have nothing in common at all. Therefore, only A elements are in the same class.
Ideally, we want the algorithm to split on the left side of the first A. This will maximize the order on the right side of the page, or otherwise trap all the As. Doing an entropy measure on the right side of that split (with the sign switched for ease), we get 5/10 lg 5/10 + 5 (1/10 lg 1/10) = 2.66. Now let's compare that with the split at the second A: 4/8 lg 4/8 + 4(1/8 lg 1/8) = 2.5. Apparently here, it's better to split at the second one because we're reducing the number of noise elements in the system. Also, there is no advantage to having larger classes, just less proportional ones.
So I've been playing with different ways of modifying the algorithm. Essentially, you want to make the noise elements worth less in determining entropy.
Friday, July 13, 2007
Reflection-based Visitor
My existing problem is that I have several different types of visitors that are used for different purposes. So let's create an example - say I have a graph based model of something, with types Graph, Node, and Edge. Then I have a GraphVisitor that has three methods: visit(Graph), visit(Node), and visit(Edge). The accept methods in Graph, Node, and Edge take care of the default traversal. Now conveniently, I can print out all the edges in the graph or something else without having to care about the actual graph structure in my visitor code. Great...simple visitor pattern.
Now in my library, I have several different types of models, lets say I have another one that's Tree. Both Tree and Graph inherit the Model interface. Now I want to create a visitor that visits a list of Models without caring what type they are. We sort of have a problem. I can create a ModelVisitor, add an accept(ModelVisitor) method in the Model interface, and then check the type of the ModelVisitor and if it's also a GraphVisitor for instance, then do the component traversal. But that's messy. For the last long while, that's what I've been doing. However, in more interesting cases, it gets silly. You have to write accept methods for every type of visitor that could visit the node. Being as this is a visitor-centric architecture, this gets ridiculous even with inheritance. For instance, my TextNodeImpl class recently looked like:
public class TextNodeImpl implements TextNode {
...
public void accept(TextVisitor visitor) {
visitor.visit(this);
}
public void accept(DataSourceVisitor visitor) {
visitor.visit(this);
}
public void accept(TreeVisitor visitor) {
visitor.visit(this);
}
public void accept(MutableTextVisitor visitor) {
visitor.visit(this);
}
}
Kinda crazy, right?
A while back I played around with reflection-based visitors, but I think I took it way too far (I also made the accepting objects reflection-based) and it got hard to understand because there was very little type-checking. However, I revisited it last night and I think I found a good balance.
All the visitors implement the SuperVisitor interface. You can either use a more type-safe generics version:
public interface SuperVisitor {
public boolean canVisit(Class<?> dataClass);
public <T> DataSourceVisitor<T>
getVisitor(Class<T> dataClass);
}
where DataSourceVisitor just has one "visit(T object) " method.
or the more handy and yet not as clean way:
public interface SuperVisitor {
public void visit(Object object);
}
where the visitor will attempt to visit any object and just fail quietly if it can't.
Then in AbstractSuperVisitor which takes care of the reflection, you initialize a Map with a mapping of requested Class to DataSourceVisitor. It initially puts in all the visit methods that are found in the runtime class. Then when a visitor for a particular runtime class is requested, you walk up the class hierarchy tree to find the appropriate visit method(s) that has already been initialized. Then you insert that new mapping into the tree. This caching is necessary to cut down lookup time, since reflection is pretty slow. And while this does create an extra object for each visit method, plus a Map for every Visitor pattern, you rarely have so many visitors that this would matter too much. So the primary cost is the cost of reflection based invocation. If this cost is important, I advise you look into articles that discuss how to replace reflection with byte code generation, which would work perfectly here.
So what this allows is that I can now create a Visitor that can visit any type of object on the fly. Let's say I want to print out everything this visitor comes across. Just put in a visit(Object o) method with a print line in the method and all the visit calls will be resolved to this method.
In addition, I now only need a single accept method in all my model classes, which calls the generic visit on itself and then passes the visitor to its components.
public class Graph {
...
public void accept(SuperVisitor visitor) {
visitor.visit(this);
for (Edge edge : edges) {
edge.accept(visitor);
}
}
}
The design loss is in the loss of cohesion between the visitor and the things it's visiting. However, this can be mitigated by still using particular visitor interfaces that are sort of fake visitor guidelines for methods it should implement. So I can still have a GraphVisitor which enforces some methods on the visitor implementation, however, the Graph object will only see it as a SuperVisitor.
Overall, I think this is a neat solution to a model-based architecture where you have a lot of visitors interesting in different aspects of complex models. I also added a preVisit and a postVisit which gives me more control over post-children processing for instance. Hopefully this will be a good enough model for WebSeer.
Monday, June 18, 2007
Detecting Web Page File Size
But the full web page file size includes scripts, styles, images, etc. Ok. So let's parse out all the links and retrieve those with HttpURLConnection. Hmmm, that's a little tricky, because now we have to recursively step through frames and imported stylesheets, but we know HTML so it's possible. Phew, disaster averted.
But uh oh, these pages use a script to preload some images. Now what? Put in Rhino JS interpreter and look for those calls? Yech, now it's just getting nasty.
My solution is in a way nastier and yet more elegant as a whole. Run the page through a renderer, a la WebRenderer, and care less about the technology it uses. Instead, point the renderer at a proxy server on a different thread (finding a simple java proxy server is a whole different post), and count the file size on the proxy. This way, we make sure we trap every little piece of the page for file size counts. Even those little browser icons.
Oh, and just as a note, remember to disable any cache in your renderer and proxy server so it doesn't bypass cached fetches. And obviously this solution would become much more complex in a multi-threaded application. Offhand, if the proxy was lightweight enough, I'd probably just open a different proxy on a different port for every thread. This way you could still track sizes in a consistent way.
Look for a proxy server post coming up where I examine some pros and cons of different free solutions for doing this and certain other tasks I need a proxy for.
Friday, June 15, 2007
Using SVG in XML Thesis
The first challenge that I looked into was graphs and charts; hopefully I found a procedure that I'm ok with. I wanted to use SVG, which Prince supports, as a scalable image source. To generate the SVG, there's really two feasible options in my mind:
- Embed the data directly into the source XML file and preprocess with an XSLT file, a la here. This is a really cool option, but you have to have a smooth XSLT translation and you'd lose fine-grained control over a lot of the elements without specifying a lot of presentation information in your data. However, I will play with that a bit to see if I can get something working.
- Use a SVG graph/chart exporter to create the SVG files. This is actually harder to do then you'd think. No MS product does this, and while OpenOffice can theoretically, its a two program task (Calc->Draw->SVG) and I personally couldn't get the text to output. I finally found GraPL, which in my opinion was the best way to get a graph into SVG from a client-side spreadsheet and looking good. The trial version actually lets you import data and export to SVG, so if you are ok with not saving your graph settings you can do it for free.
Like I said, I'll try the first method first, but I'm honestly expecting it not to live up to my ideals. I'll give updates on how this goes, in addition to my explorations with other neat things, like an XML-driven JSP thesis-o-meter.
Choosing a Thesis Editor
- Simplicity - I can very easily get visual feedback on my document as I write it. While I really do appreciate the theoretical idea of separating presentation and content, its kind of annoying for me sometimes as a visual person.
- Collaboration - Versioning and edit tracking in Word is very good. Similar concepts in LaTeX just don't exist, you have to use custom commands or use very basic text diffing, which is not nearly as convenient.
- Portability - Word is very common and even if it wasn't, it's a simple to understand package. LaTeX is a "system" and a "language" and requires a lot more for a new user to understand.
However, the main disadvantage is that I have to spend several hours minimum switching a paper format when I resubmit it to a different conference. And that's with using all the Word helpers, like fully Styling the paper. With a thesis that is going to be much larger, I could definitely see things getting sort of hairy and presentation becoming inconsistent as I lose track of the whole document.
So I started off this morning by finally considering LaTeX. However, while it is very accepted and seems to do a decent job, it would essentially require me learning a new language. So before I ran off to B&N to sit down and start grinding through another presentation language, I looked into some other options and ended up at Prince XML.
Prince XML is a product that is developed by YesLogic in Australia. It takes XML and CSS, combines it, uses a couple of proprietary CSS tags, and generates a print-ready PDF file. The reason why this appeals to me is several-fold:
- It's a more web-specific technology. My research and personal interest is the web area so why I shouldn't I also use it for my thesis?
- I already know the technology. I've been writing CSS and XML for years, so the fundamental technology I already understand. Furthermore, I'm proficient in analyzing text in XML documents so I can more easily write cool tools like a thesis-o-meter.
- It's popular. XML and CSS are very well known, popular technologies. I can leverage that with a better suite of tools for source editing, display, and version tracking.
My only concern is that I won't have as much control over the output as with LaTeX, so I'll either have to spend time writing hacks or actually changing the content to make it easier for Prince. But at least for now I'm going to attempt to do it this way and see how far I get before I run into problems.
Thursday, June 7, 2007
Review: Reproduced and Emergent Genres of Communication on the World Wide Web
The first paper I've chosen to focus on is a web genre classification paper. Honestly, I've read less of these papers than anything else, mainly because there are so few. If you expand the field to general text genre classification, it gets a bit larger, but the core web genre classification field is limited to less than 25 peer-edited unique papers, and I believe I'm being very liberal with my categorization.
Reproduced and Emergent Genres of Communication on the World Wide Web was one of the first web specific genre papers out there. It was first published at the Hawaii International Conference on System Science in 1997, but a more thorough version of the paper appeared in The Information Society journal in 2000. It was written by Kevin Crowston and Marie Williams of Syracuse University (go SUNY!) who are both information scientist, not computer scientists. This is common to the field currently. It doesn't really become a CS field until it actually gets implemented or the machine learning takes place.
The core idea of the paper is that the genres that exist in traditional print and electronic text were changing as they moved onto the Web. New ones were being created and old ones were being changed. They showed this by sort of random sampling 100 URLs and then classifying each one into a genre and finally examining the genre list. Honestly, the mutation of communication patterns across media is really not my thing (a little too abstract for my CS liking), but I loved the paper as a first step toward establishing genre on the Web. Again, my interest in genre is purely from its applications to assisting the information search process.
Most of the papers they have written since then are somewhat of an extension of this work and some of them have more interesting things to talk about in my opinion, but I chose this piece to highlight because its the first paper that really brings genre classification to the Web specifically. Previous papers may have used the Web as an example or even gotten data from the Web, but usually the genres were text-specific. This was the first place we see Web-specific genres like home pages, hot lists (which are generally known as hubs in the HTML community), and server statistics used as a genre type. Frankly, I think some later papers that use non-Web specific genres should have taken a clue from this direction to do their classification studies.
The original paper was interesting because it actually lists the 100 pages they hand-classified and the genre they chose. Sometimes its refreshing to see actual data coding examples so you can get a more concrete grip on what genre really meant to the paper. However, it was kind of a small sample, they didn't cross-code for consistency, and they didn't do any work with organizing the genres, so the analysis feels a bit immature.
Those things were all improved in the later version of this paper. The sampling method was to actually get random URLs from Altavista, which was an improvement over the earlier manner. They also increased the sample size to 1000. This is a much better number when you figure that at the fine-grained granularity they define their genres to have, they had almost as many genres as documents in the 100 sample. They also did a much more thorough analysis of the genre types and the role they play followed by a very large future works section. In particular, I enjoyed the discussion on their sampling methodology, since I've dealt with the problem of creating good sampling methods using limited web databases (see my paper last year). Also, the point on pages as the unit of representation in genre classification is of particular interest to me in my research. The only problems I had with the later version was that it felt bloated in my mind. Included statistics on basic link and domain distribution had very little to do with the paper, nor did the section on web site design. Go write a web survey paper if that was your intention.
So just to summarize the main contribution: transitioned the concept of genre to Web-specific genres, creating the first Web-specific genre palette.
Tuesday, June 5, 2007
On Thumbnail Preview
More worried because (at the risk of sounding like a promoter) the feature is incredibly useful. By mousing over the links I can very quickly weed out spam-like sites that I'm not really interested in and get to the more appropriate ones that I need. Search UI assistance is particularly useful in dealing with broad queries, people who don't know how to form queries, and one word queries. In all of these cases, the results will be quite varied and manual filtering is required. By putting a thumbnail preview on the page, I take out a step for all of those query types. I don't have to page back and forth through results, looking for the one nugget of goodness.
However, I'm reassured because people are lazy and the process isn't perfect. First of all, sometimes the information is not there or is old or couldn't be rendered correctly for the preview. However, any classification system would suffer from similar issues, so I'll not fault them there. The fundamental problem is that the searcher still has to do the work. The positive aspect of this is that in general humans are smarter than machines. On the otherhand, they still have to go through the same process, it's just faster. In essence, the tool isn't really helping anyone find information, it's just making the task quicker.
Other aspects of Ask.com and Google do make my life easier. When they detect that I'm searching for a city and give me weather and maps and all sorts of neat things, that's taking a good guess at what I want to know. I think that genre classification falls more into this type of category. And that is what strengthens my resolve as I progress my research in that area.
Welcome
As way of introduction, I'm a PhD student in Information Retrieval at Binghamton University. I'm currently doing a dissertation on applying higher order HTML document features to several different fields of HTML analysis, currently focusing on genre classification. Genre classification is looking at consistent themes in web page design and information in order to help guide a user to the information they are looking for.
The main point of this blog is to serve as a brainstorming pad for my ideas on using visual features. A large area of research in academia is devoted to analyzing the text of an HTML web page, yet VERY little is devoted to analyzing the layout of the page. This is primarily because text documents were all the rage 50 years ago when most of the research was being done. Furthermore, in several tasks like content classification, its not really as important. However, in other fields like the up and coming genre classification, it can be much more useful.
While I believe starting a blog about academic research is in some ways dangerous (there's this pervasive academic atmosphere of quasi hiding what you're doing until you publish it or at least have the resources to do it better than anyone else), I think it should spur me to be quicker in my implementation of ideas.