Monday, December 17, 2007

Eclipse SWT-XPCOM Renderer

I made the most amazing (re)discovery last night. SWT makes the Mozilla DOM easily available via Java XPCOM. It solves all of my rendering analysis problems in one swoop; I really wish I had looked more into this a year ago, it would have saved me about a month of hassle messing with WebRenderer. Just to be fair to WebRenderer who gave me a free academic license for their product, they do make some of the basic stuff easier for non-hackers. But their API has some pretty serious deficiencies.

At first I was turned off because it appeared fairly heavyweight in dependencies and confusion. However, the swt library is one jar, the XPCOM libraries are one jar, and then you need a GRE installed (through XULRunner). Then use both the snippets from the SWT site and the XPCOM interface APIs and you can do almost everything you can from within native Mozilla code through Java fairly easily.

This is sort of bad timing for me since I'm currently in the middle of lockdown mode where I'm trying not to go off on coding binges and spend more time writing my dissertation. However, I may have to take a small detour, just for the sake of the library.

Just to go over several advantages of SWT-XPCOM over WebRenderer (sorry guys):
1) It gives me access to file size information through the http cache Mozilla has
2) It shuts down, which WebRenderer doesn't do...they hold open a native library listener until the system exits
3) It's easier to upgrade the mozilla version, you just get a new GRE
4) It doesn't require me to pass interesting page calls through serialized Javascript, I can directly access the native objects through JNI interfaces
5) Their JNI interfaces appear to be better written, some calls in WebRenderer take forever

Disadvantages:
1) Requires a GRE installed
2) Requires some understanding of how Mozilla code works (runtime interface casting)

Sunday, December 2, 2007

Plugin Frameworks

Lately I've been finalizing some of the architecture for WebSeer. It's built on Nutch, so it proscribes to the plugin framework that Nutch uses. However, I wanted to tweak the way it loads plugins slightly and found that Nutch is very statically tied to its plugin manager. Therefore, I went on a quest to find a better plugin framework.

I quickly found JPF, which is beautiful in the way they take the plugin architecture to the extreme. However, in my opinion, architecturally they took it too far when they allowed plugin initialization methods. This enables them to start applications without having a main class really. The entire application is just plugins and then you point a configuration file to the plugin that is the "main" plugin. Again, I appreciate the elegance of the solution, but realistically I think people like seeing the main application code. The one thing that JPF does that Nutch doesn't do is create their own classloaders, so the code spaces for plugins is separate. This is a really neat feature that I'd like to add to WebSeer eventually, but....

I ended right back at the Nutch plugin framework. It's not as powerful but I think it's easier to understand. This application has to be understandable by IR students and having core functionality wrapped in plugins would be hard for the ridiculously poor students I normally deal with to understand. Core interfaces and the code that brings those interfaces together to run something will be in the main library and implementations of those interfaces will be the plugins.

On a side note, anyone know the history of these plugin frameworks? JPF and the Nutch frameworks are suspiciously similar, Nutch is just a dumbed down version. I didn't look into the histories of these projects to see how they overlapped by one was definitely inspired by the other.

Sunday, October 7, 2007

Webseer Architecture

I believe I finally have a final architecture for Webseer. Inspired by the architecture of Nutch, I have ripped almost everything out of Webseer and turned it into plugins. It turns out I couldn't use all the IO stuff in Nutch, since their APIs are set up to handle text parsing only, so I have defined the following extension points that define the core of Webseer:
FeatureExtractor - a visitor that collects features on a model
DataSourceIO - turns bytes into models
Transformer - transforms models into other models

I haven't figured out the classification side yet, but I need to abstract the common classification tasks into an interface. Something with more a little more power than just being able to plug in weka classes.

Thursday, September 20, 2007

Nutch Rearchitecture

After first examining Nutch to use it for its crawling code, I've decided to rearchitect WebSeer to use the same basic architecture of reflection based extension points and a lot of their IO code. The codebase is almost an identical setup, with dynamic protocol/content type/parsing handling, only theirs has a plugin capability which is much more distribution friendly than WebSeer currently is. I believe this will enable me to finally focus WebSeer on what it actually is, an HTML analysis/feature extraction/classification library, not an IO library.

Sunday, September 2, 2007

Mozilla Flaw

Yeah, they probably know about this already and I'm too lazy to go check, but Firefox doesn't save web pages correctly, and therefore WebRenderer doesn't either. CSS pages are not localized, so they still point to relative references that don't exist. In addition, the web page isn't encoded correctly so on reload it has those non-renderable characters. I've fixed the problem somewhat by post-processing the files with pattern replacement to localize some of the CSS, but there are definitely multiple cases that slip by that, so it would be much better if they took care of it.

Monday, August 6, 2007

WebRenderer

WebRenderer was kind enough last month to give me a license for my academic work on WebSeer. They have one of the best Java-based HTML renderers. The secret is that it uses the actual Gecko codebase on the backend to do the rendering, so its fast, machine dependent, and exactly what you'd see in a Firefox browser.

However, up until this year, I wasn't interested. Why? To be honest, their API and event model was fairly annoying and hard to use. They don't give you very sophisticated DOM element methods, so you're forced to reconstruct JS calls on the document in order to get things like background color on a table cell. Their event model is shaky, though to be fair, it's a hard problem. Determining when a complex document is loaded is not easy I guess. Often I won't get document complete events and sometimes I'll get multiple if there are behind the scenes redirects. I understand that it's hard to avoid this with what HTML can do these days, but I want some of this encapsulated better.

The thing that made me reexamine WebRenderer is that they finally made it lightweight. The advantage of the "pure Java" implementations out there (IceSoft Browser for instance) is that they are lightweight so can be drawn on top of and tossed around without visualizing. (The main disadvantage is that the rendering is often pretty bad.) The main version of WebRenderer doesn't allow this type of rendering. It essentially dedicates a particular part of the screen to allowing Mozilla code to draw the HTML document. You can't touch it after it hands over control really, besides some integrated event handling.

It's nice to be able to highlight portions of the screen and interact with the actual document and the tool that I'm building within WebSeer to do labeling needs that ability. So I'm a strong proponent of WebRenderer now for this reasion. I am concerned that it still has some major flaws from a document analysis perspective. However, if you're just looking for integrated lightweight Java rendering, I'd highly recommend it.

Tuesday, July 31, 2007

Entropy Segmentation

I'm currently working on a model for page segmentation based on entropy. There's a paper already on this using the DOM, but I'm using a visual rendering so it's much more resilient. Anyway, the problem I'm having is that traditional entropy isn't really working for me. Web pages are very, very noisy visually speaking, more so than using a DOM. The goal is to attempt to cluster similar pieces together. If you do a top-down recursive algorithm to determine possible gaps and then apply entropy to pick a gap, at some point you end up favoring a split that reduces overall noise as opposed to clustering similar objects. An example is in order:

Elements are laid out horizontally on a page in this pattern: XXAXAXAXAXAX. For simplicity, A elements are all the same type of element and X elements have nothing in common at all. Therefore, only A elements are in the same class.

Ideally, we want the algorithm to split on the left side of the first A. This will maximize the order on the right side of the page, or otherwise trap all the As. Doing an entropy measure on the right side of that split (with the sign switched for ease), we get 5/10 lg 5/10 + 5 (1/10 lg 1/10) = 2.66. Now let's compare that with the split at the second A: 4/8 lg 4/8 + 4(1/8 lg 1/8) = 2.5. Apparently here, it's better to split at the second one because we're reducing the number of noise elements in the system. Also, there is no advantage to having larger classes, just less proportional ones.

So I've been playing with different ways of modifying the algorithm. Essentially, you want to make the noise elements worth less in determining entropy.