Saturday, November 20, 2010

Webseer Update

It's been several years since I started this blog and I've been working at ITA Software in that time. My PhD project, webseer, which is what this blog was mainly about has been progressing slowly. I think sisyphean would be a brilliant word to describe the project as I've restarted it so many times I've lost count. I still have a core set of utility functions and the data analysis bits that I wrote my initial papers with, but the actual library/app has gone through so many changes it's barely recognizable from the original Swing app that I wrote to transform HTML sources into visual models that I could write segmentation algorithms with.

The latest effort that I've been working on the last two years on my weekends and vacations has been a web application. The primary goal of the web application is to be a visual programming language (think Yahoo Pipes) that allows 1) storage of large sets of data, 2) type-specific visualization of said data, and 3) complete transparency with regards to data sourcing. This language would allow people to generate their own plugin functions as uploaded Java classes that would allow them to easily customize how datasets are generated to be able to compare different methods of data collection/extraction. The theory being that many classification problems are more dependent on proper feature generation than on algorithm selection. In particular, HTML pages tend to have many more dimensions than traditional text classification or measurement problems. Additionally, measuring the efficacy of these different approaches are often fraught with problems in program consistency - "the devil is in the details".

Implementation-wise, the main problems I've faced have been 1) spending way too much on the UI and 2) making the language very complete - particularly with regards to multiple inputs/outputs on generic functional units. Doing the synchronization on streams while maintaining a stored data trail is a bitch of coding. I finally got past a lot of the first problem by shrinking the UI into a single main interface and giving up the concept of visualization being built into the graph language. I might revisit at some time in the future, but it was getting messy.

The second problem is still something I struggle with, trying to balance all of the tracking with a framework that allows easy plugin creation.