Tuesday, April 3, 2012
Day 1: Feature Extraction on the Web
Today I start writing my dissertation in earnest. I have a fairly regimented schedule that I plan on keeping over the next 25 days, including reading and writing, arith...programming. The dissertation has the working title of "Feature Extraction on the Web" and will focus on different methods for quantifying web pages for machine learning. The current plan is to first write an abstract to focus the papers direction, then write a rough outline, then to start piecing together my previously published works into a framework with which to fill in the gaps with a more cohesive story. I will be updating this blog every night with my progress, to give myself an opportunity to make sure I'm still on track. With such a tight timeframe, it's important that I don't get sidetracked on tangents or distractions.
Monday, February 20, 2012
Thesis Tools
So I've decided that I'm finally going to write a thesis. Will it be awesome and groundbreaking? Probably not. But at this point, I just want some closure to my degree and I'll settle for mildly interesting to some one in the same field. And I think when I have everything that I've already done written into a single document, it will become a bit obvious which things I need to shore up with more data.
A while back I investigated thesis writing options and gloriously put down some money for Oxygen XML editor. I'm not really sure why, but I think I thought it would be really cool to write in XML as an alternative to LaTeX, which I think is an amazingly awesome, yet completely overdone formatting (read programming) language. However, at some point I realized that what I needed was actually better tools to organize my thoughts and papers and not a document formatter. An XML editor, no matter how WYSIWYG and nifty, is just not set up for that task.
So I began to investigate document workflows on the Mac, which I have embraced as my OS of choice in the last several years. I downloaded MANY free trials and went through sample thesis generation with a bunch of them. And to cut to the chase, I'm going with Papers + Scrivener + Word. Yes, you probably could technically write with freebie latex tools and a large TextEdit window for 0 dollars, but I'd rather not waste my life on learning LaTeX and being employed has certain monetary advantages.
Papers ($50) - a research tool for organizing papers/bibliographies. It's a pretty slick interface that combines a lot of the bibliography management of completely overpriced tools like EndNote with a better UI, better reading/sharing ability with my iPhone, and a pretty intelligent citation injection system. With a thesis, where you are likely to have close to a 100 references, having something besides Word to manage them, index them, and make sure you don't miss something is pretty important. I've tried several other bib tools and this is definitely the easiest to use. My one complaint is that while it lets you essentially use Google Scholar to import papers, it could beef up its area of paper discovery and import a lot. In my opinion, this is one of the better things it brings to the table and the query language is sort of hidden and new paper notification/searching on more detailed fields would really knock it out of the park.
Scrivener ($45) - I thought a lot harder about this tool, because it's either a complete waste of money or an incredibly time-saving word processor depending on how you work. It focuses more on the front-side of the writing as opposed to the finished document layout. You write sections and then re-arrange, group, include resources to those sections. However, because my thesis is going to be a decent length and because it's mostly an amalgamation of past work, being able to push sections around will most likely be very important. And this tool sort of delinearizes the overall process, which I think will result in a better document. It's likely most of my writing will be done here.
Word ($100) - Ironically, even though this is probably the most powerful tool, it's the one I feel most ripped off about. However, when it comes to final layout power and reviewers/publishers needing docs in a common format, it's hard to beat. Yes, it's a MS product and it's cool to hate MS, but Word is ancient and polished and OpenOffice/Pages just doesn't beat it for most of the things it does. The final editing will be done here.
So I'll weight back in when I have more chance to use these in the next couple months. Wish me luck.
A while back I investigated thesis writing options and gloriously put down some money for Oxygen XML editor. I'm not really sure why, but I think I thought it would be really cool to write in XML as an alternative to LaTeX, which I think is an amazingly awesome, yet completely overdone formatting (read programming) language. However, at some point I realized that what I needed was actually better tools to organize my thoughts and papers and not a document formatter. An XML editor, no matter how WYSIWYG and nifty, is just not set up for that task.
So I began to investigate document workflows on the Mac, which I have embraced as my OS of choice in the last several years. I downloaded MANY free trials and went through sample thesis generation with a bunch of them. And to cut to the chase, I'm going with Papers + Scrivener + Word. Yes, you probably could technically write with freebie latex tools and a large TextEdit window for 0 dollars, but I'd rather not waste my life on learning LaTeX and being employed has certain monetary advantages.
Papers ($50) - a research tool for organizing papers/bibliographies. It's a pretty slick interface that combines a lot of the bibliography management of completely overpriced tools like EndNote with a better UI, better reading/sharing ability with my iPhone, and a pretty intelligent citation injection system. With a thesis, where you are likely to have close to a 100 references, having something besides Word to manage them, index them, and make sure you don't miss something is pretty important. I've tried several other bib tools and this is definitely the easiest to use. My one complaint is that while it lets you essentially use Google Scholar to import papers, it could beef up its area of paper discovery and import a lot. In my opinion, this is one of the better things it brings to the table and the query language is sort of hidden and new paper notification/searching on more detailed fields would really knock it out of the park.
Scrivener ($45) - I thought a lot harder about this tool, because it's either a complete waste of money or an incredibly time-saving word processor depending on how you work. It focuses more on the front-side of the writing as opposed to the finished document layout. You write sections and then re-arrange, group, include resources to those sections. However, because my thesis is going to be a decent length and because it's mostly an amalgamation of past work, being able to push sections around will most likely be very important. And this tool sort of delinearizes the overall process, which I think will result in a better document. It's likely most of my writing will be done here.
Word ($100) - Ironically, even though this is probably the most powerful tool, it's the one I feel most ripped off about. However, when it comes to final layout power and reviewers/publishers needing docs in a common format, it's hard to beat. Yes, it's a MS product and it's cool to hate MS, but Word is ancient and polished and OpenOffice/Pages just doesn't beat it for most of the things it does. The final editing will be done here.
So I'll weight back in when I have more chance to use these in the next couple months. Wish me luck.
Sunday, January 2, 2011
Assembled Structural Typing
I think I finally came up with the concept of typing I'm going to use in webseer. I've essentially been staring at the screen for the past several days doing nothing and trying to figure out how to make the system as powerful and usable as I imagine in my head. I've been reading programming language books and Wikipedia, in particular studying multiple inheritance, structural typing, duck-typing, along with a host of other interesting typing possibilities. I've been search for a typing system that is much more flexible than a traditional typing system. Obviously, there is no ONE way to do something, everything has benefits and liabilities. But I've finally settled on something that doesn't make me squirm for my own vision. It's going to be something like structural typing, only facilitating generic type structures more easily. Otherwise, the inputs to the functional units will be typed, but typed only in their labeled structure. Types can be saved with names to facilitate ease of usage in the function implementations, but in the graphical editor, the types can be generically assembled on the way in to the function to give you the ability to morph object structures to fit the function. In complex data types, this type of structure induction would be very difficult, but for most of the cases I'm considering it's quite straightforward.
Saturday, November 20, 2010
Webseer Update
It's been several years since I started this blog and I've been working at ITA Software in that time. My PhD project, webseer, which is what this blog was mainly about has been progressing slowly. I think sisyphean would be a brilliant word to describe the project as I've restarted it so many times I've lost count. I still have a core set of utility functions and the data analysis bits that I wrote my initial papers with, but the actual library/app has gone through so many changes it's barely recognizable from the original Swing app that I wrote to transform HTML sources into visual models that I could write segmentation algorithms with.
The latest effort that I've been working on the last two years on my weekends and vacations has been a web application. The primary goal of the web application is to be a visual programming language (think Yahoo Pipes) that allows 1) storage of large sets of data, 2) type-specific visualization of said data, and 3) complete transparency with regards to data sourcing. This language would allow people to generate their own plugin functions as uploaded Java classes that would allow them to easily customize how datasets are generated to be able to compare different methods of data collection/extraction. The theory being that many classification problems are more dependent on proper feature generation than on algorithm selection. In particular, HTML pages tend to have many more dimensions than traditional text classification or measurement problems. Additionally, measuring the efficacy of these different approaches are often fraught with problems in program consistency - "the devil is in the details".
Implementation-wise, the main problems I've faced have been 1) spending way too much on the UI and 2) making the language very complete - particularly with regards to multiple inputs/outputs on generic functional units. Doing the synchronization on streams while maintaining a stored data trail is a bitch of coding. I finally got past a lot of the first problem by shrinking the UI into a single main interface and giving up the concept of visualization being built into the graph language. I might revisit at some time in the future, but it was getting messy.
The second problem is still something I struggle with, trying to balance all of the tracking with a framework that allows easy plugin creation.
The latest effort that I've been working on the last two years on my weekends and vacations has been a web application. The primary goal of the web application is to be a visual programming language (think Yahoo Pipes) that allows 1) storage of large sets of data, 2) type-specific visualization of said data, and 3) complete transparency with regards to data sourcing. This language would allow people to generate their own plugin functions as uploaded Java classes that would allow them to easily customize how datasets are generated to be able to compare different methods of data collection/extraction. The theory being that many classification problems are more dependent on proper feature generation than on algorithm selection. In particular, HTML pages tend to have many more dimensions than traditional text classification or measurement problems. Additionally, measuring the efficacy of these different approaches are often fraught with problems in program consistency - "the devil is in the details".
Implementation-wise, the main problems I've faced have been 1) spending way too much on the UI and 2) making the language very complete - particularly with regards to multiple inputs/outputs on generic functional units. Doing the synchronization on streams while maintaining a stored data trail is a bitch of coding. I finally got past a lot of the first problem by shrinking the UI into a single main interface and giving up the concept of visualization being built into the graph language. I might revisit at some time in the future, but it was getting messy.
The second problem is still something I struggle with, trying to balance all of the tracking with a framework that allows easy plugin creation.
Tuesday, March 25, 2008
Optimizing Reflective Visitor with Runtime Compilation
Last night I had the pleasure of doing some optimizations I'd been wanting to do for some time to the core visitor model relationship in webseer. Basically, all of webseer runs on the visitor pattern and a specially magical reflective visitor (see previous blog on this) that does runtime method lookup and invocation. As a quick reminder, I want to be able to write
public class MyVisitor extends ReflectiveSuperVisitor {
public void visit(HTMLTag tag) {
// do something
}
}
pass this to any object that accepts visitors and have it just run that code on anything that matches a HTMLTag. It's just a generic visitor pattern with the added coolness that there is no compile-time linking so there is no need to implement a particular visitor interface or have the model know about what visitors will visit it. In my mind, this is how the visitor pattern should be. Otherwise, it becomes too tedious for its own good in large systems.
I cache the runtime method lookups (essentially the same lookup that Java does at compile time to link to the right method) so that part isn't that intensive over many calls, but you still have to deal with a reflective method invocation. Because this is so core to all the transformations and feature extraction in webseer, I knew this had to be fundamentally faster for anyone to take it seriously.
A while back I read this great article that describes how to convert reflective calls into runtime compiled method calls. It fit perfectly and with about 15 lines of code of Javassist I was able to dramatically cut down the time of method invocation. I didn't get quite as dramatic speedups as his results (he was doing reflective lookups and several method invocations and it's possible that JVMs are faster at this now), but in my simple test it more than doubled the speed of the calls. I also changed the method invocations to take both the visitor and the acceptor so there can be one for each class in a static cache as opposed to per visitor like it was before - this cuts down both on memory overhead as well as GC time. Essentially I now dynamically generate classes that look like this:
public class MyVisitorvisitHTMLTag implements MethodCaller {
public void callFor(ReflectiveSuperVisitor visitor, Acceptor tag) {
((MyVisitor)visitor).visit((HTMLTag)tag);
}
}
Now when you call visitor.visit(this); in an accept(SuperVisitor) method, after the first lookup, it performs the following steps:
public class MyVisitor extends ReflectiveSuperVisitor {
public void visit(HTMLTag tag) {
// do something
}
}
pass this to any object that accepts visitors and have it just run that code on anything that matches a HTMLTag. It's just a generic visitor pattern with the added coolness that there is no compile-time linking so there is no need to implement a particular visitor interface or have the model know about what visitors will visit it. In my mind, this is how the visitor pattern should be. Otherwise, it becomes too tedious for its own good in large systems.
I cache the runtime method lookups (essentially the same lookup that Java does at compile time to link to the right method) so that part isn't that intensive over many calls, but you still have to deal with a reflective method invocation. Because this is so core to all the transformations and feature extraction in webseer, I knew this had to be fundamentally faster for anyone to take it seriously.
A while back I read this great article that describes how to convert reflective calls into runtime compiled method calls. It fit perfectly and with about 15 lines of code of Javassist I was able to dramatically cut down the time of method invocation. I didn't get quite as dramatic speedups as his results (he was doing reflective lookups and several method invocations and it's possible that JVMs are faster at this now), but in my simple test it more than doubled the speed of the calls. I also changed the method invocations to take both the visitor and the acceptor so there can be one for each class in a static cache as opposed to per visitor like it was before - this cuts down both on memory overhead as well as GC time. Essentially I now dynamically generate classes that look like this:
public class MyVisitorvisitHTMLTag implements MethodCaller {
public void callFor(ReflectiveSuperVisitor visitor, Acceptor tag) {
((MyVisitor)visitor).visit((HTMLTag)tag);
}
}
Now when you call visitor.visit(this); in an accept(SuperVisitor) method, after the first lookup, it performs the following steps:
- invoke ReflectiveSuperVisitor.visit(Acceptor) - SUPER FAST
- get hashcode of runtime visitor class - SUPER FAST
- lookup MethodCaller in hashtable - FAST
- invoke MethodCaller - SUPER FAST
- invoke correct visit method - SUPER FAST
- invoke correct visit method - SUPER FAST
Monday, December 17, 2007
Eclipse SWT-XPCOM Renderer
I made the most amazing (re)discovery last night. SWT makes the Mozilla DOM easily available via Java XPCOM. It solves all of my rendering analysis problems in one swoop; I really wish I had looked more into this a year ago, it would have saved me about a month of hassle messing with WebRenderer. Just to be fair to WebRenderer who gave me a free academic license for their product, they do make some of the basic stuff easier for non-hackers. But their API has some pretty serious deficiencies.
At first I was turned off because it appeared fairly heavyweight in dependencies and confusion. However, the swt library is one jar, the XPCOM libraries are one jar, and then you need a GRE installed (through XULRunner). Then use both the snippets from the SWT site and the XPCOM interface APIs and you can do almost everything you can from within native Mozilla code through Java fairly easily.
This is sort of bad timing for me since I'm currently in the middle of lockdown mode where I'm trying not to go off on coding binges and spend more time writing my dissertation. However, I may have to take a small detour, just for the sake of the library.
Just to go over several advantages of SWT-XPCOM over WebRenderer (sorry guys):
1) It gives me access to file size information through the http cache Mozilla has
2) It shuts down, which WebRenderer doesn't do...they hold open a native library listener until the system exits
3) It's easier to upgrade the mozilla version, you just get a new GRE
4) It doesn't require me to pass interesting page calls through serialized Javascript, I can directly access the native objects through JNI interfaces
5) Their JNI interfaces appear to be better written, some calls in WebRenderer take forever
Disadvantages:
1) Requires a GRE installed
2) Requires some understanding of how Mozilla code works (runtime interface casting)
At first I was turned off because it appeared fairly heavyweight in dependencies and confusion. However, the swt library is one jar, the XPCOM libraries are one jar, and then you need a GRE installed (through XULRunner). Then use both the snippets from the SWT site and the XPCOM interface APIs and you can do almost everything you can from within native Mozilla code through Java fairly easily.
This is sort of bad timing for me since I'm currently in the middle of lockdown mode where I'm trying not to go off on coding binges and spend more time writing my dissertation. However, I may have to take a small detour, just for the sake of the library.
Just to go over several advantages of SWT-XPCOM over WebRenderer (sorry guys):
1) It gives me access to file size information through the http cache Mozilla has
2) It shuts down, which WebRenderer doesn't do...they hold open a native library listener until the system exits
3) It's easier to upgrade the mozilla version, you just get a new GRE
4) It doesn't require me to pass interesting page calls through serialized Javascript, I can directly access the native objects through JNI interfaces
5) Their JNI interfaces appear to be better written, some calls in WebRenderer take forever
Disadvantages:
1) Requires a GRE installed
2) Requires some understanding of how Mozilla code works (runtime interface casting)
Sunday, December 2, 2007
Plugin Frameworks
Lately I've been finalizing some of the architecture for WebSeer. It's built on Nutch, so it proscribes to the plugin framework that Nutch uses. However, I wanted to tweak the way it loads plugins slightly and found that Nutch is very statically tied to its plugin manager. Therefore, I went on a quest to find a better plugin framework.
I quickly found JPF, which is beautiful in the way they take the plugin architecture to the extreme. However, in my opinion, architecturally they took it too far when they allowed plugin initialization methods. This enables them to start applications without having a main class really. The entire application is just plugins and then you point a configuration file to the plugin that is the "main" plugin. Again, I appreciate the elegance of the solution, but realistically I think people like seeing the main application code. The one thing that JPF does that Nutch doesn't do is create their own classloaders, so the code spaces for plugins is separate. This is a really neat feature that I'd like to add to WebSeer eventually, but....
I ended right back at the Nutch plugin framework. It's not as powerful but I think it's easier to understand. This application has to be understandable by IR students and having core functionality wrapped in plugins would be hard for the ridiculously poor students I normally deal with to understand. Core interfaces and the code that brings those interfaces together to run something will be in the main library and implementations of those interfaces will be the plugins.
On a side note, anyone know the history of these plugin frameworks? JPF and the Nutch frameworks are suspiciously similar, Nutch is just a dumbed down version. I didn't look into the histories of these projects to see how they overlapped by one was definitely inspired by the other.
I quickly found JPF, which is beautiful in the way they take the plugin architecture to the extreme. However, in my opinion, architecturally they took it too far when they allowed plugin initialization methods. This enables them to start applications without having a main class really. The entire application is just plugins and then you point a configuration file to the plugin that is the "main" plugin. Again, I appreciate the elegance of the solution, but realistically I think people like seeing the main application code. The one thing that JPF does that Nutch doesn't do is create their own classloaders, so the code spaces for plugins is separate. This is a really neat feature that I'd like to add to WebSeer eventually, but....
I ended right back at the Nutch plugin framework. It's not as powerful but I think it's easier to understand. This application has to be understandable by IR students and having core functionality wrapped in plugins would be hard for the ridiculously poor students I normally deal with to understand. Core interfaces and the code that brings those interfaces together to run something will be in the main library and implementations of those interfaces will be the plugins.
On a side note, anyone know the history of these plugin frameworks? JPF and the Nutch frameworks are suspiciously similar, Nutch is just a dumbed down version. I didn't look into the histories of these projects to see how they overlapped by one was definitely inspired by the other.
Subscribe to:
Posts (Atom)