I believe I finally have a final architecture for Webseer. Inspired by the architecture of Nutch, I have ripped almost everything out of Webseer and turned it into plugins. It turns out I couldn't use all the IO stuff in Nutch, since their APIs are set up to handle text parsing only, so I have defined the following extension points that define the core of Webseer:
FeatureExtractor - a visitor that collects features on a model
DataSourceIO - turns bytes into models
Transformer - transforms models into other models
I haven't figured out the classification side yet, but I need to abstract the common classification tasks into an interface. Something with more a little more power than just being able to plug in weka classes.
Showing posts with label architecture. Show all posts
Showing posts with label architecture. Show all posts
Sunday, October 7, 2007
Thursday, September 20, 2007
Nutch Rearchitecture
After first examining Nutch to use it for its crawling code, I've decided to rearchitect WebSeer to use the same basic architecture of reflection based extension points and a lot of their IO code. The codebase is almost an identical setup, with dynamic protocol/content type/parsing handling, only theirs has a plugin capability which is much more distribution friendly than WebSeer currently is. I believe this will enable me to finally focus WebSeer on what it actually is, an HTML analysis/feature extraction/classification library, not an IO library.
Friday, July 13, 2007
Reflection-based Visitor
In WebSeer, the program that does my web analysis, the main architecture is based on the concept of visitors analyzing models. So you go to a URL, convert the data stream into a local, easy-to-use object, and then run some analysis depending on the application. If you don't understand what the visitor pattern is, I advise you Google it and come back. I've been using the pattern for five years very heavily so I sort of take it for granted in my writing.
My existing problem is that I have several different types of visitors that are used for different purposes. So let's create an example - say I have a graph based model of something, with types Graph, Node, and Edge. Then I have a GraphVisitor that has three methods: visit(Graph), visit(Node), and visit(Edge). The accept methods in Graph, Node, and Edge take care of the default traversal. Now conveniently, I can print out all the edges in the graph or something else without having to care about the actual graph structure in my visitor code. Great...simple visitor pattern.
Now in my library, I have several different types of models, lets say I have another one that's Tree. Both Tree and Graph inherit the Model interface. Now I want to create a visitor that visits a list of Models without caring what type they are. We sort of have a problem. I can create a ModelVisitor, add an accept(ModelVisitor) method in the Model interface, and then check the type of the ModelVisitor and if it's also a GraphVisitor for instance, then do the component traversal. But that's messy. For the last long while, that's what I've been doing. However, in more interesting cases, it gets silly. You have to write accept methods for every type of visitor that could visit the node. Being as this is a visitor-centric architecture, this gets ridiculous even with inheritance. For instance, my TextNodeImpl class recently looked like:
Kinda crazy, right?
A while back I played around with reflection-based visitors, but I think I took it way too far (I also made the accepting objects reflection-based) and it got hard to understand because there was very little type-checking. However, I revisited it last night and I think I found a good balance.
All the visitors implement the SuperVisitor interface. You can either use a more type-safe generics version:
where DataSourceVisitor just has one "visit(T object) " method.
or the more handy and yet not as clean way:
where the visitor will attempt to visit any object and just fail quietly if it can't.
Then in AbstractSuperVisitor which takes care of the reflection, you initialize a Map with a mapping of requested Class to DataSourceVisitor. It initially puts in all the visit methods that are found in the runtime class. Then when a visitor for a particular runtime class is requested, you walk up the class hierarchy tree to find the appropriate visit method(s) that has already been initialized. Then you insert that new mapping into the tree. This caching is necessary to cut down lookup time, since reflection is pretty slow. And while this does create an extra object for each visit method, plus a Map for every Visitor pattern, you rarely have so many visitors that this would matter too much. So the primary cost is the cost of reflection based invocation. If this cost is important, I advise you look into articles that discuss how to replace reflection with byte code generation, which would work perfectly here.
So what this allows is that I can now create a Visitor that can visit any type of object on the fly. Let's say I want to print out everything this visitor comes across. Just put in a visit(Object o) method with a print line in the method and all the visit calls will be resolved to this method.
In addition, I now only need a single accept method in all my model classes, which calls the generic visit on itself and then passes the visitor to its components.
The design loss is in the loss of cohesion between the visitor and the things it's visiting. However, this can be mitigated by still using particular visitor interfaces that are sort of fake visitor guidelines for methods it should implement. So I can still have a GraphVisitor which enforces some methods on the visitor implementation, however, the Graph object will only see it as a SuperVisitor.
Overall, I think this is a neat solution to a model-based architecture where you have a lot of visitors interesting in different aspects of complex models. I also added a preVisit and a postVisit which gives me more control over post-children processing for instance. Hopefully this will be a good enough model for WebSeer.
My existing problem is that I have several different types of visitors that are used for different purposes. So let's create an example - say I have a graph based model of something, with types Graph, Node, and Edge. Then I have a GraphVisitor that has three methods: visit(Graph), visit(Node), and visit(Edge). The accept methods in Graph, Node, and Edge take care of the default traversal. Now conveniently, I can print out all the edges in the graph or something else without having to care about the actual graph structure in my visitor code. Great...simple visitor pattern.
Now in my library, I have several different types of models, lets say I have another one that's Tree. Both Tree and Graph inherit the Model interface. Now I want to create a visitor that visits a list of Models without caring what type they are. We sort of have a problem. I can create a ModelVisitor, add an accept(ModelVisitor) method in the Model interface, and then check the type of the ModelVisitor and if it's also a GraphVisitor for instance, then do the component traversal. But that's messy. For the last long while, that's what I've been doing. However, in more interesting cases, it gets silly. You have to write accept methods for every type of visitor that could visit the node. Being as this is a visitor-centric architecture, this gets ridiculous even with inheritance. For instance, my TextNodeImpl class recently looked like:
public class TextNodeImpl implements TextNode {
...
public void accept(TextVisitor visitor) {
visitor.visit(this);
}
public void accept(DataSourceVisitor visitor) {
visitor.visit(this);
}
public void accept(TreeVisitor visitor) {
visitor.visit(this);
}
public void accept(MutableTextVisitor visitor) {
visitor.visit(this);
}
}
Kinda crazy, right?
A while back I played around with reflection-based visitors, but I think I took it way too far (I also made the accepting objects reflection-based) and it got hard to understand because there was very little type-checking. However, I revisited it last night and I think I found a good balance.
All the visitors implement the SuperVisitor interface. You can either use a more type-safe generics version:
public interface SuperVisitor {
public boolean canVisit(Class<?> dataClass);
public <T> DataSourceVisitor<T>
getVisitor(Class<T> dataClass);
}
where DataSourceVisitor just has one "visit(T object) " method.
or the more handy and yet not as clean way:
public interface SuperVisitor {
public void visit(Object object);
}
where the visitor will attempt to visit any object and just fail quietly if it can't.
Then in AbstractSuperVisitor which takes care of the reflection, you initialize a Map with a mapping of requested Class to DataSourceVisitor. It initially puts in all the visit methods that are found in the runtime class. Then when a visitor for a particular runtime class is requested, you walk up the class hierarchy tree to find the appropriate visit method(s) that has already been initialized. Then you insert that new mapping into the tree. This caching is necessary to cut down lookup time, since reflection is pretty slow. And while this does create an extra object for each visit method, plus a Map for every Visitor pattern, you rarely have so many visitors that this would matter too much. So the primary cost is the cost of reflection based invocation. If this cost is important, I advise you look into articles that discuss how to replace reflection with byte code generation, which would work perfectly here.
So what this allows is that I can now create a Visitor that can visit any type of object on the fly. Let's say I want to print out everything this visitor comes across. Just put in a visit(Object o) method with a print line in the method and all the visit calls will be resolved to this method.
In addition, I now only need a single accept method in all my model classes, which calls the generic visit on itself and then passes the visitor to its components.
public class Graph {
...
public void accept(SuperVisitor visitor) {
visitor.visit(this);
for (Edge edge : edges) {
edge.accept(visitor);
}
}
}
The design loss is in the loss of cohesion between the visitor and the things it's visiting. However, this can be mitigated by still using particular visitor interfaces that are sort of fake visitor guidelines for methods it should implement. So I can still have a GraphVisitor which enforces some methods on the visitor implementation, however, the Graph object will only see it as a SuperVisitor.
Overall, I think this is a neat solution to a model-based architecture where you have a lot of visitors interesting in different aspects of complex models. I also added a preVisit and a postVisit which gives me more control over post-children processing for instance. Hopefully this will be a good enough model for WebSeer.
Subscribe to:
Posts (Atom)