Thursday, October 18, 2012

Programming is getting easier

So I haven't actually tried to work on the software I wrote for my PhD in about a year and a half.  So imagine my surprise when yesterday I created a new Mercurial project in Eclipse, connected to my source code repository that I actually had to search for to find, cloned the repository, and everything worked literally right out of the box.  My application was running within 10 seconds of me finishing the Eclipse wizard....that's crazy.

I can't explain the number of times I've started the project over practically from scratch because I couldn't reproduce the infrastructure that held all the feature extraction components together.  I usually can track down the code for the actual extraction, but it's usually very difficult to get everything running again.  Things in the world of IDE and collaboration are definitely moving in the right direction.

Now, before I found my repository, I was glancing at THE FUTURE.  The future of programming as I see it is a couple of things:

  1. Instant data feedback - seeing the effect the code you right has on logic flows instantaneously...this will close the debug/test loop almost completely, though I'm curious to see how it scales.  For an example of an IDE that champions this cause, see Light Table.
  2. Online coding - this is something like the Eclipse Orion project or Koding.  Not having to install an IDE and having a common environment is going to be incredible for team development.  Sorry, emacs is not the future..though I'm sure there's a macro for it.
  3. Cloud services - I think the future result of programming in general is web based services, both backend and frontend, but that's not a very novel vision...most people think this.  Amazon is practically betting half their company on it.  It's just stagnating waiting for critical adoption, simplicity, and standards to come along.
Those concepts combine to make a nasty combination of fast development, collaboration, and deployment....mmm...I hope I'm around long enough to see where this leads.

Friday, April 27, 2012

Day 19: It's not the worst thing I've read

Current Paper - 26,800 words/80 references

I did a mostly complete read-through of my paper yesterday.  I came up with almost two hundred things that still need to be done, even assuming I'm not fleshing out certain sections that I don't have data for.  That sounds daunting, but most of them were small.

The best thing was that it's not bad.  I know from experience that other readers will find certain sections incomprehensible and it's pretty light on meat.  A lot of content is just discussion and unproven best practice, which is less than ideal for a computer science dissertation.  But if it gets to where I see it in my head with some of the diagrams and lists and some more of the interesting detail stuff, it definitely won't be the worst dissertation I've read.

Most importantly, it's really the dissertation that I wanted to write.  It's the topics that I care about instead of the topics that I think would look the most impressive.  My academic strength are not discovering novel information retrieval models, or finding ways to improve machine learning algorithms.  My strength is that I was a web programmer before I was a web researcher and I respect the value of the complexity of a Web page.

My plan is to spend the rest of the day trying to knock off that to-do list and then spend tomorrow working on diagrams and figures that will go a long way toward breaking up some of the page-long descriptions...people love pictures.  If all goes well, I'll send a rough draft of the paper to my advisor on Sunday and see what she thinks.

The best part about this is that it's now to the point where I would have to be an idiot to fail.  I have 77 pages (which may be closer to 100 after formatting and diagrams, etc.) of fairly cohesive text that will serve as a framework to put things even if it needs to get longer.  There is no 10-ton albatross hanging around my neck anymore.  It's more like a finch....ok, maybe a raven.

Thursday, April 26, 2012

Day 18: Academic Duels


Current Paper - 26,500 words/75 references

Everyday I'm reminded how much I don't want to be in academia.  I don't really like writing very much, but love coding, and love having tangible things to get working.  However, a lot of the days over the course of this month I have felt a strong desire to publish.  Probably for the wrong reason...anger.

Many, many days over the last month I have read papers that say stupid things that really, really makes me want to send them emails.  The email addresses are right there, on their papers.  I've even started a couple.  They just sound so flamey and arrogant that I close them all after two sentences.

I generally don't approve of direct paper bashing, but I want to exemplify my point by ranting about a paper I actually read yesterday that I had only skimmed before.
<rant>
This paper is written by Serge Sharoff and published in 2010.  Now, I happen to know for a fact that he knows about my research because he was one of the editors on this paper in that I wrote in 2008 which I obtain better accuracy than his paper obtains on the KI-04 document base by using visual features along with a bunch of other HTML features.  There is very secondary mention of HTML features and zero mention of anything I've done on one of the same tasks he experiments on.  Because some of the corpora he's evaluating are flawed, he decides to completely abandon any of those HTML/visual features and convert everything to text.  Now, I happen to know that several of the genre classification tasks the paper performs can obtain better accuracy numbers than most of the text accuracies that are presented using just the URL, so this move is insane in my opinion.  Not to mention, the paper is called "Web Library of Babel" and there is nothing "Web"y about any of the documents the way they are being compared.

He then goes on to proclaim that the entire field is doomed because of lack of correlation between current document bases.  That may actually be true, I definitely agree with the overall point about the lack of good corpora, and the idea of measuring the SVM vector correlations was a good one, but without what I consider to be a complete set of features (that I've proven have significance in two separate papers) with which to measure that correlation, the results are meaningless or at least not meaningful enough to make such a huge declaration.  It only proves that the features that have been chosen don't generalize across experiment because they are overfitting on stupid textual clues because everything else was thrown away.  And because human beings were doing the choosing, it makes intuitive sense to me that features that measuring whatever visual cues they were choosing to pick the class would probably be useful and more likely to cross-generalize.
</rant>
That gives you a good idea of what is going through my head when I read these papers.  I am just itching to run a parallel experiment that shows that this conclusion was not only overstated, but actually flawed.

People are clearly not reading the things that I write, even years after I've published them in the same field.  Or they are and they don't trust me.  I'm not sure which is worse.  I don't have many points that I try to make in my papers, so I think I'm being pretty clear.  I'm pretty much a one-tricky pony - "Stop treating Web pages as text documents, they are not the same thing!"

Over the course of academic history, there is a way that people have channeled this anger productively.  They write a paper duplicating what the person has done, and proving that they are wrong.  I don't have time to do that right now and I know my anger will fade as I go back to focusing on work...so I have decided we need to reinstate a different time honored tradition.  I will challenge authors I disagree with to a duel.  High noon.  Woburn.  If you're not there, I win and you have to retract your conclusions.

Wednesday, April 25, 2012

Day 17: Barebones WebSeer

Current Paper - 25,500 words/70 references

WebSeer is the name that I branded the software that performed my web survey and maybe the work for my masters, I forget.  Over the last eight years I have lived with this software that was constantly changing, never done, and never useful to anyone else.  At several stable points in the past, I attempted to teach other graduate students to use it and it was too abstract to grasp.  I started talking about code elegance and reflection-based visitor patterns or dataflow programming and the several people who had a vested interest in understanding couldn't use it.  Some of that was lack of documentation and "getting started" type stuff.  But in the end, the core problem was that I was obsessed with building frameworks as opposed to getting stuff working.  My overhead on actual functionality has been generally about 90% of my effort I'd say.

So given that I only had what amounts to several days of time to work on programming this month, I chose to start as barebones as I could get, while giving some thoughts to the next logical step.  As someone who has been through a lot of frameworks in my days, mostly web ones, frameworks suck.  These days I usually start off looking for the bare simplest solution I can find that looks sufficiently robust and go from there.  I think a lot of people already have that figured out, but it took me a while.

By the end of this week, I should have completed a fairly small and straightforward library of classes that have simple transformation methods.  They input one thing and output another.  Generally, we start with a URL and eventually end up with features (string->double) measurements with enough of these methods.  The things that I input and output right now are protocol buffer structures.  This was useful because it avoided me having to maintain data structure classes.  Since I use a bunch of different models, this would have been wasted time writing boilerplate.

So for starters, you will be able to download these simple little classes (which sometimes wrap more complex libraries) and quickly have a whole lot more Web page features than you know what to do with, incorporating textual features, tag features, and visual features.  Great, simple case solved and if everything else goes to hell, something will be usable.

From there we grow from reasonable all the way up to unreasonable:
1. A demo in a browser that you type in a URL and you have a way to show the measurements that are generated.  These will be uncheckable by model and feature generation function so you can filter down the list and be able to interpret the mass of features in a piece-meal manner.  I can do this with a single page I think...keep it simple.

2. A way to take whatever features you've checked and download just a self-contained set of jars and a feature generator facade that will just produce those features, for running real systems.  Some of the features require several libraries and often in your experiments you find that certain classes of features are just not that useful, so this will get you up and running quickly without doing any real effort.

3. Web services wrapping these little transformation functions so that they can be used from other languages.  I have several ideas in mind, but haven't settled on anything.  WSDL is more supported and has service addressing built in if its on top of HTTP, but I hate XML pretty intensely.

The eventual goal is to have these written in multiple languages and to have the methods and data structures exposed in a more language-unspecific documentation and implementation manner so that language doesn't have to be a barrier.

Tuesday, April 24, 2012

Day 16: Bibliographic Nightmares

Current Paper - 24,900 words/72 references

New version of Papers today led me into a 2-hour slog of bibliography management.  The bibliography was a bit of a mess in the thesis.  I've been importing papers fairly rapidly and not fixing them at all, so titles are often wrong, author names are sometimes missing, and a lot of the conference/journal names were incorrect.

Bibliographies in general are somewhat of a nightmare in computer science and I imagine all of academia.  There are several good formats for expressing the meta-information in a common manner so you can import it, but while most papers and citations are available, their meta information is consistently different depending on where you see the citation.  So one citation might call it the Proceedings of the 12th annual conference of topic X and another might call it WEBX '13 or some abbreviation or different representation.  Most databases are not totally comprehensive both in meta-information and paper coverage, so it's not possible to just use one completely.

Recent Author Publications
I use Papers for my bibliography management and it has some serious strengths and some annoying weaknesses (actually, most of them are just bugs).  It recently put back in recent author search which is a pretty nice feature that does a lookup by year when you click on the author name.  This lets me make sure I haven't missed a closely related paper by an author I know writes in my area.  Note the buggy layout in the search results window.  It also does this for journals, but in computer science 90% of the publications are in proceedings so it's not quite as useful for me.

It also has an integrated general search that piggy backs on online databases to allow direct meta-information import into the program.  That's nice and I use it a lot but strangely there is no way to do a Google Scholar locate to directly download the PDF after you have the meta-information.  Half the time I need to download the file myself and attach it to the record.  And sometimes I even have to get it from the publisher's site, which requires going through my university's library.  The weird double-standard with regards to copyrights on academic papers strikes again.

It also has primary entities for conferences and authors and periodicals so you can pivot and organize things all nicely, so that's what I spent most of my time doing this morning - opening up cited papers and fixing meta information that either OCR messed up or was incomplete in the database that it was retrieved from.  I do believe this stuff is getting better than it was ten years ago, but it's a pretty slow improvement path for such a fast-moving field.

Monday, April 23, 2012

Day 15: Last Push Panic

Current Paper - 23,700 words/67 references

This is the last week!

My friend Mike Oltmans told me on Saturday that I didn't appear depressed and stressed enough to be working on a PhD.  So I just want to assure everyone that right now I'm feeling that trapped, panicky feeling in my mind and the pit of my stomach.  I only have a week left to pull together a lot of the parts of the thesis that are still in fragments or grossly underdeveloped.  Then there is this little voice in the back of my mind that says "Wait, this isn't at all what I was going to do, this is completely not going to work".

Right now the core chapter that I've been working on is starting to look like a pretty good and comprehensive survey of the field with regards to Web page models and features, better than any I've read (in scope at least).  It still is lacking my own contributions in those sections though besides some side comments about first-hand usefulness, which is not good.

Additionally, my papers on genre classification need to be rewritten to make more sense in the context of the rest of the paper.  Right now they are sort of jammed in there without too much linear flow.

To get more data I need to program, which as I've said before is a dangerous path.  I already spent this morning coding when I should have been writing.  It happened to be a pretty fruitful session, but I still can't afford it.  I need to get this thing into readable state and chances are I won't have data by the end of this week anyway so I need to continue to enforce the lower priority of that.  I plan on abandoning some of my pace and luxuries I use to keep myself going and just sprinting for the end.  I should be able to handle that for a week.  Wish me luck!

Friday, April 20, 2012

Day 14: Channeling the Inner Writer

Current Paper - 22,700 words/65 references

As I've mentioned before, I'm not much of a writer by nature.  Whenever I start writing, I itch to go write code.  It's so much more satisfying to build something functional, I never really get the urge to influence people with words.  However, you can't get a PhD with code alone, sadly.  So I've had to really work hard to channel my inner writer this month.

Along the way, I've come up with several tricks for writing my dissertation that I thought I would share for those who have to write a nonfiction book or dissertation or any large-scoped piece of writing.


  1. Read something - if I hit a roadblock in my writing, I go read something.  Generally it inspires me to want to react to that thing that I've read by filling out a certain part of the thesis.
  2. Break it down - just like code, it's good to break the overall thing down into manageable pieces.  Whenever I look at a large blank section, it's intimidating.  Our brains aren't meant to handle scale like that.  So you create smaller sections that contain pieces that are much more on a detailed level that you can write a couple paragraphs about.  On the right, you can see what my paper looks like in Scrivener.  It's a huge outline view.  Generally, I pick a section to work on and fill it out.
  3. Write in a small font - when you write in a large font, you feel too good about the amount of content you're writing.  It's misleading.
  4. Scan the overall export once a day - If I only ever looked at it in outline form, I wouldn't get a sense of how the flow is for another person.  You need to constantly look at the paper as if you are a new reader.
  5. Don't sweat the figures/images - These take a bit of time to look right and make the paper too pretty.  Again, misleading confidence.  Give them placeholders and save them for the end if they aren't already created.  Then you'll be pleasantly surprised when it looks better.

Thursday, April 19, 2012

Day 13: Musical Stimulants

Current Paper - 21,500 words/59 references

I was going to write on a different topic this morning, but I happened to buy a new pair of headphones (which I apparently overpaid for now that I see that link) this morning and thought I'd share how awesome that was.

I've talked a lot about routine and pace and all sorts of boring yet important stuff so far with regards to my process.  I haven't really talked about passion.  Control and passion are opposite sides of the same coin, ego and id, but I'd argue both are important to any process.  Most good teams that I meet working at Google have both, either in separate people that balance the team or in a single person who drives the team in a controlled way.

Personally, my passion is driven by self-delusion.  Seriously.  It's driven by the thought that I'm doing something amazing, revolutionary, ground-breaking, and that I am a hero to the entire world.  Maybe it was reading too much epic fantasy as a child, but this is what gives me my narcissistic rush.  And unfortunately, the more I live in the real world, the more this just doesn't make any sense.

This is where music comes in.  When I put on headphones, especially when I'm around other people, it feels like I'm surfing the sea of humanity.  Feeling people move around me but not requiring actual interaction makes you feel powerful, like Neo wandering through a human simulation.  I've always wanted to watch my brain waves when I listen to music at good moments.  I can feel the chemicals and thoughts swirling in my mind.

Regardless, music is my drug and its way more powerful than caffeine for stimulating effort.  However, it is dangerous.  Like I said, my passion is related to divorcing my senses from reality.  And reality keeps you focused on writing that next chapter, keeps you from reading useless papers because they sound so AWESOME!, and keeps you from spending too much time on useless introspection...oh, shoot.  Stupid headphones.

Wednesday, April 18, 2012

Day 12: The Approach

Current Paper - 21,200 words/59 references

As a follow-up to yesterday's background, my main goal with my dissertation is to bring it back to what I started with, ways to represent Web pages. To that end, the core of the paper is going to be about models to represent Web pages and the ways in which they can be used to extract features and how those features can be useful. I hope the paper to be useful to anyone, but particularly new researchers, as a source for ideas when they start to pull information out of Web pages.

As an introduction to these models and extraction methods, I currently have a chapter on the actual methodology for model transformation and extraction. I propose an identifier-driven approach for model structure (which I did in my last publication), where a URI represents a data structure as well as imposed semantics. Transformations are themselves also defined by identified specification. I see something like this as being crucial for research on Web pages moving forward because there is just too much complexity of information and so much is wasted effort right now and not verifiable or comparable because of the complexity of implementation. It's too common to read a paper that writes something like, "Rather than compare this against X (where X is what they've introduced as state of the art) we evaluated this against Y, because 1. we couldn't get in touch with the authors, 2. the implementation was too complex, or 3. we ran out of time, leave for future work" :).

My goal with the programming is actually to begin exposing some of these models and extraction methods in implementation as web services and a web server to browse the specifications. There are all sorts of questions about how I would actually expose these, control access to these, etc. but all I care right now is getting them working.

I haven't figured out a good way to evaluate features in a more task-insensitive way and I think it may be an impossible problem, so I'm considering throwing the feature sets at a couple of very researched problems (standard topical clustering of Web documents, an IR problem, and a genre classification problem) to just serve as examples. Once I have the feature services in place, I can download the corpuses, run them through, and throw the data into Weka for some fast and sloppy numbers.

This informal survey will then serve as a lead-in to the genre classification chapter where I really drill into comparing feature sets and aggregating feature sets in a more serious way.

The thesis will survive without those feature evaluations, but that chapter is going to read like a survey chapter, even if I'm throwing in additional features and modeling approaches that I've found useful.

Tuesday, April 17, 2012

Day 11: In the Beginning...

Current Paper - 20,500 words/53 references

I realized today that I've been very careful to keep this blog meta, which is probably safer but a bit out of scope - it should be called "My PhD Journey" or something. So I'm going to refocus by explaining a bit of the history of what I'm working on.

When I started doing IR research way back around 2003-4, it disturbed me that Web pages were being treated as text documents. At that time, I was very familiar with Web technologies and I knew that these pages were getting very sophisticated and yet most research I was reading was still all about the text-driven approaches, tf idf, etc. Because of styles and script at the time, the text that appeared on the page could theoretically be entirely different or emphasized in a completely different way than what was being ripped by a parser that just ignored the HTML tags.

At that time, styles weren't quite as sophisticated or used and I was a little overconfident and thought I could write a Web browser in order to render (draw) the page to get around some of these problems. The results of that approach was my masters thesis. It has some good points, but since then I've realized that I should have started off wrapping a real browser and so when I read it now it looks like a lot of wasted effort.

However, even though I failed to write a browser on my own, I now had experience and some code that could do some interesting things to Web pages. I replaced the rendering piece with a JRex Mozilla wrapper that actually did some rendering with the Mozilla rendering engine and used it to do some surveys of the Web. The idea was to prove to people that these Web technologies were being used a lot and should be paid attention to. The result of that was my first published paper.

About this time, I began to realize that even if I could generate these interesting bits of information about the Web page, a lot of them were completely uninteresting to IR tasks. However, I also stumbled onto genre classification as a subfield of Web page classification. Here was a field that was meant for what I was working on. Web page genres are incredibly based on the functionality and visual display of the document as opposed to its text. So I wrote a couple of papers on that and continued to refine my models of Web pages.

Then I went off and tried to write a UI to experiment with transforming these models to extract features to feed into a Weka sandbox to experiment with. This would let people play with extracting Web page features in much the same way that Weka allows newbies to machine learning to play with machine learning tasks. That was a mistake and cost me several years. The idea was not bad in theory, but I wrote it top-down and got hung up on the UI and the framework and forgot the goal wasn't the system but making it easier for people to render and create these models. I have a bit of a hobby of visual programming languages and at one point I almost had a complete UI driven asynchronous dataflow programming language. It was pretty cool but crazily out of scope.

So now here I am, with most of my published research in genre classification and most of my actual time spent writing a visual programming language when really I'm academically interested in Web page representational models - the best ways to extract and use all the rich information that authors are creating. This was the challenge I faced when I started this month.

Tomorrow I'll talk about how I'm trying to address this in my dissertation.

Monday, April 16, 2012

Day 10: Pace

Today is Marathon Monday or the day of the Boston Marathon for those outside of Boston. My friend Josh Ain is running in the marathon this year. He told everyone that he'd be running an 8 minute pace so we could look for him at certain times. As a fairly driven person I never really understood pace. What if you feel great that day? Why not run as hard as you can? Ok, I haven't really had that thought since I was a child losing races at the bus stop, but it has taken a longer time to get used to the idea of pace in other areas of my life.

In college, I was a master of the all-nighter. I loved all-nighters. Coding deep into the night to finish that project, often doing the work for a whole team of people who you didn't trust to write the code...hmm...ok, actually nothing about that is admirable. But it was amazing to see what you could get done in a very short amount of time when your whole mind was dedicated to a task.

You can't pull all-nighters every night for a month. Even if you could physically force yourself to sit in front of a screen for a month straight, you would find yourself writing worse and worse and slower and slower and then you would panic and then your brain would start scrambling and you would burn out, decide it wasn't worth it, take several days off, and repeat the whole cycle. I've experienced these things before, they're not fun.

Writing a dissertation for me is about pace, routine, and control. Every day I wake up at the same time and go to the gym. The gym is important, it definitely makes me feel good about myself even if I'm not writing as well as I want.

Then I head to my office. My office rotates between several libraries, Starbucks, and Paneras. I am never in one location for more than three hours. Sometime between 2 and 3 hours I start losing focus, almost without exception. So I need to switch, this is important.

Each segment has a focus that I decide either at the start of the day or at the end of the last segment. This is also important, a segment without a focus is a segment that will be more likely to fail. I generally have three of these segments during the day. They tend to get less productive over the course of the day. I'm brilliant in the morning and moronic near the end of the day. It's much better for me to take an afternoon off and work on Saturday morning if I can afford it.

In the end I probably average about 7 hours of work a day, if you include the half an hour that I tend to spend on writing this blog and don't include the couple hours that I program every night. That's not a very impressive amount of time. But in the end, the amount of time that I've spent staring at the screen not knowing what to write is very, very small. And I have yet had to force myself to work; I'm still pretty excited about what I'm doing. And that, I count as impressive.

Friday, April 13, 2012

Day 9: Information Overload = N * E

Current Paper - 18,600 words/43 references

I've started including the number of references even though it's sort of a silly metric, especially since several of them don't even have a good description yet. They're just being sorted into sections by where they belong.

Dealing with the large amount of papers in a dissertation is quite a challenge. I'm a bit curious how people managed before databases, lots of index cards I guess. I use Papers to do my bibliography management and paper browsing, which is a pretty awesome program. All the papers that I find relevant go in there and get keywords assigned to them like "ML Techniques", "Model:Visual", "Search Results Ranking". These keywords are what help me find what I need amongst the current 110 documents I have looked at, but it's also neat to sort them by publication date to get an overview of how my field evolved. A large majority of these will eventually find their way into my paper.

Some will not. I've run into several cases now of papers that are really related to what I'm doing but are really awful papers. In particular, they have one of the cardinal sins that I see in paper writing of completely useless formalization. This is my biggest pet peeve. I have an aversion to useless greek letters/variables that is pretty intense. There are many places where they are useful because it's impossible to clearly state what you mean without a complex paragraph of English text, but you don't need them to establish how smart you are and you should think a bit before you introduce them. Everyone knows HTML is a hierarchical tree-like tag language, some examples or a drawing could probably do the trick. If you redefine HTML as a (V,E) graph language and you don't need that abstraction to prove something or to use some interesting graph algorithm, then it's not actually helping the reader understand HTML.

Thursday, April 12, 2012

Day 8: Literature Review Panic

Current Paper: 18,500 words/43 references

Up until now, I've really gotten away with much heavy reading. I've pretty much been pushing text around that I already had and figuring out where I wanted to put future content. I'm starting to hit a wall with that though and now I really need inspiration from other sources. Especially since a large part of my thesis is a good survey of feature extraction methods on the Web, I will have a lot of discussion about other people's techniques.

So yesterday I began reading and searching for literature. This is mostly because I really haven't looked at the state-of-the-art since I left graduate school. At that time, not only was there less literature on the area I'm interested in, but the publication channels didn't get picked up quite so fast by search tools. So there is a lot of research that was going on at the time that I wrote some of my early papers that very much parallelled what I was doing that I never knew about.

When I find a new paper that is in my area that I didn't know about, especially later on in my research, I go through a very predictable series of emotions.

First, there is panic: "This paper title is exactly what I'm doing. Oh no, my whole research is redundant. I'm going to have to start over."

Then, I start reading the paper and my defensiveness kicks in: "Wait, this is nothing like the way I did it and the scope of the paper is so much smaller than the title. Why would you do it that way? This is so flawed."

Finally, there is acceptance: "Ok, I see how this could be useful. Maybe I should adapt the way I word my approach to this specific part to point out the benefits versus this existing method."

Which is pretty much the entire reason for a literature review, so you get to that eventual state where you can give the reader a better understanding of in what ways your approach/thought/algorithm is an improvement to the world versus the existing state.

Wednesday, April 11, 2012

Day 7: My Free Time

Current Paper - 18,500 words/36 references

I only have a couple hours of work today because of various family commitments, so I'm going to give a brief overview of what I've been doing with my evenings.

I've been spending about two hours each evening doing some programming in support of my thesis. The state of my code as I last left it was terrifying when I opened up the project, so I pretty much restarted with the primary goal being simple and immediately useful.

In a very short amount of time, I already have managed to return rendered HTML documents, mostly due to the awesome Webkit wrapper phantomjs. If that rendering tool had existed when I started my PhD, it most likely would have literally saved me six months of effort over the last several years as I floundered about with other tools that only half worked and even tried to write my own browser.

So I've managed to get a protocol buffer server (thanks to protobuf-rpc-pro and no thanks to Google for not providing a binary protobuf compiler for OSX) running in Amazon EC2 that serves me position annotated HTML documents for a given URL. There's still a couple of timing issues with AJAX evaluation, but to get as far as I have in about ten hours is quite awesome. By the end of the week, I hope to be able to generate a rough feature vector for given URLs, which should enable me to lace some of my content with some interesting new sampling data. When I feel more confident in the completeness of that data coming back after I've added more feature calculations from old code, I'll move into spending some time running learning experiments.

Tuesday, April 10, 2012

Day 6: Information Models

Current Paper - 17,700 words/33 references

I'm sort of at a dangerous place right now in my thought process. The outline of my paper is roughly:

Introduction
Data Survey
Feature Discussion
Genre Classification Experiments
Conclusion

The data survey and classification experiments are at least cohesive and written, barring some extra data I might throw in there if I have time. The biggest wildcard in that outline are the feature discussion chapters.

The feature stuff is sort of what I'm interested in and what I've spent a bulk of my time over my research actually doing. But I don't have a lot of content written and I don't really have a good undirected evaluation of these features or any sort of formal models to describe how those features are extracted. Without those things, it will mean the core of the thesis is really weak and I'll need to make the thesis about genre classification where I have actual meat.

So I'm trying to figure out some good ways to discuss representational models of information, without much of a background to work with. I read a thesis on information retrieval models which gave me a good starting point. However, modeling interactions between word sequences is significantly easier to grasp and describe than modeling linear streams representing hierarchical information that is interpreted to a 2d visual screen.

I also have spent a bunch of time really figuring out what information entropy is and I'm getting closer. The thought there is that I can somehow quantify the difference between feature sets derived from different document models by measuring cluster entropy in an unsupervised classification. Models that better represent the underlying information will result in less entropy in their derived feature sets. This is because the randomness of information is decreased in more accurate models because a human being can bring more semantic patterns as probability estimates. One important thing is in representing the unused information as an entropy gain in the less-good models. If you don't do this correctly, then the better models or feature sets could potentially look worse just because they fundamentally are including more information. It just sounds messy and error-prone.

Anyway, I'm hoping to find some information about this on the Web, because its starting to go over my head. If anyone can understand what I'm talking about and point me in a good direction, I'm more happy to just run the experiments and not try to prove a theoretical point that a smarter person already made.

Monday, April 9, 2012

Day 5: Arrogance and Insecurity

Current Paper - finally got the bibliography generating correctly

One of my friends, Bryan Clark, noted last week in g+ that I should begin posting copies of the drafts of my dissertation as I'm working for accountability purposes. Now honestly, I'm quite sure no one is reading them, but he had a good point to be more externally open. With that goal, last week I sent a copy of my abstract and outline to a member of my committee and he got back to me with good feedback.

Technically, finishing a PhD requires writing a dissertation that sums up your research and establishes you as an expert in your chosen field. More practically, it really just requires convincing several smart people that make up your committee that you're also smart and have checked all the boxes off. A good dissertation will help with this, but the sign-off is far more crucial.

Viewed in this fairly cynical light, the goal of writing a dissertation should not be to change the world or prove something revolutionary. For one, if you are counting on a single document you write ever doing that, you're delusional. Even great thinkers and inventors of the past didn't work this way.

So if the goal is to convince your committee members of your worth, the best approach is to involve them in the process as much as you can without sounding annoying or useless. This may sound somewhat mercenary with that goal in mind, but the truth is it's crucial in any field to get this feedback. I think that sounds obvious to a lot of people, but I've avoided collaboration in the past for several reasons:
  • I don't want to ever look dumb; I'm very proud
  • I don't want to give up my ideas
  • I'm worried others won't see my vision and will ruin what I see as the eventual goal
Over time I've gradually realized these were all really bad reasons:
  • You become smarter much faster by asking stupid questions and asking for help. The short term loss of pride is made up for by a much better longer term gain.
  • No one wants your ideas. You have to prove they are good before people will even grudgingly admit they are ok.
  • If your ideas can't stand up to criticism, they weren't worthwhile to start off. You should abandon that vision and find another one that you can better defend.
These ideas are actually all offshoots of a single sense of insecurity that I often have: What if this is my best idea and I can't come up with any more?

This combination of arrogance and insecurity is the cause for a lot of problems in the world, not just writing a PhD. Too many smart people in the world fall into one of these traps, and a lot of them are academics.

Friday, April 6, 2012

Day 4: Only 15,000 words?

Current Paper

I don't consider myself to be a good writer. My mom always said I wrote too conversationally, which I took to mean that I shoved together sentence fragments with punctuation and used a conversational tone. But I do think that I have the ability to write densely. When I wrote academic papers I never had the problem of too many pages. I would write a couple sentences and then go back and delete half of them as unimportant and distracting. I have a tendency to assume my reader has the exact same background as myself and thus trim out what I consider to be obvious points so you're only left with the "true genius" of the paper :). You could probably argue that that is almost always the wrong approach. And with a dissertation, that's especially true.

After I pulled in most of the text minus the introductions from my papers, my thesis is currently at 15,000 words. I don't really care much about word count as an end goal, but it does give you a rough idea of the type of content length and depth that you're shooting for. I've read that most science theses are somewhere on the order of 35-40k words and humanities, because they are less data driven are somewhere more around the 70-80k word counts.

If you look at the current state, you'll see a lot of the content is currently graphs and tables, which doesn't do much to help word counts but is pretty crucial to the importance and readability of a document like this. In my own mind I was roughly thinking that I would be at least half done with the raw content after I pulled in my papers. It seems this was a bit of an overestimate, but luckily I still do have a bunch of ideas on where I want the paper to go and sections I want to be fleshed out with background and referenced work, so I'm not worried yet.

Thursday, April 5, 2012

Day 3: Pulling it all Together

Current Working Paper

There's the link for accountability purposes. However, I advise no one actually look at it for at least another week even though I will post a new version every day. It is currently in a fairly unbalanced state. This is because my first task I gave myself was pulling in the papers that I have already published into a consistent text.

Like I said in previous posts, I've already published one poster, two papers, and one journal article. The dissertation is going to be an amalgamation of these papers with a different balance to tie them all together. To give you a bit of flavor of what I'm planning on doing, the summary of my papers/articles are:
  1. A data survey of technology usage (HTML/CSS/JavaScript) on the Web.
  2. A study on using visual feature extraction (what the page looks like) to improve genre classification (figuring out what the "type" of a Web page is).
  3. A framework for evaluating whether this more complex information is "worth it" to calculate for Web pages or whether more simple information will be good enough.
All of this work really revolves around the idea that Web pages have more going on than plain text and ways that you can use this information to improve the way we analyze Web pages. So the point of the dissertation is to take that core idea and expand on it. I'll do this by focusing on the actual features that all of my papers used as well as explaining how others in the field have used this extra information. Then I'll finish it all by using genre classification as a use case where I have experimental proof that this extra information is definitely useful in improving precision. By focusing primarily on the features, I think I can add something that I didn't cover exhaustively in my papers and possibly even have something that the reader can use in their own experiments. My only concern is that this will read a bit like a reference manual, so I'm going to try to balance technical detail by putting the really technical stuff in the appendix.

I currently have two of the papers copied/segmented into the dissertation with all their references corrected. I have one more to piece in and then hopefully by the end of today I will have a chance to rearrange the documents to put the overlapping pieces into their appropriate sections. Luckily, the editor I'm using, Scrivener is really made for large-scale document rearrangement.

Wednesday, April 4, 2012

Day 2: Focus on Focus

Yesterday was a very productive day. But the first day is never the problem.

There are two major slumps to look out for with motivation. First, the day after a productive day you usually feel pretty good about yourself and it's harder to motivate that extra effort. Second, there always a slump in the middle of a time period when you doubt your entire direction and think you have too much to finish. One of the main things you use to avoid these slumps are schedules that have detailed focus and scope.

My research has often had a focus problem. This is very evident in reading my PhD prospectus/proposal. A prospectus to those unfamiliar with PhD work is where you state what you've done already and what you're planning to include in your dissertation. It's usually done fairly soon before you actually write the document so it's mostly there to sort of sign off on the dissertation before you put a lot of work into it. My prospectus has a future works section that includes the following:
  • Run a sampling survey of the entire Web and produce another set of technology statistics akin to the first paper I ever published as a followup
  • Write a feature extraction framework, including a formal methodology and UI-driven application
  • Create a demo of a genre-annotated web search index and then perform a user study on how genre affects web searching
Yeah. And I wonder why I haven't gotten anywhere yet. Each one of those items is at least a single paper in itself if not a PhD, not to mention the programming work that is sort of mentioned offhand and ended up being the black hole of time over the last several years. This is a good example of how not to write a prospectus. It was written this way because I was unsure of the scope of a prospectus, afraid it would not look impressive enough and so I just included everything I could think of that I wanted to do.

So my new plan for my dissertation is to focus on focus. I have a body of work already that I could probably stretch into a dissertation, without any extra programming/data effort. The goal is to do that with the bulk of my time and then have a single targeted goal to give a little extra juice to the dissertation. It will be worked on in the evening time, as an optional more fun thing. Not gating the bulk of your work on a potential black hole of effort is utterly crucial and the one thing I wish I could go back in time and convince myself of.

Tuesday, April 3, 2012

Day 1: Feature Extraction on the Web

Today I start writing my dissertation in earnest. I have a fairly regimented schedule that I plan on keeping over the next 25 days, including reading and writing, arith...programming. The dissertation has the working title of "Feature Extraction on the Web" and will focus on different methods for quantifying web pages for machine learning. The current plan is to first write an abstract to focus the papers direction, then write a rough outline, then to start piecing together my previously published works into a framework with which to fill in the gaps with a more cohesive story. I will be updating this blog every night with my progress, to give myself an opportunity to make sure I'm still on track. With such a tight timeframe, it's important that I don't get sidetracked on tangents or distractions.

Monday, February 20, 2012

Thesis Tools

So I've decided that I'm finally going to write a thesis. Will it be awesome and groundbreaking? Probably not. But at this point, I just want some closure to my degree and I'll settle for mildly interesting to some one in the same field. And I think when I have everything that I've already done written into a single document, it will become a bit obvious which things I need to shore up with more data.

A while back I investigated thesis writing options and gloriously put down some money for Oxygen XML editor. I'm not really sure why, but I think I thought it would be really cool to write in XML as an alternative to LaTeX, which I think is an amazingly awesome, yet completely overdone formatting (read programming) language. However, at some point I realized that what I needed was actually better tools to organize my thoughts and papers and not a document formatter. An XML editor, no matter how WYSIWYG and nifty, is just not set up for that task.

So I began to investigate document workflows on the Mac, which I have embraced as my OS of choice in the last several years. I downloaded MANY free trials and went through sample thesis generation with a bunch of them. And to cut to the chase, I'm going with Papers + Scrivener + Word. Yes, you probably could technically write with freebie latex tools and a large TextEdit window for 0 dollars, but I'd rather not waste my life on learning LaTeX and being employed has certain monetary advantages.

Papers ($50) - a research tool for organizing papers/bibliographies. It's a pretty slick interface that combines a lot of the bibliography management of completely overpriced tools like EndNote with a better UI, better reading/sharing ability with my iPhone, and a pretty intelligent citation injection system. With a thesis, where you are likely to have close to a 100 references, having something besides Word to manage them, index them, and make sure you don't miss something is pretty important. I've tried several other bib tools and this is definitely the easiest to use. My one complaint is that while it lets you essentially use Google Scholar to import papers, it could beef up its area of paper discovery and import a lot. In my opinion, this is one of the better things it brings to the table and the query language is sort of hidden and new paper notification/searching on more detailed fields would really knock it out of the park.

Scrivener ($45) - I thought a lot harder about this tool, because it's either a complete waste of money or an incredibly time-saving word processor depending on how you work. It focuses more on the front-side of the writing as opposed to the finished document layout. You write sections and then re-arrange, group, include resources to those sections. However, because my thesis is going to be a decent length and because it's mostly an amalgamation of past work, being able to push sections around will most likely be very important. And this tool sort of delinearizes the overall process, which I think will result in a better document. It's likely most of my writing will be done here.

Word ($100) - Ironically, even though this is probably the most powerful tool, it's the one I feel most ripped off about. However, when it comes to final layout power and reviewers/publishers needing docs in a common format, it's hard to beat. Yes, it's a MS product and it's cool to hate MS, but Word is ancient and polished and OpenOffice/Pages just doesn't beat it for most of the things it does. The final editing will be done here.

So I'll weight back in when I have more chance to use these in the next couple months. Wish me luck.