Thursday, October 18, 2012

Programming is getting easier

So I haven't actually tried to work on the software I wrote for my PhD in about a year and a half.  So imagine my surprise when yesterday I created a new Mercurial project in Eclipse, connected to my source code repository that I actually had to search for to find, cloned the repository, and everything worked literally right out of the box.  My application was running within 10 seconds of me finishing the Eclipse wizard....that's crazy.

I can't explain the number of times I've started the project over practically from scratch because I couldn't reproduce the infrastructure that held all the feature extraction components together.  I usually can track down the code for the actual extraction, but it's usually very difficult to get everything running again.  Things in the world of IDE and collaboration are definitely moving in the right direction.

Now, before I found my repository, I was glancing at THE FUTURE.  The future of programming as I see it is a couple of things:

  1. Instant data feedback - seeing the effect the code you right has on logic flows instantaneously...this will close the debug/test loop almost completely, though I'm curious to see how it scales.  For an example of an IDE that champions this cause, see Light Table.
  2. Online coding - this is something like the Eclipse Orion project or Koding.  Not having to install an IDE and having a common environment is going to be incredible for team development.  Sorry, emacs is not the future..though I'm sure there's a macro for it.
  3. Cloud services - I think the future result of programming in general is web based services, both backend and frontend, but that's not a very novel vision...most people think this.  Amazon is practically betting half their company on it.  It's just stagnating waiting for critical adoption, simplicity, and standards to come along.
Those concepts combine to make a nasty combination of fast development, collaboration, and deployment....mmm...I hope I'm around long enough to see where this leads.

Friday, April 27, 2012

Day 19: It's not the worst thing I've read

Current Paper - 26,800 words/80 references

I did a mostly complete read-through of my paper yesterday.  I came up with almost two hundred things that still need to be done, even assuming I'm not fleshing out certain sections that I don't have data for.  That sounds daunting, but most of them were small.

The best thing was that it's not bad.  I know from experience that other readers will find certain sections incomprehensible and it's pretty light on meat.  A lot of content is just discussion and unproven best practice, which is less than ideal for a computer science dissertation.  But if it gets to where I see it in my head with some of the diagrams and lists and some more of the interesting detail stuff, it definitely won't be the worst dissertation I've read.

Most importantly, it's really the dissertation that I wanted to write.  It's the topics that I care about instead of the topics that I think would look the most impressive.  My academic strength are not discovering novel information retrieval models, or finding ways to improve machine learning algorithms.  My strength is that I was a web programmer before I was a web researcher and I respect the value of the complexity of a Web page.

My plan is to spend the rest of the day trying to knock off that to-do list and then spend tomorrow working on diagrams and figures that will go a long way toward breaking up some of the page-long descriptions...people love pictures.  If all goes well, I'll send a rough draft of the paper to my advisor on Sunday and see what she thinks.

The best part about this is that it's now to the point where I would have to be an idiot to fail.  I have 77 pages (which may be closer to 100 after formatting and diagrams, etc.) of fairly cohesive text that will serve as a framework to put things even if it needs to get longer.  There is no 10-ton albatross hanging around my neck anymore.  It's more like a finch....ok, maybe a raven.

Thursday, April 26, 2012

Day 18: Academic Duels


Current Paper - 26,500 words/75 references

Everyday I'm reminded how much I don't want to be in academia.  I don't really like writing very much, but love coding, and love having tangible things to get working.  However, a lot of the days over the course of this month I have felt a strong desire to publish.  Probably for the wrong reason...anger.

Many, many days over the last month I have read papers that say stupid things that really, really makes me want to send them emails.  The email addresses are right there, on their papers.  I've even started a couple.  They just sound so flamey and arrogant that I close them all after two sentences.

I generally don't approve of direct paper bashing, but I want to exemplify my point by ranting about a paper I actually read yesterday that I had only skimmed before.
<rant>
This paper is written by Serge Sharoff and published in 2010.  Now, I happen to know for a fact that he knows about my research because he was one of the editors on this paper in that I wrote in 2008 which I obtain better accuracy than his paper obtains on the KI-04 document base by using visual features along with a bunch of other HTML features.  There is very secondary mention of HTML features and zero mention of anything I've done on one of the same tasks he experiments on.  Because some of the corpora he's evaluating are flawed, he decides to completely abandon any of those HTML/visual features and convert everything to text.  Now, I happen to know that several of the genre classification tasks the paper performs can obtain better accuracy numbers than most of the text accuracies that are presented using just the URL, so this move is insane in my opinion.  Not to mention, the paper is called "Web Library of Babel" and there is nothing "Web"y about any of the documents the way they are being compared.

He then goes on to proclaim that the entire field is doomed because of lack of correlation between current document bases.  That may actually be true, I definitely agree with the overall point about the lack of good corpora, and the idea of measuring the SVM vector correlations was a good one, but without what I consider to be a complete set of features (that I've proven have significance in two separate papers) with which to measure that correlation, the results are meaningless or at least not meaningful enough to make such a huge declaration.  It only proves that the features that have been chosen don't generalize across experiment because they are overfitting on stupid textual clues because everything else was thrown away.  And because human beings were doing the choosing, it makes intuitive sense to me that features that measuring whatever visual cues they were choosing to pick the class would probably be useful and more likely to cross-generalize.
</rant>
That gives you a good idea of what is going through my head when I read these papers.  I am just itching to run a parallel experiment that shows that this conclusion was not only overstated, but actually flawed.

People are clearly not reading the things that I write, even years after I've published them in the same field.  Or they are and they don't trust me.  I'm not sure which is worse.  I don't have many points that I try to make in my papers, so I think I'm being pretty clear.  I'm pretty much a one-tricky pony - "Stop treating Web pages as text documents, they are not the same thing!"

Over the course of academic history, there is a way that people have channeled this anger productively.  They write a paper duplicating what the person has done, and proving that they are wrong.  I don't have time to do that right now and I know my anger will fade as I go back to focusing on work...so I have decided we need to reinstate a different time honored tradition.  I will challenge authors I disagree with to a duel.  High noon.  Woburn.  If you're not there, I win and you have to retract your conclusions.

Wednesday, April 25, 2012

Day 17: Barebones WebSeer

Current Paper - 25,500 words/70 references

WebSeer is the name that I branded the software that performed my web survey and maybe the work for my masters, I forget.  Over the last eight years I have lived with this software that was constantly changing, never done, and never useful to anyone else.  At several stable points in the past, I attempted to teach other graduate students to use it and it was too abstract to grasp.  I started talking about code elegance and reflection-based visitor patterns or dataflow programming and the several people who had a vested interest in understanding couldn't use it.  Some of that was lack of documentation and "getting started" type stuff.  But in the end, the core problem was that I was obsessed with building frameworks as opposed to getting stuff working.  My overhead on actual functionality has been generally about 90% of my effort I'd say.

So given that I only had what amounts to several days of time to work on programming this month, I chose to start as barebones as I could get, while giving some thoughts to the next logical step.  As someone who has been through a lot of frameworks in my days, mostly web ones, frameworks suck.  These days I usually start off looking for the bare simplest solution I can find that looks sufficiently robust and go from there.  I think a lot of people already have that figured out, but it took me a while.

By the end of this week, I should have completed a fairly small and straightforward library of classes that have simple transformation methods.  They input one thing and output another.  Generally, we start with a URL and eventually end up with features (string->double) measurements with enough of these methods.  The things that I input and output right now are protocol buffer structures.  This was useful because it avoided me having to maintain data structure classes.  Since I use a bunch of different models, this would have been wasted time writing boilerplate.

So for starters, you will be able to download these simple little classes (which sometimes wrap more complex libraries) and quickly have a whole lot more Web page features than you know what to do with, incorporating textual features, tag features, and visual features.  Great, simple case solved and if everything else goes to hell, something will be usable.

From there we grow from reasonable all the way up to unreasonable:
1. A demo in a browser that you type in a URL and you have a way to show the measurements that are generated.  These will be uncheckable by model and feature generation function so you can filter down the list and be able to interpret the mass of features in a piece-meal manner.  I can do this with a single page I think...keep it simple.

2. A way to take whatever features you've checked and download just a self-contained set of jars and a feature generator facade that will just produce those features, for running real systems.  Some of the features require several libraries and often in your experiments you find that certain classes of features are just not that useful, so this will get you up and running quickly without doing any real effort.

3. Web services wrapping these little transformation functions so that they can be used from other languages.  I have several ideas in mind, but haven't settled on anything.  WSDL is more supported and has service addressing built in if its on top of HTTP, but I hate XML pretty intensely.

The eventual goal is to have these written in multiple languages and to have the methods and data structures exposed in a more language-unspecific documentation and implementation manner so that language doesn't have to be a barrier.

Tuesday, April 24, 2012

Day 16: Bibliographic Nightmares

Current Paper - 24,900 words/72 references

New version of Papers today led me into a 2-hour slog of bibliography management.  The bibliography was a bit of a mess in the thesis.  I've been importing papers fairly rapidly and not fixing them at all, so titles are often wrong, author names are sometimes missing, and a lot of the conference/journal names were incorrect.

Bibliographies in general are somewhat of a nightmare in computer science and I imagine all of academia.  There are several good formats for expressing the meta-information in a common manner so you can import it, but while most papers and citations are available, their meta information is consistently different depending on where you see the citation.  So one citation might call it the Proceedings of the 12th annual conference of topic X and another might call it WEBX '13 or some abbreviation or different representation.  Most databases are not totally comprehensive both in meta-information and paper coverage, so it's not possible to just use one completely.

Recent Author Publications
I use Papers for my bibliography management and it has some serious strengths and some annoying weaknesses (actually, most of them are just bugs).  It recently put back in recent author search which is a pretty nice feature that does a lookup by year when you click on the author name.  This lets me make sure I haven't missed a closely related paper by an author I know writes in my area.  Note the buggy layout in the search results window.  It also does this for journals, but in computer science 90% of the publications are in proceedings so it's not quite as useful for me.

It also has an integrated general search that piggy backs on online databases to allow direct meta-information import into the program.  That's nice and I use it a lot but strangely there is no way to do a Google Scholar locate to directly download the PDF after you have the meta-information.  Half the time I need to download the file myself and attach it to the record.  And sometimes I even have to get it from the publisher's site, which requires going through my university's library.  The weird double-standard with regards to copyrights on academic papers strikes again.

It also has primary entities for conferences and authors and periodicals so you can pivot and organize things all nicely, so that's what I spent most of my time doing this morning - opening up cited papers and fixing meta information that either OCR messed up or was incomplete in the database that it was retrieved from.  I do believe this stuff is getting better than it was ten years ago, but it's a pretty slow improvement path for such a fast-moving field.

Monday, April 23, 2012

Day 15: Last Push Panic

Current Paper - 23,700 words/67 references

This is the last week!

My friend Mike Oltmans told me on Saturday that I didn't appear depressed and stressed enough to be working on a PhD.  So I just want to assure everyone that right now I'm feeling that trapped, panicky feeling in my mind and the pit of my stomach.  I only have a week left to pull together a lot of the parts of the thesis that are still in fragments or grossly underdeveloped.  Then there is this little voice in the back of my mind that says "Wait, this isn't at all what I was going to do, this is completely not going to work".

Right now the core chapter that I've been working on is starting to look like a pretty good and comprehensive survey of the field with regards to Web page models and features, better than any I've read (in scope at least).  It still is lacking my own contributions in those sections though besides some side comments about first-hand usefulness, which is not good.

Additionally, my papers on genre classification need to be rewritten to make more sense in the context of the rest of the paper.  Right now they are sort of jammed in there without too much linear flow.

To get more data I need to program, which as I've said before is a dangerous path.  I already spent this morning coding when I should have been writing.  It happened to be a pretty fruitful session, but I still can't afford it.  I need to get this thing into readable state and chances are I won't have data by the end of this week anyway so I need to continue to enforce the lower priority of that.  I plan on abandoning some of my pace and luxuries I use to keep myself going and just sprinting for the end.  I should be able to handle that for a week.  Wish me luck!

Friday, April 20, 2012

Day 14: Channeling the Inner Writer

Current Paper - 22,700 words/65 references

As I've mentioned before, I'm not much of a writer by nature.  Whenever I start writing, I itch to go write code.  It's so much more satisfying to build something functional, I never really get the urge to influence people with words.  However, you can't get a PhD with code alone, sadly.  So I've had to really work hard to channel my inner writer this month.

Along the way, I've come up with several tricks for writing my dissertation that I thought I would share for those who have to write a nonfiction book or dissertation or any large-scoped piece of writing.


  1. Read something - if I hit a roadblock in my writing, I go read something.  Generally it inspires me to want to react to that thing that I've read by filling out a certain part of the thesis.
  2. Break it down - just like code, it's good to break the overall thing down into manageable pieces.  Whenever I look at a large blank section, it's intimidating.  Our brains aren't meant to handle scale like that.  So you create smaller sections that contain pieces that are much more on a detailed level that you can write a couple paragraphs about.  On the right, you can see what my paper looks like in Scrivener.  It's a huge outline view.  Generally, I pick a section to work on and fill it out.
  3. Write in a small font - when you write in a large font, you feel too good about the amount of content you're writing.  It's misleading.
  4. Scan the overall export once a day - If I only ever looked at it in outline form, I wouldn't get a sense of how the flow is for another person.  You need to constantly look at the paper as if you are a new reader.
  5. Don't sweat the figures/images - These take a bit of time to look right and make the paper too pretty.  Again, misleading confidence.  Give them placeholders and save them for the end if they aren't already created.  Then you'll be pleasantly surprised when it looks better.