Fun: The AVC Word Cloud

Happy 2014! In between celebrating Christmas, hanging with family, and ringing in the New Year I managed to put together a visualization of the words used on avc.com. AVC, written by Fred Wilson, is probably one of the most popular “start up” blogs on the Internet. It covers a wide array of topics from “MBA Mondays”, USV portfolio companies and of course general startup and technology news. Given the range of topics and and that the blog has been active since 2003, it naturally seemed like generating a word cloud would produce interesting results. With the goal of generating word clouds in mind, I set off the day after Christmas.

Checkout the finished product at http://symf.setfive.com/d3_avc_blog_cloud/. I actually decided to use Scala to scrape and process the data, look for a follup post on coming to Scala from PHP.

Taking a quick glance at the clouds, a few things do jump out:

  • "Android" enters the top 100 in 2010 and has remained there since.
  • Amazon is surprisingly absent past 2007
  • Apple hasn't made the top 100 in any year.
  • It's interesting to see when USV portfolio companies like Disqus and Zemanta enter and exit.
  • Bitcoin makes the list for 2013
  • Blackberry, one and done
  • Facebook peaked in 2007 and then steadily declines until it drops out this year
  • Google hits the list for every year
  • Twitter gets in at 2007 and sticks through this year

Boston Tech Startup Spotlight: Recorded Future

Boston is one of the most active places in the US for technology innovation and home to hundreds of exciting young companies with incredible new ideas. In support of the Boston tech startup scene, I have been publishing a series of short blog posts spotlighting some of our most interesting neighbors.

Due to our continued fascination with big data and support for companies playing in the space it seemed only logical to write about Recorded Future for this edition.  These guys are also headquartered in Cambridge, with offices in Göteborg, Sweden and Arlington, VA.

They constantly collect real-time data from web sources such as news, blogs, and public social media and use their technology to analyze trends and identify past, present, and future events. These events are then linked to the people, places, and organizations that matter to their clients, who include Fortune 500 companies and leading government agencies.

Recorded Future’s team of computer scientists, statisticians, linguists, and technical business people offer up an array of software products and services centered around web intelligence. They also provide the Recorded Future API, a web service that allows developers to get in on the action by accessing Recorded Future’s index for large scale analysis of online media flow.

If you’re interested, there’s lots more about their products and services on their website.

Stay tuned for the next startup spotlight.

PrestoDB: Running PrestoDB on Amazon EMR

A weeks ago, Facebook released a new open source project called PrestoDB which they billed as a market improvement over Hive and Hadoop. According to the PrestoDB site, Presto is a real time query engine that supports a SQL like syntax, similar to Hive. However, unlike Hive, Presto doesn’t execute queries using MapReduce jobs but instead uses its own internal distribution mechanism. According to the Presto site and current users, most queries will see an order of magnitude speedup compared to Hive. And the best part? PrestoDB can read metadata from Hive’s metastore and read files off HDFS just like Hive - pretty wild.

Anyway, since I love new toys (who doesn’t!?) I decided to try setting up PrestoDB on Amazon EMR to see how difficult it was and also experience the speedups. Turns out, once you have an Amazon EMR cluster running getting PrestoDB up is almost trivial. Just follow the PrestoDB deploying directions to get yourself situated. Make sure you create *all* the files or you’ll get some necessarily cryptic errors along the way.

The config files I ended up using were:

# etc/config.properties
coordinator=false
datasources=jmx,hive
http-server.http.port=8080
presto-metastore.db.type=h2
presto-metastore.db.filename=/mnt/presto/db
task.max-memory=1GB
discovery-server.enabled=false

# etc/jvm.config
-server
-Xmx16G
-XX:+UseConcMarkSweepGC
-XX:+ExplicitGCInvokesConcurrent
-XX:+CMSClassUnloadingEnabled
-XX:+AggressiveOpts
-XX:+HeapDumpOnOutOfMemoryError
-XX:OnOutOfMemoryError=kill -9 %p
-XX:PermSize=150M
-XX:MaxPermSize=150M
-XX:ReservedCodeCacheSize=150M

# etc/log.properties
com.facebook.presto=DEBUG

# etc/node.properties
node.environment=production
node.id=ffffffff-ffff-ffff-ffff-ffffffffffff
node.data-dir=/mnt/presto/data

# etc/catalog/hive.properties
connector.name=hive-hadoop1
hive.metastore.uri=thrift://10.29.191.137:10004

You’ll need to create the “/mnt/presto” directory and also make it accessible to whatever user you plan to run the daemon under.

The one huge gotcha I ran into was that I couldn’t figure out what port Hive’s Thrift service was running on. For some reason, it’s notably absent from Amazon’s documentation and I couldn’t find the hive-site.xml file on the EMR EC2. Completely randomly, I ran across this manual page from Jaspersoft enumerating which ports different versions of Hive run Thrift on when you use EMR. Turns out, its different per Hive version but 0.11.0 will use 10004.

Once you have everything configured, just follow the docs to start the server and you’ll be ready to query. One thing to note though is that you’ll need to setup PrestoDB manually on the rest of your machines and also enable the discovery service for this to “really” work.

Anyway, happy querying!

Musing: Should everyone learn to code?

Last week, President Obama made headlines by suggesting that every American in school should learn how to code. Predictably, the comment sparked some heated discussion across the web from Fred Wilson’s blog to several threads on Hacker News. Surprisingly, some of the viewpoints were extremely polarized ranging from “its useless, some people will never get it” to “of course!”. Personally, I think everyone should definitely be exposed to some form of programming while they’re in school.

An inescapable reality is that in 2013 computers are a part of everyone’s personal and professional day to day. From non-technical roles in technical fields like account managers or project managers to traditionally non-technical jobs, like teachers, everyone is ultimately interacting with computers on a daily basis. With that in mind, having a basic understanding of how computing abstractions and programming work will benefit everyone. From being able to modify a VBA macro to construct a complex Gmail search query, having a basic understanding of how the pieces fit together certainly can’t hurt.

Looking back at high school, drawing an analogy between studying programming and studying a foreign language isn’t really accurate. A better analogy is really the general experience people have studying math in middle and high school. For people that don’t take a math class in college, that’ll normally be the last time they study math in an academic setting. Although most people forget most of the details they learned, they still retain the overarching fundamentals of how things like algebra and geometry work. Because of this, when people are faced with a basic math problem they generally know what they need to look up in order to solve it. Extending this, if people were introduced to basic programming early on they’d have a sense that there might be an easier way to approach certain tasks. Need to format a list of names in Excel? There might be a function for that.

So how can we make this happen? The good news is there’s already a push to make high quality, programming focused education material available to everyone. There are already dozens of masively online open course projects including Khan Academy, Coursera, and Code Academy providing free, interactive, computer science resource for everyone. The next step is pushing states and school systems to actively adopt CS education for their middle school and high school students. Hopefully it’ll prove and easy and effective step to keeping everyone competitive in an increasingly technology powered workplace.

Hive: Hive in 15 minutes on Amazon EMR

As far as “big data” solutions go, Hive is probably one of the more recognizable names. Hive basically offers the end user an abstraction layer to run “SQL like” queries as MapReduce jobs across data that they have in HDFS. Concretely, say you had several hundred million rows of data and you wanted to count the number of unique IDs Hive would let you do that. One of the issues with Hadoop and by proxy Hive is that it’s notably difficult to setup a cluster to try things out. Tools like Whirr exist to make things easier they’re, a bit rough around the edges and in my experience hit up against “version hell”. One alternative that I’m surprised isn’t more popular is using Amazon’s Elastic Map Reduce to bootstrap a Hadoop cluster to experiment with.

Fire up the cluster

The first thing you’ll need to do is fire up an EMR cluster from the AWS backend. It’s mostly just point and click but the settings I used were:

  • Termination protection? No
  • Logging? Disabled
  • Debugging? Off since no logging
  • Tags - None
  • AMI Version: 2.4.2 (latest)
  • Applications to be installed:
  • Hive 0.11.0.1
  • Pig 0.11.1.1
  • Hardware Configuration:
  • One m1.small for the master
  • Two m1.small for the cores

The “security and access” section is important, you need to select an existing key pair that you have access to so that you can SSH into your master node to use the Hive CLI client.

Then finally, under Steps since you’re not specifying any pre-determined steps make sure you mark “Auto-terminate” as “No” so that the cluster doesn’t terminate immediately after it boots.

Click “Create Cluster” and you’re off to the races.

Pull some data, and load HDFS

Once the cluster launches, you’ll see a dashboard screen with a bunch of information about the cluster including the public DNS address for the “Master”. SSH into this machine using the user “hadoop” and whatever key you launched the cluster with:

ashish@ashish:~$ ssh hadoop@ec2-107-20-21-245.compute-1.amazonaws.com
Linux (none) 3.2.30-49.59.amzn1.x86_64 #1 SMP Wed Oct 3 19:54:33 UTC 2012 x86_64
--------------------------------------------------------------------------------

Welcome to Amazon Elastic MapReduce running Hadoop and Debian/Squeeze.
 
Hadoop is installed in /home/hadoop. Log files are in /mnt/var/log/hadoop. Check
/mnt/var/log/hadoop/steps for diagnosing step failures.

The Hadoop UI can be accessed via the following commands: 

  JobTracker    lynx http://localhost:9100/
  NameNode      lynx http://localhost:9101/
 
--------------------------------------------------------------------------------
Last login: Thu Dec 12 02:53:54 2013 from 50.136.18.114
hadoop@ip-10-29-191-137:~$ 

Once you’re in, you’ll want to grab some data to play with. I pulled down Wikipedia Page View data since it’s just a bunch of gzipped text files which are perfect for Hive. You can pull down a chunk of files using wget, be aware though that the small EC2s don’t have much storage so you’ll need to keep an eye on your disk space.

hadoop@ip-10-29-191-137:~$ wget -r --no-parent --reject "index.html*" http://dumps.wikimedia.org/other/pagecounts-raw/2013/

Once you have some data (grab a few GB), the next step is to push it over to HDFS, Hadoop’s distributed filesystem. As an aside, Amazon EMR is tightly integrated with Amazon S3 so if you already have a dataset in S3 you can copy directly from S3 to HDFS. Anyway, to push your files to HDFS just run:

hadoop@ip-10-29-191-137:~/dumps.wikimedia.org/other/pagecounts-raw/2013$ cd ~/dumps.wikimedia.org/other/pagecounts-raw/2013/
hadoop@ip-10-29-191-137:~/dumps.wikimedia.org/other/pagecounts-raw/2013$ hadoop dfs -mkdir /mnt/pageviews-2013-12
hadoop@ip-10-29-191-137:~/dumps.wikimedia.org/other/pagecounts-raw/2013$ hadoop dfs -put * /mnt/pageviews-2013-12/

Build some tables, query some data!

And finally, it’s time to query some of the pageview data using Hive. The first step is to let Hive know about your data and what format it’s stored in. To do this, you need to create an external table that points to the location of the files that you just pushed to HDFS. Start the Hive client by running “hive” and then do the following:

hadoop@ip-10-29-191-137:~/dumps.wikimedia.org/other/pagecounts-raw/2013$ hive

Logging initialized using configuration in file:/home/hadoop/.versions/hive-0.11.0/conf/hive-log4j.properties
Hive history file=/mnt/var/lib/hive_0110/tmp/history/hive_job_log_hadoop_21902@ip-10-29-191-137.ec2.internal_201312120412_1458262789.txt
hive> CREATE EXTERNAL TABLE page_views (
    >     project STRING, title STRING,
    >     req_count STRING, pg_size STRING
    > ) ROW FORMAT DELIMITED 
    >    FIELDS TERMINATED BY ' '
    >    LINES TERMINATED BY '\n'
    >    STORED AS TEXTFILE
    > LOCATION '/mnt/pageviews-2013-12';
OK
Time taken: 0.313 seconds

Now select some data from your newly created table!

hive> select * from page_views limit 10; 
OK
*.b	Ingl\xC3\x01\x00\x00\x00\x00Z\xBB\xB9\x01\x00\x00\x00\x00;kGS\xF6\xC3aH\x0D\x00\x00	1	325
*	100013-11-30T23:59:57.3	1	325
*	100\x00\x00\x00\x00\x00\x00\x00\xA1\x0D\xA1\x01\x00\x00\x00\x00\x10\x00\x00\x00\x00	1	325
*	\x5Cx-\x22\x5Cx/(\x5Cx/*\x5Cx/(\x5Cx/\x5Cx/0\x82\x01\xB3\x06\x09`\x86H\x01\x86\xFDl\x01\x010\x82\x01\xA40:\x06\x08+\x06\x01\x05\x05\x07\x02\x01\x16.http://www.digicert.com/ss	1	325
*	\x5Cx-\x22\x5Cx/(\x5Cx/*\x5Cx/(\x5Cx/\x5Cx/\x00\x00\x00\x00\xDE~\xC4\xD1\x91\xA5\x09\x00\x0A\x00\x00\x00\x00\x00\x00\x00\x19%~\x01\x00\x00\x00\x00\x0A\x00\x00\x00\x00\x00\x00\x00%%~\x01\x00\x00\x00\x004Qm\x01\x00\x00\x00\x00\x08\xB5'\xAB\x00\x00\x00	1	325
AR	%D9%82%D8%A7%D8%A6%D9%85%D8%A9_%D8%A3%D9%84%D8%B9%D8%A7%D8%A8_%D8%A5%D9%8A_%D8%A2%D9%8A%D9%87_%D9%84%D9%88%D8%B3_%D8%A3%D9%86%D8%AC%D9%84%D9%88%D8%B3	1	44480
De.mw	De	1	10302
De	Customer-Relationship-Management	3	96858
De	Include	3	24679
EN.mw	EN	1	4693
Time taken: 14.766 seconds, Fetched: 10 row(s)

Pretty sweet huh? Now feel free to run any arbitrary query against the data. Note: since we used m1.small EC2s the performance of Hive/Hadoop is going to be pretty abysmal. But hey, give it a shot:

hive> select count(*) AS c, title from page_viewsn group by title order by c desc limit 1000;
Total MapReduce jobs = 2
Launching Job 1 out of 2
Number of reduce tasks not specified. Estimated from input data size: 6
In order to change the average load for a reducer (in bytes):
  set hive.exec.reducers.bytes.per.reducer=<number>
In order to limit the maximum number of reducers:
  set hive.exec.reducers.max=<number>
In order to set a constant number of reducers:
  set mapred.reduce.tasks=<number>

Anyway, don’t forget to tear down the cluster once you’re done. As always, let me know if you run into any issues!