Tuesday, March 1, 2011

a course about mobile computer vision

Just come across a course about mobile computer vision, taught by Silvio Savarese in U. Michigan. If you are an active computer vision researcher, you know who he is.

Advanced Topics in Mobile Computer Vision

There are lots of useful resource on their course site, include a brief introduction for Android and development therein. There are also several project reports, which can be used a reference for someone interested in topic.

Wednesday, February 23, 2011

DIY telepresence robot

Since I may need to work far away from home (from 200+miles, e.g., Minneapolis, MN, to 2000 miles, e.g., Los Angeles, CA) in the near future, I am considering to build a telepresence robot to keep in touch with my wife and kids at our Iowa home.

I first got the idea of telepresence robot from an issue of IEEE Spectrum a few months ago. Here is the link, A DIY Telepresence Robot, and a website mentioned in this IEEE article, Sparky Jr., dedicated to DIY, open-source mobile telepresence. Today, I found another post about this idea,titled "Google Engineer Builds an Affordable DIY Telepresence Robot To Keep In Touch With Remote Fiancee", and the Google engineer's Jonny Lee's website that explains his approach is Procrastingeering. The total cost of Lee's telepresence robot is about $500. If you don't make you hand dirty, iRobots is going to release an app platform, AVA, so you can forget the DIY work on the hardware and focus on coding on this new tech gadget now.

I think the cost can be further reduced if I replace the $250 netbook with a Chinese Shan-Zhai netbook or APad (Pad running Android). Nowadays, there are lots of APad under 1000 RMB (~150 US dollar) with touch screen. For example, 7寸Android 安卓系统的MID APAD 平板电脑 1000元以下, 1000以下平板电脑排行. I believe my kids will love the touch screen, and in turn love their daddy! In addition, it is very cool to connect a Kinect on the remote site, so I can control the robot at remote home using my hands and pose.

This idea is great, at least for me. I am planning to do it right after submitting ICCV paper.

The following are a list of useful websites for DIY robot I discovered after I first post this blog:
http://www.robotshop.com/store
http://www.parallax.com/
http://www.servocity.com/
http://www.surveyor.com/SRV_info.html

Friday, February 4, 2011

a new book about computer vision

 Computer Vision: Algorithms and Applications by Richard Szeliski at Microsoft Research. The author provides free PDF version to download. I should read it when possible.

Tuesday, January 25, 2011

(ZT) How technology will change our mind and brain?

An article on my favorite Chinese online community "cchere", discussing how technology changes our minds, our ways to use our brain or even our brain itself. The article is written in Chinese. 


科技改变大脑 (上)
科技改变大脑(下)

My abstract is as follows:
  • Part 1: how our brain/mind can be changed by new technologies
    • It is found Nietzsche changed his writing style after using typewritter
    • Google enable us to access to vast amount information so easily that it is unnecessary for us to memorize many things. As a result, our skills of memorization is impaired.
    • After more and more reading on the Web, we lose our patient to read a long article from top to bottom: we keep jump from one point to another and all we remember is a pile of small pieces of information rather the whole set of them. Consequently, our capability of reading and understanding degenerates. 
    • Our brain can be change due to our style of using it (Neuroplasticity)
  •  Part 2: how control our world by directly using our brain and how we can control our brain
    • we have already technology to issue some simple command using our brainwave, which is mostly used by disabled person
    • DARPA's project "Silent Talk" on using brainwaves to communicate in battlefield
    • DARPA's project on repairing our brain, which can be used by soldiers in battlefield
    • 2008 NSA has a report "Emerging Cognitive Neuroscience and Related Technologies"
    • Will the scenarios in "Matrix" become true? Very likely.

read your mind on iphone?

Recently, PLX devices release a new product called Xwave, which outputs eight EEG signals about your brainwave: Delta, Theta, Low Alpha, High Alpha, Low Beta, High Beta, Low Gamma and Mid Gamma and two easy-to-interpret values: attention and meditation that are derived from the Beta and Alpha wave respectively. The main advantage of the this device is its low price ($99) and easiness to use. In addition, Xwave has a SDK such that iphone developers can implement their ideas quickly. Currently, there has been a few games using this device. One of them is to try to float a ball using your attention strength. More complex and practical games are proposed in its developer guide, including
  • MindWrestle – It’s the same as arm wrestling, however it’s done with the mind.
  • Wormhole Disk - The more you relax or (meditate) the disks float into these wormholes
    which pop up from the bottom of the screen. The more tense you are, the disks just float there and takes longer for the disks to find their way into the wormhole.
  • useTheForce - use your mind's force to throw your enemy or objects
  • Yoga
  • Archery/Shooting
  • MusicMatch - compare your brainwave with your friend's when you are listening to the same song
  • CoupleSync - compare your brainwave with your loved one when you are doing the same activity
  • BrainExercise

By the nature of the EEG signals, we can not use this types of device to do activities that require accurate handling or localization, such as drive a car. However, Xwave does bring us a huge potential on a lots of futuristic games and activities. To some extents, some scenarios in sci-fi will become true. 

Here are a few more interesting applications using Xwave that come to my mind:
  • youLie - use your brainwave to tell whether you are lying
    • your wife or your fiancee will love it!
  • findMyFavorite - find your favorite picture/food/game/product by reading your mind 
    • you can rely on your instinct now
  • Zen - do Zen medication and let Xwave tell you how good you are (similar to Yoga)
    • test your brainwave when you medicate and listen to relaxing music, so that you can find the best music to bring you peace
To me, Xwave is not just an entertaining device. I think it will be also useful in my research on computer vision. Several researchers have started to investigate the mapping between images and brain activity. For example, Fei-Fei Li's group has a project on scene classification using fMRI. Kewei Tu's blog "Mind Reading" also mentioned a few research work on mapping brain activities to words/images. Here are two papers referred in his blog: 
I would like to tap into this research topic when time permits me to do so. 

Wednesday, December 8, 2010

Ph.D. proposal exam passed

I just passed my Ph.D. proposal exam on Monday, Dec. 6th, 2010. The title of the proposal is "Bridging the Semantic Gap: Image and Video Understanding by Integrating Vision and Language". The hypothesis of this thesis is that many vision problems can not be solved solely based on the visual data. Knowledge and reasoning process need to be integrated in the loop, which are provided by language. This idea is embodied by several recent projects, as described in my research web page.


Now I need to work hard to finish the rest work towards my Ph.D. Hopefully, they can be done in next year.

Twenty Questions Game and Object Recognition

Objects can be defined by many features/parts/attributes, each of which can be viewed as a test. The problem of object recognition/detection is then solved by combining outputs of these tests. Instead of performing all possible tests, a smart way is to select a small set of tests without sacrificing the recognition quality/accuracy. The process of selectinvision.ucsd.edu/sites/default/files/Visipedia20q.pdfg the right tests can be formulated as a 20-question game, and the recognition of object is achieved by sequentially asking a question to an Oracle, and analyzing the results returned by the Oracle. The criterion of selecting next question is the information gain brought by the answer of the question. This approach is also called "Active Testing" in "An Active Testing Model for Tracking Roads in Satellite Images", PAMI 1996.

So far, the earliest work using this idea for object recognition is due to Donald Geman of JHU, described in his 1993 technical report "Shape Recognition and Twenty Questions". Each test is a local functional of the image loosely corresponding to configurations (vertex labels) resembling "endings", "junctions", and "turns", or a invariant relations (relational labels) between two vertex labels, i.e., "same class", "same orientation".

The most recent work is "Active Testing for Face Detection and Localization", PAMI 2010, "Visual Recognition with Humans in the Loop", ECCV2010a, and "Indoor Scene Recognition Through Object Detection Using Adaptive Objects Search", ECCV 2010b. In the PAMI 2010 paper, the tests are specific type of image functional (i.e., proportion of edges in particular orientation and scale) within a local region. In the ECCV 2010b paper, the tests are object detectors. In the ECCV 2010a paper, the tests are object attributes while the Oracle is human.

This idea can be extended in many aspects. In the application domain, it can be used in scene and activity recognition; regarding the questions to ask, we can ask many richer questions besides What, e.g., Where, How Many, How Big, etc. We are currently investigating these problems.

Wednesday, December 1, 2010

Microsoft Kinect: the next generation of HCI device?

With the release of Kinect, Microsoft becomes a star in the eyes of researchers of computer vision and HCI. It is really a cool idea to control your computer with your hands and body, without attaching/holding any other devices. It provides us with numerous possibilities. I think there will be boom of games and VR/AR applications using Kinect in the next few years.

Kinect for Xbox 360 review

Open source Kinect driver

Kinect's open-source ambitions

Wednesday, November 17, 2010

Two extreme views on object attributes in the community

I have a paper on attribute-based transfer learning for object categorization in ECCV this year. So I am very curious about the views or attitudes of the community on this topic. During this ECCV, there is a one-day workshop on this topic. At the end of this workshop, there is a panel discussion about this topic among five leading researchers in the computer vision community. It turns out that there are two views on this topic which occupy the two extremes of the spectrum. The followings are summaries of their personal views on object attributes: 
  • Malik doesn’t favorite attributes. He said “vision should not be hijacked by language”
  • Mata doesn’t favorite attributes. He said “my dog can recognize as good as the state-of-the-art computer vision algorithms or even better without language”
  • Hoiem considers attributes a way to go beyond recognition for image understanding, i.e., describing objects and scene
  • Fei-Fei considers attributes as a knowledge
  • Lampert considers attributes as a way to transfer knowledge to the vision system
Overall, there is neither clear definition on attributes nor consensus in the community. It is still a controversial topic. But  it may be a hot research topic in the next a few years. In this ECCV, there are three papers about attributes.
  • Automatic Attribute Discovery and Characterization from Noisy Web Data 
    • Idea:  mining text and image data sampled from the Internet
    • Motivation: product images online are often accompanied texts describing their attributes, such as color, parts, functions, etc.
  • A Discriminative Latent Model of Object Classes and Attributes
  • Attribute-based Transfer Learning for Object Categorization with Zero or One Training Example (my paper)