Monday, March 29, 2010

Whack Gestures: Inexact and Inatentive Interaction with Mobile Devices

Scott E. Hudson

Chriss Harrison

Beverly L. Harrison

Human-Computer Interaction Institute Carnegie Mellon University, Pittsburgh, PA 15213

Anthony LaMarca

Intel Labs Seattle, WA 98105



Comments:


Manoj's Blog

Frank's Blog


Summary:


Whack Gestures is an imprecise inattentive interaction technique which enables the user to interact with devices without the use of detained vision or fine motor skill. An application of this methodology would be interacting with a mobile device without getting it out with crude eye movements or a grasping action. The utility of this approach is that ongoing activities do not have to be interrupted. The purpose of the article is to present such a system and demonstrate its feasibility.

Advances in mobile device technology have broadened the application of mobile computing devices. New interaction methods are required to maximize their effectiveness, and avoid frustrations, which are typified by cell phones ringing at inappropriate times. In situations where interaction with a mobile device creates an interruption it is important that interaction can be carried out with minimal attention being drawn away from the primary task. The current study examines the use of a set of crude communication gestures, some of which do not require the device to be pulled out.

A small vocabulary of gestures to deal with typical events such as silencing a ringing cell phone, responding to a message with yes, no, forward, etc was developed. The gestures were recorded with an accelerometer and consisted of hitting or whacking the mobile device with either the heel or the palm of the hand. Whacks were chosen as they can be organized highly accurately, and can be recorded by accelerometers which are increasingly present in mobile devices.

Whacks by themselves are hard to distinguish from bumps, so to minimize false positives a pair of whacks was used to frame a gesture as follows . In this way by requiring the whacks to be of similar magnitude and temporally close to the gesture a high signal to noise ratio could be achieved.


The three gestures used for evaluation:

whack-whack An empty gesture.

whack-whack-whack Whack as a signal gesture.

whack-wiggle-whack Shaking as a signal gesture.


To implement and test this technique a mobile sensor platform containing a variety of resources including a 3D accelerometer was used. Data was samples at 256Hz and stored on flash memory for post processing. A recognizer was developed to differentiate above background acceleration changes along the approximate vector of impact. Whacks are differentiated by subtracting an exponential decay average from the raw data. The whacks are then characterized by short sharp pulses. Once a whack has been detected the next 300ms is examined. A framing whack must occur within three seconds and with a magnitude of plus or minus 33%. If more than one are detected the final one is treated as the framing closure. The signal portion between the frames is passed to a secondary recognizer. Average energy is used to identify whack wiggle whack, and peak detection with a min separation of 200ms is used for whack whack whack.



Evaluation:


11 volunteers subjects were used.

Given the MSP and shown how to attach it and instructed to go about their normal activities.

In this way 22hours of baseline data was collected.


Subjects were then given a brief instruction of the three gestures, and were asked to perform them three times each. This data was to be used to train the recognizer.

6 of the subjects data was withheld as a validation set.

The ability to avoid false positives was the primary concern.

Detecting the framing whacks was first determined.



Data:


1 false positive / 12hr of use under normal activity from the validation set.

No false positives from the training data?

100% correct classification of the original signals performed artificially by the recognizer.

97% correct classification overall by the recognizer.



Discussion:


The need for gesture based communication with devices which does not require a large amount of attention is a most interesting idea. The choice of hitting the device is sound. However, I am not sure how one would wiggle the device without pulling it out.

A comparison between alternative gestures is mentioned regarding video and survey studies but no information is given. Only that striking or whacking was rated most positively amongst the six alternatives. It would have been interesting to have presented the details of the alternatives as well as the study.

Finally, no information was given on how the evaluation was conducted to collect data during the two hour period:

What sort of stimuli were the subjects to respond to.

How many were presented.

Were the subjects tested under similar conditions.

what were the qualitative reactions to the system.

The applications of a gesture based communication system with minimal attentional demand is of great interest with many applications.

User-Defined Gestures for Surface Computing

Jacob O. Wobbrock

The InformationSchool DUB Group, University of Washington, Seattle, WA 98195 USA

wobbrock@u.washington.edu

Merideth Ringel Morris,

Andrew D. Wilson

Microsoft Research, One Microsoft Way, Richmond, WA 98052 USA

{merrie, awilson}@microsoft.com



Comments:


Frank's Blog



Summary:


The article presents a gesture based surface computing system which is reflective of user behavior. This system is based on eliciting gestures and then asking the user to perform its effect. The study revealed that desktop idioms strongly influence user mental model, with some commands eliciting little or no gesture agreement. A complete user defined gesture set is presented. The results will hopefully help create better gesture sets.

The exploration of interactive surface tops has revealed a preference for multitouch over the traditional mouse based input. Typically, surface gestures are predefined by the system designers. In contrast, the proposed system allows the user to express their own interpretation. Users rather than principals are used to determine which gestures are chosen.


Eliciting Input from Users:

Participatory design is not a new concept. It has been used successfully in the development of many naturalistic gestures. Other examples include the choice of speech commands from listening to verbal exchanges during similar collaborative tasks.


Developing a User-Defined Gesture Set:

Each of the participants were given a voice command and then saw the effect of a gesture, a block moving across the screen, and was then invited to perform a gesture which would cause the effect. A think aloud protocol was used in addition to videotaping. A wizard of OZ approach was employed with particular attention given to the think-aloud data. Gestures with the highest common use would ultimately be assigned to the particular task.


Procedure:

Each of the 20 participants was presented with 27 random referents. A one handed followed by a two handed gesture was made. The participant was finally asked to evaluate the effectiveness and ease of the gesture before moving on the next one. The gestures were classified according to form, nature, binding, and flow. Within each of these categories there were further subdivisions.

Once the data had been collected an agreement score was calculated to reflect the degree of consensus.



Discussion:

Of the three experts directing the project, their individually generated gesture sets only covered 43.5% of the gesture set generated by the 20 participants in fact the combination of all three individually generated sets still only managed to cover 60.9%. A priori it was by no means clear that the participants gestures would generate a coherent set. Additionally, a combination of widgets and gestures would be effective for the efferents where imaginary ones were used.


Evaluation:

One hand is preferred over two.

Surprising consistency in gesture selection.

Number of fingers are used is not important.

The greater the complexity of the referent the greater the planning time of the gesture.


Data:

1080 gestures

20 participants

22 out of 27 referents were assigned to gestures.

2 referents were combined.

4 were not assigned as they compromise other more primitive gestures or relied on imaginary widgets.

The resulting gesture set covers 57% of all proposed gestures.


Conceptual complexity and therefore planning time inversely correlated with goodness.

Gesture articulation time did not affect goodness.



Discussion:

A very powerful approach to uncovering a more naturalistic gesture set.

The conception and testing was particularly impressive.

However, I am surprised that the investigators were not able to come up with a more representative gesture set as the gestures described seem very familiar.


Wednesday, March 24, 2010

Device Agnostic 3D Gesture Recognition using Hidden Markov Models

Anthony Whitehead,

Kaitlyn Fox,

School of InformationTechnology

Carleton University

1125 Colonel By Drive,

Ottawa, ontario, Canada


Comments:


Manoj’s Blog

Franck’s Blog



Summary:


Introduction:

This paper seeks to identify the necessary elements to successfully use Hidden Markov Models for 3D gesture recognition regardless of sensor devices being used. A variety of input sensors in place of a single one results in increasing the alphabet of the model. However, the study has shown that an alphabet larger than 27 becomes computationally too expensive to allow real time interactivity.


Gestures in Training:

The training set was generated from several users wearing an accelerometer on their wrists. They all performed seven gestures.


The Number of HMM States:

A balance between false negatives and false positive needs to be obtained for the model to be viable. 27 states yielded the best performance. The smaller number of the states had a benefit of increased computational performance.


Number of Samples in a Training Set:

250 samples are sufficient as a training set for a system with 27 states.


Culling Training Data:

During training there were some inconsistencies caused by individual interaction with the hardware. To address this the longest and shortest data elements were discarded. A 1.5 standard deviation rule was used to decide which sequences to discard.


Left vs. Right Hand Training:

Gestures performed by the left and right hands were compared. Gestures performed by the left hand were not recognized, this was attributed to a possible difference in the angle of tilt of the accelerometer. There was an almost a 50% decline in performance from the right to the left hand. Interestingly this was not the case with unidirectional gestures.


Results and Conclusion:

Overall:

91.6% correct recognition for gestures in the training set.

86.4% correct recognition for gestures not in the training set.


Discussion:

The simplicity and adaptability of a wide selection of input sensors was particularly appealing. The comparison between right and left handedness is interesting. There is no mention of the difference in dexterity differences demonstrated by most subjects. A more representative evaluation may have been to test the performance of right handed trained models on left hand gestures performed by left handed subjects.


Wiizards: 3D Gesture Recognition for Game Play Input

Louis Kratz,

Frank J. Lee,

Dept. of Computer Science

Matthew Smith,

Digital Media Labs

Frank J. Lee,

Drexel Universoty

3141 Chesnut Street

Philadelphia, PA 19104


Comments:


Manoj’s Blog

Franck’s Blog



Summary:


Introduction:

This paper explores the use of gestures for game play input using a 3D accelerometer as an input device. Two types of gestures are defined: Static, and dynamic. Bayesian methods will be used to classify the accelerometer data directly in place of path tracking. This is seen as more efficient than path shape matching. The Hidden Markov Models approach used in this paper is additionally well suited to the noisy sensor data provided by the accelerometer.


Implementation:

Wiizards in a two player zero sum game where opponents cast spells at each other. As each player performs a gesture it is placed on a queue so that a series of gestures leading to a cumulation is possible. Several forms of visual feedback are provided. A visual representation of the spell queue which serves as a reminder. A indication of how long until a spell becomes available.

A nintendo Wii controller and a gesture recognition system are used in the implementation. Gestures are represented in the form a collection of vectors. A model is created for each gesture and then a probabilistic matching approach is used as calculated by the Viterby Algorithm.


Results:

7 Users were tested.

Each gesture was performed over 40 times.

Multiple states were used.

90% correct recognition was obtained with only 10 states.

93% with 15 states.

250 gestures per second were possible with a 2.66Ghz PC.

The time of training significantly increases with the number of states.


Conclusions:

The implementation and structure of the game allowed it to adapt with the users varying levels. As the individual’s skill increased so could the complication of the gestures.


Discussion:

The system can run in real time but is encumbered by a 10 second machine training session, which is seen as a limiting factor.

The adaptive nature of the system is very appealing as it should maintain interest as the users skill improves.


Gameplay Issues in the Design of 3D Gestures for Video Games

John Payne, Paul Keir, Jocelyn Elgoyhen, Mairghread McLundle, MartinNaef, Martyn Horner, Paul Anderson

Digital Design Studio

Galsgow School of Art

Glasgow, UK G41 5BW


Comments:


Manoj’s Blog

Franck’s Blog



Summary:


Introduction:

This article identifies points to be considered in the development of 3D gestures as a means of interacting with video games. Four game scenarios using different gesture characteristics were used to identify gameplay issues that have an impact on the design of 3D gestures. The use of gestures offers a natural and intuitive alternative to cumbersome controller mechanisms. However, in spite of the benefits the implementation of a gesture based system is not without difficulty.

To address this 3motion which is a development kit allows 3D gestures to be easily defined and implemented was developed. This is the platform on which user experiences with intuitive gestures in gaming applications will be measured.

The implementation of 3D gestures brings to light several problems:

How to present 3D gesture feedback.

User performance differences.

What are familiar semiotics for 3D gestures.

How to control menu screens in 3D.

The lack of anything to hold.


Testing Rationale:

In the initial phases of the development of 3motion it was subjected to a wide range of users from different backgrounds enabling informal observations of user behavior to be made. From this, several contrasting gesture types were explored:

Direct mapping to actions.

Symbolic use.

Tight , highly controlled precise movements.

Broad gestures.

Speed/repetition.

Accuracy.


System Description & Hardware:

The system consisted of a video camera, laptop with the demonstration software, and the 3motion hardware.

A pretest interview was conducted where the users were introduced to the controls and given a brief description of the games. They were then allowed to experiment with each game under supervision while their comments and performance recorded. Once they were comfortable, they were left to play several rounds with each of the games. During the trials the users were asked to think aloud describing their experience. At the end a post test session was conducted to identify the positive and negative elements of the experience.


User feedback:

Attention was given to the users ability to calibrate their movements to obtain the best results.

There was a preference for simple gestures, with some users commenting that they were unsure what movements exactly constituted particular gestures. The users ability to learn and differentiate gestures was less than expected. In the wizard game the use of horizontal and vertical mirroring was also frequently confused.

Users with a lot of gaming experience found the use of gestures lacking in precision. In contrast, users without much gaming experience were impressed with the intuitive nature of the 3motion setup.


Conclusions:

Although the tests were somewhat preliminary, the intuitive nature of gestures was preferred to button based interfaces. The role of initial instruction was significant in the users overall enjoyment. In particular, user feedback in gesture games is closely linked to the type of gameplay being designed. Future work will focus on this last point.


Discussion:

There is a general lack of descriptive detail in this article. In particular the user study does not present any quantitative data. It is however interesting that individuals who have not had a lot of gaming experience seemed more openminded to the new gesture based approach. Whereas confirmed gamers seemed more set in their ways comparing gestures with familiar interfaces.


Wednesday, March 10, 2010

An Architecture for Gesture-Based Control of Mobile Robots

Soshi Iba, Michael Vande Weghe Christiaan J.J. Paredis, and Pradeep K. Khosla

The Robotics Institute

The Institute for Complex Engineered Systems

Carnegie Mellon University

Pittsbergh, PA 15213


Comments:


Drew’s Blog

Manoj’s Blog



Summary:


Introduction:

This article presents a gesture based method for controlling mobile robots. Hidden Markov Models are used to spot and recognize six gestures reliably, with the use of a wait state to differentiate non-gestures. The gestures are mapped onto global and local modes of operational control. The two modes refer to the frame of reference from which the command is based.


The current state of the art is based on iconic programming. This is a method by which information can be communicated through human demonstration. In this way the programming burden is transferred from robot experts to task experts. The interface has to be intuitive and be able to cope with potentially vague information. Examples of data input are vision, data glove, and tactile sensing. The challenge is to interpret rather than mimic the input data, interpreting the intent as it were.


System Description & Hardware:

Data is collected through a combination CyberGlove and Polhemus 6DOF position sensor. The mobile robot is tracked with a geolocation system that measures both position and orientation.


Hardware:

P5 Data glove was chosen due to its economic cost and integrates position tracking. The finger flexion data was fairly reliable, unlike the position data which was needed additional processing in order to be usable.

Onboard sensors:

8 sonar sensors

7 IR obstacle detectors

A black & white camera with radio transmitter

Stereo mocrophones

position encoders on the tread drive mechanisms

on-board PC104-based i486 running linux


The onboard system is responsible for motion and camera control

A CyberRAVE client server is used to manage the gesture spotter/interpreter and geoposition system.


Gesture Recognition:

The HHM was chosen to take advantage of the temporal component of the gestures. Data is preprocessed in two stages. First, the 18-dimensional joint vector is reduced to a 10-dimensional feature vector. This is augmented with its first derivative to produce a 20-dimensional column vector. The second stage reduces this a 1-dimensional code word. through vector quantization. The codebook is trained offline with a vocabulary of 32 codewords.


Six Gestures:

OPENING: closed fist to flat hand.

OPENED: flat hand.

CLOSING: flat hand to closed fist.

POINTING: moving from a flat hand to a pointing index finger.

WAVING LEFT: fingers extended waiving to the left.

WAIVING RIGHT: fingers extended waiving tot he right.


Local Robot Control:

Closing decelerates the robot.

Opening, Opened maintains the current speed.

Pointing, accelerates the robot.

Waiving Left/RIGHT increases the rotational velocity in the appropriate direction.


Global Robot Control:

Closing decelerates and eventually stops the robot.

Opening, Opened maintains the current speed.

Pointing, “go there”

Waiving Left/RIGHT increases the rotational velocity in the appropriate direction.


Discussion:

The use of hand gestures to control a mobile robot takes advantage of the rich and natural vocabulary of hand gestures. A particularly interesting feature of the implementation is the use of a wait state to segment gestures and discriminate between gestures and non gestures. This produced a reduction in false positive identifications by an order of magnitude in comparison to traditional HHM recognizers. In future research this will be extended to multi robot systems.


Human-Centered Interaction with Documents

Andreas Dengel, Stefan Agne, Bertin Klein

Knowledge Management Lab, DFKI GmbH Kaiserslautern, Germany

{dengel,agne,klein}@dfki.de

Achim Ebert, Matthias Deller

Intelegent Visualization Lab, DFKI GmbH Kaiserslautern, Germany

{ebert,deller}@dfki.de


Comments:


Manoj’s Blog

Franck’s Blog



Summary:


Introduction:

This article presents a new user interface for organizing and visualizing documents in 3D. In the last decade documents have ceased to be tangible objects. With this transition some of their defining qualities embodied in their physical layout have also been lost.

A collection of documents can be viewed as an information space. A virtual environment that resembles a real space can be more readily interpreted by users without prior computer knowledge. Qualities such as size, and relation to other documents to name a few can be visually represented. Additionally, the user experience can also be more fun.

Documents in this implementation are presented as if they are arranged in a book case. A search can be invoked by a gesture. Preselected documents, appear with greater detail, occupy a higher zoom domain. Several viewing modes are available. Pulsation, to draw attention to a document, and color in the form of yellowing to illustrate age are used. The arrangement of documents also reflects their relevance to each other.


Interaction:

The most natural way to manipulate objects is with ones hands. Hands are used to grab, and move, or manipulate objects in other ways. In the interest of minimizing the cognitive load on the user a gesture recognition engine that recognizes natural hand gestures is employed. It has to be compatible with multiple devices, operated in multiple environments.


Hardware:

P5 Data glove was chosen due to its economic cost and integrates position tracking. The finger flexion data was fairly reliable, unlike the position data which was needed additional processing in order to be usable.


Posture, Gesture Recognition & Learning:

Postures are learned by performing them and giving them a name. The system is intended to provide realtime functionality on an average PC without taking up too much processing power.


Recognition Process:

Recognition is done in a two step process with data acquisition and gesture management.


Gesture Recognition:

Gestures are seen as a series of successive postures, in this way dynamic gestures can be perceived. Posture change events are used to segment the end of one gesture and the beginning of another.


Implementation & Results:

A SeeReal C-I 3D display was used to present a stereo image to the user.

Semiotic gestures are used to communicate information and ergotic gestures are used are used to manipulate ones surroundings. A calendar and pin board were used to allow users to experiment with the interface. They were able to manipulate objects and replace existing gestures with new ones. Several users, from a range of backgrounds, were tested in moving and browsing through documents. After a short adaptation period to the glove they were able to successfully use naturalistic gesture. However, leafing through lengthy documents proved difficult.


Discussion:

The loss of non verbal information through the transition to non tangible electronic media is frequently underestimated. The layout of documents was an element that I had not thought of, The attempt to provide a more tangible environment all be it a virtual one in which to manipulate and store documents is logically very sound.

The user study was rather brief and not commensurate with the effort put into the implementation.