Pages

Tuesday, May 29, 2012

T(ether)

Artivle from http://kiwi.media.mit.edu/tether/

Authors: Matthew Blackshaw (@mblackshaw), Dávid Lakatos (@dogichow), Hiroshi Ishii, Ken Perlin

For more information please contact tether@media.mit.edu.

T(ether) is a novel spatially aware display that supports intuitive interaction with volumetric data. The display acts as a window affording users a perspective view of three- dimensional data through tracking of head position and orientation. T(ether) creates a 1:1 mapping between real and virtual coordinate space allowing immersive exploration of the joint domain. Our system creates a shared workspace in which co-located or remote users can collaborate in both the real and virtual worlds. The system allows input through capacitive touch on the display and a motion-tracked glove. When placed behind the display, the user’s hand extends into the virtual world, enabling the user to interact with objects directly.

Above: The environment can be spatially annotated using a tablet's touch screen.

Above: T(ether) is collaborative. Multiple people can edit the same virtual environment.

Technical

T(ether) uses Vicon motion capture cameras to track the position and orientation of tablets, user heads and hands. Server-side synchronization was coded using NodeJS and tablet-side code uses Cinder. The synchronization server forwards tag location to each of the tablets over wifi, which in turn renders the scene. Touch events on each tablet are broadcasted to all other tablets using the synchronization server.

Press Kit

via congeo

Monday, May 28, 2012

ΠΡΟΣΚΛΗΣΗ ΕΚΔΗΛΩΣΗΣ ΕΝΔΙΑΦΕΡΟΝΤΟΣ ΓΙΑ ΕΚΠΟΝΗΣΗ ΔΙΔΑΚΤΟΡΙΚΗΣ ΔΙΑΤΡΙΒΗΣ

Στα πλαίσια του ερευνητικού προγράμματος ΘΑΛΗΣ, η Ομάδα Επεξεργασίας Εικόνας και Πολυμέσων (http://ipml.ee.duth.gr/) του Τμήματος Ηλεκτρολόγων Μηχανικών και Μηχανικών Υπολογιστών του Δημοκριτείου Πανεπιστημίου Θράκης, θα διερευνήσει νέες μεθοδολογίες για την αποτελεσματική ανάκτηση τριδιάστατων (3Δ) αντικειμένων (στατικών και 3Δ βίντεο). Χαρακτηριστικά δείγματα σχετικών ερευνητικών αποτελεσμάτων της παραπάνω ομάδας βρίσκονται στο σύνδεσμο http://utopia.duth.gr/~ipratika/3dor/

Σε αυτό το πρόγραμμα υπάρχει μία (1) χρηματοδοτούμενη θέση για εκπόνηση διδακτορικής διατριβής.

Η έναρξη εκπόνησης της διατριβής αφορά στο ακαδημαϊκό έτος 2012-13. Ενθαρρύνονται να εκδηλώσουν ενδιαφέρον υποψήφιοι οι οποίοι διαθέτουν τα απαιτούμενα προσόντα και έχουν ολοκληρώσει τις βασικές ή μεταπτυχιακές σπουδές τους το αργότερο μέχρι τον ερχόμενο Σεπτέμβρη.

Είναι επιθυμητό οι υποψήφιοι να έχουν υπόβαθρο σε τουλάχιστον κάποιο από τα θέματα που αφορούν σε Γραφικά, Επεξεργασία Εικόνας, Οραση Υπολογιστή και Αναγνώριση Προτύπων καθώς και να έχουν εμπειρία στον προγραμματισμό είτε σε περιβάλλον Matlab είτε σε οποιοδήποτε άλλο περιβάλλον (πχ. Visual Studio, κτλ.) με γλώσσες προγραμματισμού όπως C/C++/C#.

Οι υποψήφιοι που θα επιλεγούν θα δουλέψουν σε ένα δυναμικό περιβάλλον και θα συνεργασθούν με ερευνητές ιδρυμάτων του εσωτερικού (Εθνικό Καποδιστριακό Πανεπιστήμιο Αθηνών – Τμήμα Πληροφορικής, Ε.Κ. «ΑΘΗΝΑ», ΕΚΕΦΕ «ΔΗΜΟΚΡΙΤΟΣ» - Ινστιτούτο Πληροφορικής και Τηλεπικοινωνιών) καθώς και με ερευνητές ιδρυμάτων του εξωτερικού (University of Houston – USA, Vrije Universiteit Brussel – Belgium, Utrecht University - Netherlands, Consiglio Nazionale delle Ricerche - Italy).

Οι ενδιαφερόμενοι παρακαλούνται να επικοινωνήσουν άμεσα με :

Επ. Καθηγητή Ιωάννη Πρατικάκη (http://utopia.duth.gr/~ipratika/)

Τμήμα Ηλεκτρολόγων Μηχανικών και Μηχανικών Υπολογιστών

Δημοκρίτειο Πανεπιστήμιο Θράκης

Γραφείο 1.15, Κτίριο Β’

Πανεπιστημιούπολη, Κιμμέρια, Ξάνθη

VIR tutorial by Oge Marques featuring Lire @ SIGIR 2012

Article from:http://www.semanticmetadata.net/2012/05/16/vir-tutorial-by-oge-marques-featuring-lire-sigir-2012/

Dr. Oge MarquesDr. Oge Marques, author of the book Practical Image and Video Processing Using MATLAB is giving a tutorial on Java based visual information retrieval at SIGIR 2012. Oge Marques is Associate Professor in the Department of Computer & Electrical Engineering and Computer Science at Florida Atlantic University. He has been teaching and doing research on image and video processing for more than twenty years, in seven different countries.

In his tutorial, he presents an overview of visual information retrieval (VIR) concepts, techniques, algorithms, and applications. Several topics are supported by examples written in Java, using Lucene (an open-source Java-based indexing and search implementation) and LIRE (Lucene Image REtrieval), an open-source Java-based library for content-based image retrieval (CBIR) .

Read more & register on the SIGIR 2012 web page (as soon as it is updated).

Friday, May 25, 2012

Microsoft Released Face Tracking SDK in Kinect for Windows

The Microsoft's revolutionary hardware, the Microsoft Kinect, is getting a new piece of brain. Microsoft just released Face Tracking SDK in Kinect For Windows. It can be used for 3D face tracking. It supports most facial types and works in real-time.

You can use the Face Tracking SDK in your program if you install Kinect for Windows Developer Toolkit 1.5. You need to have Kinect camera attached to your PC. The face tracking engine tracks at the speed of 4-8 ms per frame depending on how powerful your PC is.

Take a look at the following demo which shows its facial tracking capabilities, range of supported motions, real-time tracking speed and robustness to occlusions.

Here are several things that will affect tracking accuracy, provided by Nikolai Smolynskiy.

1) Light – a face should be well lit without too many harsh shadows on it. Bright backlight or sidelight may make tracking worse.

2) Distance to the Kinect camera – the closer you are to the camera the better it will track. The tracking quality is best when you are closer than 1.5 meters (4.9 feet) to the camera. At closer range Kinect’s depth data is more precise and so the face tracking engine can compute face 3D points more accurately.

3) Occlusions – if you have thick glasses or Lincoln like beard, you may have issues with the face tracking. This is still an open area for improvement. Face color is NOT an issue.

The Face Tracking SDK is based on the Active Apperance Model (See Wikipedia explanation for AAM). It also utilizes Kinect’s depth data, so it can track faces/heads in 3D. More technical publications You can be found in the following publications:

  • Iain Matthews and Simon Baker, "Active Appearance Models Revisited," International Journal of Computer Vision, Vol. 60, No. 2, November, 2004, pp. 135 - 164. pdf
  • Zhou, M., Liang, L., J. S. & Wang, Y. "AAM based face tracking with temporal matching and face segmentation,"IEEE CVPR, 2010, 701-708. pdf

To download the SDK visit here.

Wednesday, May 23, 2012

Leap

Leap represents an entirely new way to interact with your computers. It’s more accurate than a mouse, as reliable as a keyboard and more sensitive than a touchscreen.  For the first time, you can control a computer in three dimensions with your natural hand and finger movements.

This isn’t a game system that roughly maps your hand movements.  The Leap technology is 200 times more accurate than anything else on the market — at any price point. Just about the size of a flash drive, the Leap can distinguish your individual fingers and track your movements down to a 1/100th of a millimeter.

This is like day one of the mouse.  Except, no one needs an instruction manual for their hands

https://live.leapmotion.com/about/

[via]

Tuesday, May 22, 2012

[New Paper] Dynamic two-stage image retrieval from large multimedia databases

Avi Arampatzis | Konstantinos Zagoris | Savvas A. Chatzichristofis

Information Processing & Management

Content-based image retrieval (CBIR) with global features is notoriously noisy, especially for image queries with low percentages of relevant images in a collection. Moreover, CBIR typically ranks the whole collection, which is inefficient for large databases. We experiment with a method for image retrieval from multimedia databases, which improves both the effectiveness and efficiency of traditional CBIR by exploring secondary media. We perform retrieval in a two-stage fashion: first rank by a secondary medium, and then perform CBIR only on the top-K items. Thus, effectiveness is improved by performing CBIR on a ‘better’ subset. Using a relatively ‘cheap’ first stage, efficiency is also improved via the fewer CBIR operations performed.

Full-size image

Our main novelty is that K is dynamic, i.e. estimated per query to optimize a predefined effectiveness measure. We show that our dynamic two-stage method can be significantly more effective and robust than similar setups with static thresholds previously proposed. In additional experiments using local feature derivatives in the visual stage instead of global, such as the emerging visual codebook approach, we find that two-stage does not work very well. We attribute the weaker performance of the visual codebook to the enhanced visual diversity produced by the textual stage which diminishes codebook’s advantage over global features. Furthermore, we compare dynamic two-stage retrieval to traditional score-based fusion of results retrieved visually and textually. We find that fusion is also significantly more effective than single-medium baselines. Although, there is no clear winner between two-stage and fusion, the methods exhibit different robustness features; nevertheless, two-stage retrieval provides efficiency benefits over fusion.

http://www.sciencedirect.com/science/article/pii/S0306457312000489