Pages

Friday, March 29, 2013

How Well Do You Know Tom Hanks? Using a Game to Learn About Face Recognition

Oge Marques, Justyn Snyder and Mathias Lux

Human face recognition abilities vastly outperform computer-vision algorithms working on comparable tasks, especially in the case of poor lighting, bad image quality, or partially hidden faces. In this paper, we describe a novel game with a purpose in which players must guess the name of a celebrity whose face appears blurred. The game combines a successful casual game
paradigm with meaningful applications in both humanand computer-vision science. Preliminary user studies were conducted with 28 users and more than 7,000 game rounds. The results supported and extended preexisting knowledge and hypotheses from controlled scientific experiments, which show that humans are remarkably good at recognizing famous faces, even
with a significant degree of blurring. Our results will be further incorporated into research in human vision as well as machine-learning and computer-vision algorithms for face recognition.

http://faculty.eng.fau.edu/omarques/files/2013/02/Marques_Snyder_Lux_CHI2013.pdf

Wednesday, March 27, 2013

Mobile Visual Search (MVS)

Once more, one year later

An amazing presentation by  (my friend) Oge Marques!!!!

Watch the (full length !!! ) video here:

http://www.foerderverein-technische-fakultaet.at/2012/01/ruckblick-mobile-visual-search-video-slides/

imageimage

Abstract

Mobile Visual Search (MVS) is a fascinating research field with many open challenges and opportunities which have the potential to impact the way we organize, annotate, and retrieve visual data (images and videos) using mobile devices. This talk is structured in four parts:

(i) MVS — opportunities: where I present recent and relevant numbers of the mobile computing market, particularly in the field of photography apps, social networks, and mobile search.

(ii) Basic concepts: where I explain the basic MVS pipeline and discuss the three main MVS scenarios and associated challenges.

(iii) Advanced technical details: where I explain technical aspects of feature extraction, indexing, descriptor matching, and geometric verification, discuss the state of the art in these fields, and comment on open problems and research opportunities.

(iv) Examples and applications: where I show recent and significant examples of academic research (e.g., Stanford Product Search System) and commercial apps (e.g., Google Goggles, oMoby, kooaba) in this field.

Great stuff professor Marques, thanks for sharing

Tuesday, March 26, 2013

Learning to Rank Research using Terrier

Part 1: http://terrierteam.blogspot.co.uk/2013/03/learning-to-rank-research-using-terrier.html

Part 2: http://terrierteam.blogspot.gr/2013/03/learning-to-rank-research-using-terrier_26.html

In recent years, the information retrieval (IR) field has experienced a paradigm shift in the application of machine learning techniques to achieve effective ranking models. A few years ago,we were using hill-climbing optimisation techniques such as simulated annealing to optimise the parameters in weighting models, such as BM25 or PL2, or latterly BM25F or PL2F. Instead, driven first by commercial search engines, IR is increasingly adopting a feature-based approach, where various mini-hypothesis are represented as numerical features, and learning to rank techniques are deployed to decide their importance in the final ranking formulae.

The typical approach for ranking is described in the following figure from our recently presented WSDM 2013 paper:

Phases of a retrieval system deploying learning to rank, taken from Tonellotto et al, WSDM 2013.

In particular, there are typically three phases:

  1. Top K Retrieval, where a number of top-ranked documents are identified, which is known as the sample.
  2. Feature Extraction - various features are calculated for each of the sample documents.
  3. Learned Model Application - the learned model obtained from a learning to rank technique re-ranks the sample documents to better satisfy the user.
The Sample

The set of top K documents selected within the first retrieval phase is called the sample by Liu, even though the selected documents are not iid. Indeed, in selecting the sample, Liu suggested that the top K documents ranked by a simple weighting model such as BM25 is not the best, but is sufficient for effective learning to rank. However, the size of the sample - i.e. the number of documents to be re-ranked by the learned model - is an important parameter: with less documents, the first pass retrieval can be made more efficient by the use of dynamic pruning strategies (e.g. WAND); on the other hand, too few documents may result in insufficient relevant documents being retrieved, and hence effectiveness being degraded.
Our article The Whens and Hows of Learning to Rank in the Information Retrieval Journal studied the sample size parameter for many topic sets and learning to rank techniques - for the mixed information needs on the TREC ClueWeb09 collection, we found that while a sample size of 20 documents was sufficient for effective performance according to ERR@20, larger sample sizes of thousands of documents were needed for effective NDCG@20; for navigational information needs, predominantly larger samples sizes (upto 5000 documents) were needed; Moreover, the particular document representations that used to identify the sample was shown to have an impact on effectiveness - indeed, navigational queries were found to be considerably easier (requiring smaller samples) when anchor text was used, but for informational queries, the opposite was observed. In the article, we examined these issues in detail, across a number of test collections and learning to rank techniques, as well as investigating the role of the evaluation measure and its rank cutoff for listwise techniques - for in depth details and conclusions, see the IR Journal article.

Read More:

Part 1: http://terrierteam.blogspot.co.uk/2013/03/learning-to-rank-research-using-terrier.html

Part 2: http://terrierteam.blogspot.gr/2013/03/learning-to-rank-research-using-terrier_26.html

Monday, March 25, 2013

QUALINET Multimedia Databases

Article from http://multimediacommunication.blogspot.gr/

A key for current and future developments in Quality of Experience resides in a rich and internationally recognized database of content of different sorts, and to share such a database with the scientific community at large. The QUALINET Database platform takes the necessary steps to make them accessible to all researchers:

http://dbq.multimediatech.cz/ (registration is free of charge)

 

 

Currently, the QUALINET database comprises 122 multimedia databases, based on literature/Internet search and input from Qualinet partner laboratories. They're mostly image (~52) or video datasets (~68), with (~57) or w/o subjective quality rating, special content, e.g. 3D (~23), FV, eyetracking (~21), audio, audiovisual (~8), HDR, and other modalities.

The documentation of the QUALINET databases can be found on the corresponding Wiki page. For an overview, please consult the white paper on QUALINET databases (PDF) and please reference it as follows:

Karel Fliegel, Christian Timmerer, (eds.), “WG4 Databases White Paper v1.5: QUALINET Multimedia Database enabling QoE Evaluations and Benchmarking”, Prague/Klagenfurt, Czech Republic/Austria, Version 1.5, March 2013.

Finally, you're welcome to contribute to this effort, simply send an email towg4.qualinet@listes.epfl.ch and briefly describe your dataset [to subscribe, send an e-mail (its content is unimportant) to wg4.qualinet-subscribe@listes.epfl.ch, you will receive information to confirm your subscription, and upon the acceptance of the moderator will be included in the mailing-list].

Additionally, you may consider submitting a dataset paper to QoMEX or MMSys which hosts dataset tracks and accepted dataset paper will be automatically included within the QUALINET database.

Sunday, March 24, 2013

“There’s no bad data, only bad uses of data”

Article from http://www.nytimes.com/2013/03/24/technology/big-data-and-a-renewed-debate-over-privacy.html?smid=tw-share&_r=0

IN the 1960s, mainframe computers posed a significant technological challenge to common notions of privacy. That’s when the federal government started putting tax returns into those giant machines, and consumer credit bureaus began building databases containing the personal financial information of millions of Americans. Many people feared that the new computerized databanks would be put in the service of an intrusive corporate or government Big Brother.

“It really freaked people out,” says Daniel J. Weitzner, a former senior Internet policy official in the Obama administration. “The people who cared about privacy were every bit as worried as we are now.”

Along with fueling privacy concerns, of course, the mainframes helped prompt the growth and innovation that we have come to associate with the computer age. Today, many experts predict that the next wave will be driven by technologies that fly under the banner of Big Data — data including Web pages, browsing habits, sensor signals, smartphone location trails and genomic information, combined with clever software to make sense of it all.

Proponents of this new technology say it is allowing us to see and measure things as never before — much as the microscope allowed scientists to examine the mysteries of life at the cellular level. Big Data, they say, will open the door to making smarter decisions in every field from business and biology to public health and energy conservation.

“This data is a new asset,” says Alex Pentland, a computational social scientist and director of the Human Dynamics Lab at the M.I.T. “You want it to be liquid and to be used.”

But the latest leaps in data collection are raising new concern about infringements on privacy — an issue so crucial that it could trump all others and upset the Big Data bandwagon. Dr. Pentland is a champion of the Big Data vision and believes the future will be a data-driven society. Yet the surveillance possibilities of the technology, he acknowledges, could leave George Orwell in the dust.

The World Economic Forum published a report late last month that offered one path — one that leans heavily on technology to protect privacy. The report grew out of a series of workshops on privacy held over the last year, sponsored by the forum and attended by government officials and privacy advocates, as well as business executives. The corporate members, more than others, shaped the final document.

The report, “Unlocking the Value of Personal Data: From Collection to Usage,” recommends a major shift in the focus of regulation toward restricting the use of data. Curbs on the use of personal data, combined with new technological options, can give individuals control of their own information, according to the report, while permitting important data assets to flow relatively freely.

“There’s no bad data, only bad uses of data,” says Craig Mundie, a senior adviser at Microsoft, who worked on the position paper.

The report contains echoes of earlier times. The Fair Credit Reporting Act, passed in 1970, was the main response to the mainframe privacy challenge. The law permitted the collection of personal financial information by the credit bureaus, but restricted its use mainly to three areas: credit, insurance and employment.

The forum report suggests a future in which all collected data would be tagged with software code that included an individual’s preferences for how his or her data is used. All uses of data would have to be registered, and there would be penalties for violators. For example, one violation might be a smartphone application that stored more data than is necessary for a registered service like a smartphone game or a restaurant finder.

The corporate members of the forum say they recognize the need to address privacy concerns if useful data is going to keep flowing. George C. Halvorson, chief executive of Kaiser Permanente, the large health care provider, extols the benefits of its growing database on nine million patients, tracking treatments and outcomes to improve care, especially in managing costly chronic and debilitating conditions like heart disease, diabetes and depression. New smartphone applications, he says, promise further gains — for example, a person with a history of depression whose movement patterns slowed sharply would get a check-in call.

“We’re on the cusp of a golden age of medical science and care delivery,” Mr. Halvorson says. “But a privacy backlash could cripple progress.”

more: http://www.nytimes.com/2013/03/24/technology/big-data-and-a-renewed-debate-over-privacy.html?smid=tw-share&_r=0

Proximity solutions for a new way to interact with mobile devices