Showing posts with label research. Show all posts
Showing posts with label research. Show all posts

Saturday, 6 June 2015

Gianluca is back! (or: PhD course on Micro-task crowdsourcing)

It is with much pleasure that I welcome Gianluca Demartini back at Udine University, Department of Mathematics and Computer Science, for a PhD course on Micro-task crowdsourcing. Gianluca got his Master's degree under my supervision some years ago (don't ask!) and since then he has obtained several research positions abroad. He is now Lecturer at the University of Sheffield.

The lecture timetable is below, together with a tentative program; everyone is welcome!

Timetable (all lectures are in the "Aula Multimediale / DIMI"):

1. Monday 15/6 10:00 - 12:00
2. Monday 15/6 15:00 - 17:00
3. Tuesday 16/6 10:00 - 12:00
4. Tuesday 16/6 15:00 - 17:00
5. Wednesday 17/6 10:00 - 12:00
6. Wednesday 17/6 15:00 - 17:00
7. Thursday 18/6 10:00 - 12:00

Preliminary/tentative program:

Lecture 1 - Introduction to Crowdsourcing
We will start with an overview of the entire module highlighting its aims and objectives. Then, we will look at fundamental definitions and different types of crowdsourcing incentives. Finally, we will present early examples of crowdsourcing such as reCAPTCHA and the ESP game.

Lecture 2 - Introduction to Micro-task Crowdsourcing Platforms
After defining the key terminology of micro-task crowdsourcing, we will introduce popular crowdsourcing platforms such as Amazon MTurk and CrowdFlower including a demonstration on how to use such systems both as a crowd worker as well as a requester.

Lecture 3 - How to Setup a Crowdsourcing Task
In this lecture we will discuss all the dimensions involved in crowdsourcing task design such as pricing, question design, and quality assurance mechanisms (e.g., honeypots). We will also design and deploy a task during the lecture and see how to collect results back from the crowdsourcing platform.

Lecture 4 - Crowdsourcing Patterns
In this lecture we will define the concept of crowdsourcing pattern (i.e., the combination of multiple crowdsourcing tasks) and present popular example patterns. We will also discuss the concept of crowdsourcing workflows where multiple tasks as well as machine processing steps are combined together.

Lecture 5 - Hybrid Human-machine Systems
In this lecture we will see some advanced example uses of crowdsourcing applied to the database, web, and biomedical domains. We will see how systems that combine both the scalability of machines over large amounts of data as well as the quality of human intelligence can be used to improve the effectiveness of Big Data systems.

Lecture 6 - Crowdsourcing Scalability
In hybrid human-machine systems the latency bottleneck lays on the side of the crowd as human intelligence is naturally slower than machine-based computation. In this lecture we will see recent research results that proposed techniques to improve the latency of crowdsourcing platforms.

Lecture 7 - Open Research Directions in Crowdsourcing
In this lecture we will give an overview on which micro-task crowdsourcing research questions different Computer Science areas focus on including database, information retrieval, semantic web, human-computer interaction, multimedia, and bio-informatics.

Prerequisites:
No prior knowledge is needed. Having at least basic programming skills is a plus.

S.


Sunday, 13 January 2013

Cerchi lavoro?


A breve bandirò un assegno di ricerca annuale per lavorare al progetto Axiometrics: Foundations of Evaluation Metrics in IR. Il progetto è finanziato da un Google Faculty Research Award (http://research.google.com/university/relations/research_awards.html).

Abstract del progetto Axiometrics
Effectiveness evaluation is of paramount importance in the field of Information Retrieval (IR). IR is probably the most evaluation-oriented field in computer science, as witnessed by an evaluation methodology developed in the 60s during the Cranfield experiments and by several evaluation initiatives running today (TREC, CLEF, NTCIR, INEX, FIRE). One crucial aspect of evaluation are evaluation metrics. About 100 IR effectiveness metrics exist, and counting. This project aims at understanding the relationships among them, in terms of both axiomatic properties and statistical relations, for both metric science (understanding of metrics) and engineering (their development). More in detail, we aim at proposing:

  • Axioms: rules that any metric must satisfy. For example, when swapping a relevant document and a non-relevant one in the ranking, by decreasing the rank of the relevant one and increasing the rank of the non relevant one, the metric value should decrease. Axioms might be verified and compared with user intuition using crowdsourcing.
  • Desiderata: desirable properties that a metric should have, based on common sense. For example, an effectiveness value according to one metric should be affected more by a swap in earlier rank positions than a swap in later ranks.
  • Empirical properties: those that emerge from data, i.e., from actual test collections and system comparisons. These include robustness, statistical correlation between metrics, etc.


Axiometrics è una delle linee di ricerca più importanti per l’IR proposte durante il meeting SWIRL 2012 (http://www.cs.rmit.edu.au/swirl12/). L’attività è svolta in collaborazione con:

  • Stefano Mizzaro (Principal Investigator), Dept. of Maths and Computer Science, University of Udine, Italy, mizzaro_pippo@uniud.it.
  • Julio Gonzalo (co-Principal Investigator) ed Enrique Amigó, E.T.S.I. Informática de la UNED,  Madrid, Spain, julio_pippo@lsi.uned.es, enrique_pippo@lsi.uned.es.
  • Evangelos Kanoulas (Google sponsor), Google Zurigo, ekanoulas_pippo@gmail.com.

[ovviamente gli indirizzi email contengono alcuni caratteri in più da rimuovere, a meno che tu non sia uno spammer :)]
Se pensi che ti possa interessare, contattami.

Stefano Mizzaro
www.dimi.uniud.it/~mizzaro

Wednesday, 12 September 2012

Alcune bufale sull'università italiana


Alberto Baccini su ROARS elenca alcune affermazioni sull'università italiana, affermazioni che si sentono spesso e che sono, come dice appunto Baccini, false e/o distorte -- insomma, "bufale":

Il contesto. Ecco ciò-che-tutti-sanno-dell’università-italiana (poco importa che alcuni punti dello scenario siano sostanzialmente falsi e altri ingigantiti e distorti):
1.  “La ricerca scientifica attraversa un periodo di stasi”, perché l’università produce poca ricerca [si legga qui e qui], ed è avviata al declino;
2. “La ricerca scientifica deve servire alla scienza e alle esigenze nazionali. Non deve servire a creare nuove cattedre e nuovi insegnamenti.” Il declino della ricerca italiana è causato dall’autoreferenzialità dei baroni.
3. Il sistema di reclutamento è distorto e corrotto da nepotismo e clientele. Il merito è mortificato.
4. All’università italiana e alla ricerca non mancano le risorse. La ricerca condotta dai baroni è spesso inutile per la società ed autoreferenziale.
5. I baroni hanno stipendi tra i più alti al mondo.
Da questo segue che l’università italiana è irriformabile con gli strumenti legislativi usuali; c’è bisogno di una rivoluzione (lo sostiene per esempio  Andrea Ichino). La politica ha il compito di individuare una élite accademica illuminata e d’avanguardia che possa modificare dall’alto il funzionamento della università e della ricerca italiana. I due snodi fondamentali sono finanziamento e reclutamento. Lo strumento istituzionale è l’ANVUR: un organismo tecnico di nomina ministeriale, lasciato incompiuto dal governo di centro-sinistra, cui vengono attribuiti poteri (oltre a molti altri) su valutazione e criteri per il reclutamento.
S.

Wednesday, 4 July 2012

"No flashy. This is Science." [censored] -- Or: How two jetlagged IR researchers met and had an idea

[Prologue: the beginning of a beautiful friendship]

J. Hi. Can I sit here?
S. Sure. Hi. I don't think we met before. Nice to meet you, I'm S.
J. I'm J.
S. Nice place here, isn't it?
J. Yep! Beach this afternoon?
S. Sure!
J. Ok, let's pretend we do some work first. So, what are you working on?
S. Oh, several things, blablabla [snip] And you?
J. Well, I did blablabla... and blabla... and I also published some papers on formal analysis of clustering metrics.
S. Interesting. I also started something similar for IR metrics years ago, but I never managed to publish it.
J. Really?
S. Yep. I'll show you [opening his laptop]. See, I defined a framework based on measurement theory, then I defined some axioms...
J. That's crazy! I did exactly the same!
S. ... then some desiderata...
J. Exactly the same!
S. ... Yes but you published, I didn't...
J. ... because you tried to do everything, look here, that's crazy!
S. Yes, I'm a bit, sure... and then blablabla...
J. blablabla [very technical details here - ok, ok, I admit I do not remember that!]
S. blablabla [very technical details here - ok, ok, I admit I do not remember that!]
J. Well, why not proposing this as an idea for SWIRL research directions?
S. We might indeed!
J. Ok, beach time now.
S. Beach!

[Chapter 1: split groups doing Science]

M. So, let's have a round of the table and everyone presents his own idea.
[various good ideas...everyone presents just one, accordingly to the rules. But you know that Italians and Rules do not fit well in the same sentence, so...]
S. I do have two ideas. The first is not very exciting, but it's Science. The second is more exciting...
Others. Well, tell us both. [meaning: the usual Italian breaking rules...]
S. The first is about IR effectiveness metrics. We have around 100 metrics, counting the system-oriented ones only. The research project would be to find formal properties that each metric should satisfy.
Others. So do you mean...?
S. To define axioms that metrics should satisfy. For instance, first retrieved documents should weight more (or not less) than following docs, etc.
Others. Hm. What's that for?
S. Oh well, for instance we could understand why nDCG discount function is defined in that way...
Others. Cool. We need a name...
A. Axiometrics!
Everyone. That's great!!

[Chapter 2: Axiometrics]

S. Hi J.. Did you mention the idea about metrics in your group?
J. No, I didn't, and you?
S. Yes, I did, and guess it, it was selected!
J. Really?? Great!!
S. And A. invented a cool name: Axiometrics
J. That's wonderful!
S. Let's have some coffe, I'm still jetlagged.
J. Yes. And beach later. And beers.

[Chapter 3: Coffee and beers]

[after some hours, and coffee. And beach. And beers.]
S., J., A. M. N. Ok, let's write this short report about Axiometrics in SWIRL...
... well, let's have a short and last beach session first!


[Chapter 4: Planes, emails, and deadlines]


[After some long flights back home, tons of unanswered emails and student requests, expenses claim forms, etc. etc., and just a few spare days before the deadline...]
S. J., do you think it is reasonable to submit a Google research grant proposal on Axiometrics? Or is it a stupid idea?
J. That's a great idea!
S. Let's involve E. as well, he is interested.
[some hard science follows: bibliography is polished, CVs are created, margins and fonts are modified, PDFs generated...]


[Chapter 5: 4th of July]
Guys, we won!!

P.S. Thanks to Julio, Evangelos, ArjenMarteen, Nicola. And to SWIRL and Mark!


S.