It is with much pleasure that I welcome Gianluca Demartini back at Udine University, Department of Mathematics and Computer Science, for a PhD course on Micro-task crowdsourcing. Gianluca got his Master's degree under my supervision some years ago (don't ask!) and since then he has obtained several research positions abroad. He is now Lecturer at the University of Sheffield.
The lecture timetable is below, together with a tentative program; everyone is welcome!
Timetable (all lectures are in the "Aula Multimediale / DIMI"):
1. Monday 15/6 10:00 - 12:00
2. Monday 15/6 15:00 - 17:00
3. Tuesday 16/6 10:00 - 12:00
4. Tuesday 16/6 15:00 - 17:00
5. Wednesday 17/6 10:00 - 12:00
6. Wednesday 17/6 15:00 - 17:00
7. Thursday 18/6 10:00 - 12:00
Preliminary/tentative program:
Lecture 1 - Introduction to Crowdsourcing
We will start with an overview of the entire module highlighting its
aims and objectives. Then, we will look at fundamental definitions and
different types of crowdsourcing incentives. Finally, we will present
early examples of crowdsourcing such as reCAPTCHA and the ESP game.
Lecture 2 - Introduction to Micro-task Crowdsourcing Platforms
After defining the key terminology of micro-task crowdsourcing, we will
introduce popular crowdsourcing platforms such as Amazon MTurk and
CrowdFlower including a demonstration on how to use such systems both as
a crowd worker as well as a requester.
Lecture 3 - How to Setup a Crowdsourcing Task
In this lecture we will discuss all the dimensions involved in
crowdsourcing task design such as pricing, question design, and quality
assurance mechanisms (e.g., honeypots). We will also design and deploy a
task during the lecture and see how to collect results back from the
crowdsourcing platform.
Lecture 4 - Crowdsourcing Patterns
In this lecture we will define the concept of crowdsourcing pattern
(i.e., the combination of multiple crowdsourcing tasks) and present
popular example patterns. We will also discuss the concept of
crowdsourcing workflows where multiple tasks as well as machine
processing steps are combined together.
Lecture 5 - Hybrid Human-machine Systems
In this lecture we will see some advanced example uses of crowdsourcing
applied to the database, web, and biomedical domains. We will see how
systems that combine both the scalability of machines over large amounts
of data as well as the quality of human intelligence can be used to
improve the effectiveness of Big Data systems.
Lecture 6 - Crowdsourcing Scalability
In hybrid human-machine systems the latency bottleneck lays on the side
of the crowd as human intelligence is naturally slower than
machine-based computation. In this lecture we will see recent research
results that proposed techniques to improve the latency of crowdsourcing
platforms.
Lecture 7 - Open Research Directions in Crowdsourcing
In this lecture we will give an overview on which micro-task
crowdsourcing research questions different Computer Science areas focus
on including database, information retrieval, semantic web,
human-computer interaction, multimedia, and bio-informatics.
Prerequisites:
No prior knowledge is needed. Having at least basic programming skills
is a plus.
S.
Showing posts with label IR. Show all posts
Showing posts with label IR. Show all posts
Saturday, 6 June 2015
Friday, 29 March 2013
Postdoc position
In the context of a Google Faculty Research Award (http://research.google.com/university/relations/research_awards.html), we are seeking a post doctoral researcher at Udine University (Udine, Italy). The title of the project is:
Axiometrics: Foundations of Evaluation Metrics in IR.
Axiometrics is one of the most important research directions proposed during the SWIRL 2012 meeting (http://www.cs.rmit.edu.au/swirl12/). A slightly more detailed description is at the end of this announcement. This is a joint project, involving:
*** Application deadline: 16 April 2013. ***
Short project description
Effectiveness evaluation is of paramount importance in the field of Information Retrieval (IR). IR is probably the most evaluation-oriented field in computer science, as witnessed by an evaluation methodology developed in the 60s during the Cranfield experiments and by several evaluation initiatives running today (TREC, CLEF, NTCIR, INEX, FIRE). One crucial aspect of evaluation are evaluation metrics. About 100 IR effectiveness metrics exist, and counting. This project aims at understanding the relationships among them, in terms of both axiomatic properties and statistical relations, for both metric science (understanding of metrics) and engineering (their development). More in detail, we aim at proposing:
Stefano Mizzaro
www.dimi.uniud.it/mizzaro
Axiometrics: Foundations of Evaluation Metrics in IR.
- Stefano Mizzaro (Principal Investigator), Dept. of Maths and Computer Science, University of Udine, Italy, mizzaro_foo@uniud.it.
- Julio Gonzalo (co-Principal Investigator) ed Enrique Amigó, E.T.S.I. Informática de la UNED, Madrid, Spain, julio_foo@lsi.uned.es, enrique_foo@lsi.uned.es.
- Evangelos Kanoulas (Google sponsor), Google Zurigo, ekanoulas_foo@gmail.com.
*** Application deadline: 16 April 2013. ***
Short project description
Effectiveness evaluation is of paramount importance in the field of Information Retrieval (IR). IR is probably the most evaluation-oriented field in computer science, as witnessed by an evaluation methodology developed in the 60s during the Cranfield experiments and by several evaluation initiatives running today (TREC, CLEF, NTCIR, INEX, FIRE). One crucial aspect of evaluation are evaluation metrics. About 100 IR effectiveness metrics exist, and counting. This project aims at understanding the relationships among them, in terms of both axiomatic properties and statistical relations, for both metric science (understanding of metrics) and engineering (their development). More in detail, we aim at proposing:
- Axioms: rules that any metric must satisfy. For example, when swapping a relevant document and a non-relevant one in the ranking, by decreasing the rank of the relevant one and increasing the rank of the non relevant one, the metric value should decrease. Axioms might be verified and compared with user intuition using crowdsourcing.
- Desiderata: desirable properties that a metric should have, based on common sense. For example, an effectiveness value according to one metric should be affected more by a swap in earlier rank positions than a swap in later ranks.
- Empirical properties: those that emerge from data, i.e., from actual test collections and system comparisons. These include robustness, statistical correlation between metrics, etc.
Stefano Mizzaro
www.dimi.uniud.it/mizzaro
Labels:
IR,
university
Wednesday, 4 July 2012
"No flashy. This is Science." [censored] -- Or: How two jetlagged IR researchers met and had an idea
[Prologue: the beginning of a beautiful friendship]
J. Hi. Can I sit here?
S. Sure. Hi. I don't think we met before. Nice to meet you, I'm S.
J. I'm J.
S. Nice place here, isn't it?
J. Yep! Beach this afternoon?
S. Sure!
J. Ok, let's pretend we do some work first. So, what are you working on?
S. Oh, several things, blablabla [snip] And you?
J. Well, I did blablabla... and blabla... and I also published some papers on formal analysis of clustering metrics.
S. Interesting. I also started something similar for IR metrics years ago, but I never managed to publish it.
J. Really?
S. Yep. I'll show you [opening his laptop]. See, I defined a framework based on measurement theory, then I defined some axioms...
J. That's crazy! I did exactly the same!
S. ... then some desiderata...
J. Exactly the same!
S. ... Yes but you published, I didn't...
J. ... because you tried to do everything, look here, that's crazy!
S. Yes, I'm a bit, sure... and then blablabla...
J. blablabla [very technical details here - ok, ok, I admit I do not remember that!]
S. blablabla [very technical details here - ok, ok, I admit I do not remember that!]
J. Well, why not proposing this as an idea for SWIRL research directions?
S. We might indeed!
J. Ok, beach time now.
S. Beach!
[Chapter 1: split groups doing Science]
M. So, let's have a round of the table and everyone presents his own idea.
[various good ideas...everyone presents just one, accordingly to the rules. But you know that Italians and Rules do not fit well in the same sentence, so...]
S. I do have two ideas. The first is not very exciting, but it's Science. The second is more exciting...
Others. Well, tell us both. [meaning: the usual Italian breaking rules...]
S. The first is about IR effectiveness metrics. We have around 100 metrics, counting the system-oriented ones only. The research project would be to find formal properties that each metric should satisfy.
Others. So do you mean...?
S. To define axioms that metrics should satisfy. For instance, first retrieved documents should weight more (or not less) than following docs, etc.
Others. Hm. What's that for?
S. Oh well, for instance we could understand why nDCG discount function is defined in that way...
Others. Cool. We need a name...
A. Axiometrics!
Everyone. That's great!!
[Chapter 2: Axiometrics]
S. Hi J.. Did you mention the idea about metrics in your group?
J. No, I didn't, and you?
S. Yes, I did, and guess it, it was selected!
J. Really?? Great!!
S. And A. invented a cool name: Axiometrics
J. That's wonderful!
S. Let's have some coffe, I'm still jetlagged.
J. Yes. And beach later. And beers.
[Chapter 3: Coffee and beers]
[after some hours, and coffee. And beach. And beers.]
S., J., A. M. N. Ok, let's write this short report about Axiometrics in SWIRL...
... well, let's have a short and last beach session first!
[Chapter 4: Planes, emails, and deadlines]
[After some long flights back home, tons of unanswered emails and student requests, expenses claim forms, etc. etc., and just a few spare days before the deadline...]
S. J., do you think it is reasonable to submit a Google research grant proposal on Axiometrics? Or is it a stupid idea?
J. That's a great idea!
S. Let's involve E. as well, he is interested.
[some hard science follows: bibliography is polished, CVs are created, margins and fonts are modified, PDFs generated...]
[Chapter 5: 4th of July]
Guys, we won!!
P.S. Thanks to Julio, Evangelos, Arjen, Marteen, Nicola. And to SWIRL and Mark!
S.
J. Hi. Can I sit here?
S. Sure. Hi. I don't think we met before. Nice to meet you, I'm S.
J. I'm J.
S. Nice place here, isn't it?
J. Yep! Beach this afternoon?
S. Sure!
J. Ok, let's pretend we do some work first. So, what are you working on?
S. Oh, several things, blablabla [snip] And you?
J. Well, I did blablabla... and blabla... and I also published some papers on formal analysis of clustering metrics.
S. Interesting. I also started something similar for IR metrics years ago, but I never managed to publish it.
J. Really?
S. Yep. I'll show you [opening his laptop]. See, I defined a framework based on measurement theory, then I defined some axioms...
J. That's crazy! I did exactly the same!
S. ... then some desiderata...
J. Exactly the same!
S. ... Yes but you published, I didn't...
J. ... because you tried to do everything, look here, that's crazy!
S. Yes, I'm a bit, sure... and then blablabla...
J. blablabla [very technical details here - ok, ok, I admit I do not remember that!]
S. blablabla [very technical details here - ok, ok, I admit I do not remember that!]
J. Well, why not proposing this as an idea for SWIRL research directions?
S. We might indeed!
J. Ok, beach time now.
S. Beach!
[Chapter 1: split groups doing Science]
M. So, let's have a round of the table and everyone presents his own idea.
[various good ideas...everyone presents just one, accordingly to the rules. But you know that Italians and Rules do not fit well in the same sentence, so...]
S. I do have two ideas. The first is not very exciting, but it's Science. The second is more exciting...
Others. Well, tell us both. [meaning: the usual Italian breaking rules...]
S. The first is about IR effectiveness metrics. We have around 100 metrics, counting the system-oriented ones only. The research project would be to find formal properties that each metric should satisfy.
Others. So do you mean...?
S. To define axioms that metrics should satisfy. For instance, first retrieved documents should weight more (or not less) than following docs, etc.
Others. Hm. What's that for?
S. Oh well, for instance we could understand why nDCG discount function is defined in that way...
Others. Cool. We need a name...
A. Axiometrics!
Everyone. That's great!!
[Chapter 2: Axiometrics]
S. Hi J.. Did you mention the idea about metrics in your group?
J. No, I didn't, and you?
S. Yes, I did, and guess it, it was selected!
J. Really?? Great!!
S. And A. invented a cool name: Axiometrics
J. That's wonderful!
S. Let's have some coffe, I'm still jetlagged.
J. Yes. And beach later. And beers.
[Chapter 3: Coffee and beers]
[after some hours, and coffee. And beach. And beers.]
S., J., A. M. N. Ok, let's write this short report about Axiometrics in SWIRL...
... well, let's have a short and last beach session first!
[Chapter 4: Planes, emails, and deadlines]
[After some long flights back home, tons of unanswered emails and student requests, expenses claim forms, etc. etc., and just a few spare days before the deadline...]
S. J., do you think it is reasonable to submit a Google research grant proposal on Axiometrics? Or is it a stupid idea?
J. That's a great idea!
S. Let's involve E. as well, he is interested.
[some hard science follows: bibliography is polished, CVs are created, margins and fonts are modified, PDFs generated...]
[Chapter 5: 4th of July]
Guys, we won!!
P.S. Thanks to Julio, Evangelos, Arjen, Marteen, Nicola. And to SWIRL and Mark!
S.
Labels:
IR,
research,
science,
university
Friday, 24 February 2012
SWIRL 2012
Last week I attended SWIRL 2012: these are my highlights.
Facts.
SWIRL 2012 has been a workshop where around 50 jetlagged top IR researchers gathered together essentially to experience summer during winter and incidentally to discuss the future of IR research. I am not sure that I count among the top IR researchers. Actually, probably the selection mechanism was more oriented towards the most crazy IR researchers since some really top were missing and some very crazy were definitely there. So I am not sure that I count among the top IR researchers, we can discuss if I count among the top crazy ones, but I'm sure I count among the top jetlagged ones. Luckily enough, all the others were jetlagged as well, and looked like Australians in Europe, so I managed to camouflage myself somehow. Also, and surprisingly, being very jetlagged even helped, since we had long discussions during the nights, and we managed to use the nights to read our emails :) and have the days free for working.
The workshop started with the usual nice dinner in a nice hotel. Well, actually we had a prologue, with some talks at RMIT from some IR researchers. They were, guess, extremely jetlagged, so probably not all the talks were so good, and yes I was one of the speakers. You see what I mean? Anyway, I presented Readersourcing, and I learned two things:
users People, of course (fair enough, we didn't write "students"). SmartIR is the smarter term in the list. I've been working for the last 10 years, and published most of my last papers, on "(less than) zero query" [sic], although with different, and equally stupid strange labels like "query-free IR", "zero-interaction interactive IR", "browsing the virtual space by walking in the physical space", etc.
But at the end I joined the Mobile IR group. And, guess it, it was not about Mobile IR. It was about Understanding what Mobile IR is. Funny: a group that did not understand what it was about but aimed at understanding what its (mistaken) topic was about. How couldn't we succeed? Plus, when people chose their own group, strange things happened. Some people were not able to find their group; perhaps it was on the nice beach? Some people (those above plus others) joined another "wrong" group, and some of them, just to avoid being idle, decided to act as trolls. We had two trolls in the Mobile IR group, and you will understand how that made the discussion far more interesting, lively, polite, and constructive than you could ever imagine. Luckily, the night came, and after a good sleep (or perhaps some good chats, beers, barbecues, etc., since we were seriously jetlagged and can't manage to have that much sleep anyway), the morning after, Jamie and Vanessa had a clear vision of what the group was about and started drafting the report (each group was meant to write a report). Not having trolls around early morning probably helped. Did trolls suffer from jetlag? Did trolls take surfer lessons? We will never know.
As mentioned above, the other not-top-6-best-great-ideas were not killed, and some still-quite-heavily-jetlagged-IR-researchers volunteered to write a short report on those. Now, I have to mention the best name of the workshop: Axiometrics (credits to Arjen), a research line aimed at defining axioms in order to understand the about 100 IR effectiveness metrics. BTW, we might think of Anatometrics as well.
Comments?
So the workshop was great, the organizers incredibly managed to obtain something out of about-50-top-crazy-IR-researchers-that-as-you-know-were-very-jetlagged, the discussions were interesting, I managed to have some good ideas and contacts for future paper writing, and perhaps some contact for my sabbatical as well (I'm looking for places were to stay during my sabbatical; please let me know if you're interested. I promise that I won't be so jetlagged for the whole sabbatical duration.) Plus, I've been out of business recently, for several reasons including lack of funds, A.'s birth, etc., and it was really nice to meet some old good friends and some new ones.
Criticisms? I always have. We could have made use of some "Social Web/Web2.0" tools, like Gdocs, Facebook, Twitter, to have a virtual discussion as well and, for example, to vote the ideas (did I hint that the voting mechanism was a bit... "italianized"? Now, we all know about Arrow's theorem, but that was far beyond that). Someone said that s/he had the impression that we were simply drafting a report to make easier for US and Australian researchers to get funds. Someone had the impression that the meeting was a bit too "old fashioned". But as I wrote, the organizers did manage to get something out of about-50-top-IR-researchers-that-as-you-know-were-very-jetlagged, and this is an enormous success.
Quotations!
Some interesting sentences were uttered during those days, and shall never be forgotten:
Once back in Melbourne, a (randomly) selected subgroup of all those (somehow) 50 selected IR researcher had an interesting post-workshop, post-dinner, during-beer mini-workshop on p*orn and user models. I took some pictures of the participants:
As you can see, we range from someone (pretending to be?) not interested, or perhaps simply jetlagged, to someone really having fun, to a shy guy who doesn't want to be recognized, to someone counting beers (not an easy task!), to someone hiding in the shade. I learned some interesting statistics about user features. Anyway, guys, it was fun. Thanks!
S.
Facts.
SWIRL 2012 has been a workshop where around 50 jetlagged top IR researchers gathered together essentially to experience summer during winter and incidentally to discuss the future of IR research. I am not sure that I count among the top IR researchers. Actually, probably the selection mechanism was more oriented towards the most crazy IR researchers since some really top were missing and some very crazy were definitely there. So I am not sure that I count among the top IR researchers, we can discuss if I count among the top crazy ones, but I'm sure I count among the top jetlagged ones. Luckily enough, all the others were jetlagged as well, and looked like Australians in Europe, so I managed to camouflage myself somehow. Also, and surprisingly, being very jetlagged even helped, since we had long discussions during the nights, and we managed to use the nights to read our emails :) and have the days free for working.
The workshop started with the usual nice dinner in a nice hotel. Well, actually we had a prologue, with some talks at RMIT from some IR researchers. They were, guess, extremely jetlagged, so probably not all the talks were so good, and yes I was one of the speakers. You see what I mean? Anyway, I presented Readersourcing, and I learned two things:
- Now I know who was one of the referees of the famous Go*gle/PageRank paper rejected from SIGIR; and that that submitted version wasn't that good... Anyway...
- When presenting Readersourcing in these years, I managed to discuss... ehm let's say to argue with S, K, and B. I did not record the first "discussion", but I did for the second one here. Now, the point is: shouldn't I seriously consider to quit this line of research?!?
- Evaluation? We haven't had enough.
- How to wreck a nice beach (don't ask me to explain this, ask David).
- You can live your life as a party (ask Leif)
- Mobile IR
- Structured and unstructured information
usersPeople- SmartIR (aka information literacy)
- "(less than) zero query" [sic]
- and, of course, Argy-Bargy
But at the end I joined the Mobile IR group. And, guess it, it was not about Mobile IR. It was about Understanding what Mobile IR is. Funny: a group that did not understand what it was about but aimed at understanding what its (mistaken) topic was about. How couldn't we succeed? Plus, when people chose their own group, strange things happened. Some people were not able to find their group; perhaps it was on the nice beach? Some people (those above plus others) joined another "wrong" group, and some of them, just to avoid being idle, decided to act as trolls. We had two trolls in the Mobile IR group, and you will understand how that made the discussion far more interesting, lively, polite, and constructive than you could ever imagine. Luckily, the night came, and after a good sleep (or perhaps some good chats, beers, barbecues, etc., since we were seriously jetlagged and can't manage to have that much sleep anyway), the morning after, Jamie and Vanessa had a clear vision of what the group was about and started drafting the report (each group was meant to write a report). Not having trolls around early morning probably helped. Did trolls suffer from jetlag? Did trolls take surfer lessons? We will never know.
As mentioned above, the other not-top-6-best-great-ideas were not killed, and some still-quite-heavily-jetlagged-IR-researchers volunteered to write a short report on those. Now, I have to mention the best name of the workshop: Axiometrics (credits to Arjen), a research line aimed at defining axioms in order to understand the about 100 IR effectiveness metrics. BTW, we might think of Anatometrics as well.
Comments?
So the workshop was great, the organizers incredibly managed to obtain something out of about-50-
Criticisms? I always have. We could have made use of some "Social Web/Web2.0" tools, like Gdocs, Facebook, Twitter, to have a virtual discussion as well and, for example, to vote the ideas (did I hint that the voting mechanism was a bit... "italianized"? Now, we all know about Arrow's theorem, but that was far beyond that). Someone said that s/he had the impression that we were simply drafting a report to make easier for US and Australian researchers to get funds. Someone had the impression that the meeting was a bit too "old fashioned". But as I wrote, the organizers did manage to get something out of about-50-
Quotations!
Some interesting sentences were uttered during those days, and shall never be forgotten:
- Cloud computing is not transparent.
- You can't leave indoor outdoor, but you can't leave outdoor outdoor either.
- How to wreck a nice beach.
- Life as a party.
- (Julio please help with the other one)
- (anybody welcome to add)
Once back in Melbourne, a (randomly) selected subgroup of all those (somehow) 50 selected IR researcher had an interesting post-workshop, post-dinner, during-beer mini-workshop on p*orn and user models. I took some pictures of the participants:
As you can see, we range from someone (pretending to be?) not interested, or perhaps simply jetlagged, to someone really having fun, to a shy guy who doesn't want to be recognized, to someone counting beers (not an easy task!), to someone hiding in the shade. I learned some interesting statistics about user features. Anyway, guys, it was fun. Thanks!
S.
Labels:
Cultura,
IR,
lavoro,
science,
unforgettable
Friday, 13 November 2009
IIR 2010
I though that I would never have organized something again, but I was wrong. Yes: errare humanum est, sed perseverare diabolicum. But this sounds interesting.
I'm organizing the Italian Information Retrieval workshop. Feel free to contact me for further infos. And feel free to submit your papers, of course!!
S.
I'm organizing the Italian Information Retrieval workshop. Feel free to contact me for further infos. And feel free to submit your papers, of course!!
S.
Labels:
IR
Subscribe to:
Posts (Atom)