Amy Guy

Raw Blog

Showing posts with label social machine. Show all posts
Showing posts with label social machine. Show all posts

Tuesday, February 18, 2014

Paper accepted to WWW

My first paper has been accepted to the SOCM14 workshop at WWW.

That means I get to go to Seoul, South Korea, in April!

I'll post a pre-print at some point.

Thursday, July 11, 2013

#SSSW2013: Social semantics and serendipity

We started work on the serendipity project before breakfast today, although I didn't make it down as early as some of my teammates.

To start the day, Fabio Ciravenga talked about some really exciting practical applications of monitoring and analysing social media streams.  It's particularly interesting during emergencies, or large events where problems might occur.  The people on the ground make the perfect sensors if you can work out the differences between people who are saying something useful and who aren't; people who are really there, and people who are speculating or asking about the situation.  A main problem has been that people tweet crap.  They were trying to monitor a house fire, but so many people were tweeting lyrics from Adele's various singles at the time, which all apparently contain references to fire, it was almost impossible.

They also put (or tapped into existing) sensors in peoples' cars to monitor driving patterns with the aim of more fairly charging for car insurance.  I told my Mum about this the other day, and she was pretty alarmed by the idea.  Which made me wonder how they'll get mass adoption, if it's going to go anywhere.

Fabio did have some interesting things to say about using all this data ethically though, and never working for someone who is going to take that away from you.  But in case the 'bad guys' do find out about all this data you have about people, keep a magnet handy.

My notes are here.

This was followed by a hands-on session where we got to mess with a mini version of the twitter topic monitoring system that Fabio's team use at large events, to try to answer questions about the Tour de France only by manipulating the incoming social media streams and following only links which came through that.

Spanish omelette sandwiches were an amazing outdoor leisurely lunch.  We headed to the pool down the road and chilled out there for a couple of hours.  Us tough British folk found the water pleasantly tepid, whilst all those wimpy Europeans and Latin Americans shivered on the grass.  They'd made such a fuss in advance about how cold the pool was going to be.

We regrouped that afternoon to work on Project Cusack, creating a slide deck of pictures from Serendipity.  I don't like slides with too much to read on, so I enforced this.  The imagery from the movie will be lost on most people, but we have at least managed to choose pictures of John Cusack with appropriate expressions for each part of the presentation.  We worked outside in the forest, because Oscar's 3G was faster than the residence wifi.



We also brainstormed for the required short film, which we only just discovered doesn't have to be about our project.

We returned to the residence to find everyone eating ham and cheese, and attempted to get some shots for our film, but other people were unwilling to participate.

That evening we ate tasty vegetable soup, weird (in a bad way) pasta in a creamy onion sauce, and chocolatey ice cream cake.  The tutors spontaneously organised a game where students had to arrange the tutors by age, which was funny.  Someone suggested the tutors ought to play it with the students.  Obviously there were too many students, but they elected to find the youngest student, and that turned out to be me.

[Notes] Fabio Ciravenga at #SSSW2013


Make a model of what is happening.

WeSenseIt - citizen water observations.

River belongs to citizens, not authorities.
Physical sensors (hard layer) are expensive and brittle.
So use people instead (soft layer, social).

Give people small sensors.  Phones.
Then you just need software for information management.

  • capture.
  • integrate and correlate data.
  • share.

Can't rely on phones.
Old people in Doncaster.

Give them easy sensors instead.

  • camera
  • humidity
  • position GPS
  • water depth, velocity
  • rainfall via accelerometer
  • could coverage via luminocity

Costs about EUR 80.

Open Source & hackable.

Not expected to substitute professional sensors, but a way to crowdsource information you would never get.


In Delft

Give people flood preparation advice and record who ticks things off, to build a picture of who/how/when preparations take place.


The Floow Ltd

"Commercialises data solution for telematic insurance."

World divided 10x10m squares, sense things everywhere.
Traffic risks.


Sensors tell you people are going somewhere, not why.
That's what social media can tell you.



Monitoring development of a house fire via Twitter.
Seeing events through the eyes of the community.

Social streams:

  • High volume
  • Duplicated, incomplete, imprecise, incorrect
  • Time sensitive / short term
  • Informal
  • Only 140 characters
  • Spam

Large music festival.  Monitor geolocated messages, trends, topics and relations.

Most 'critical' events were management issues.
Developing system to warn you automatically about things to pay attention to.

Look/listen for event within 72 hours.  10 minutes to find out what it was.
- Simulation of station bombing.
Minute by minute description of event.
1.5 billion messages.

  • Linguistic issues
    • Alternative language
    • Negatives
    • Conditional statements
    • Hope/prayer statements
    • Irony/sarcasm
    • Ambiguity
    • Unreliable capitalisation
    • Data sparsity

Four things when monitoring:

  • What
    • Identify, classify, cluster
      • Events and sub-events
      • Involved entities
  • Who
    • Human or not?
    • Bots can be beneign, but many are a serious risk.
    • Bots that pretend to be humans.
  • When
  • Where


Big problem - people tweet crap!
People don't realise when people nearby are in danger.


Deception on social media

False crowdsourcing political support on social networks.
Smear campaigns using bots.
Bots to foster / prevent social unrest.


Identifying bots

23 behavioural features.
Feature set is open.
Recognise 90% of bots - more than humans can do.



Very small amount of tweets are geolocated, it's useless.
Have to use the text.

Timestamp is not necessarily correct.


Issues in events

No infrastructure (eg. at music festivals).
Phone signal issues, phone charging issues.

Most tweets from outside event.

Conclusions

Need to convince citizens that authorities are not spying on them.
Need to convince authorities that citizens are not all criminals.

Privacy and legality issues.

Creating a company on this research would be unethical.
Need to pass the right message.  Full disclosure.  Non-intrusive use of tweet content.

What happens when authorities demand this technology for privacy-invading stuff.

Have to be careful with what you publish.
Always assume the bad guys have thought of what you thought of.
Always be in a situation where you can destroy your data at short notice.
Bit legal barrage behind them.  Know what they are/aren't allowed, know what they do/don't have to do.
Start leading a blameless life.

Friday, March 15, 2013

Notes about Meervisage - A Community Based Annotation Tool (for the Semantic Web)

Rowe, M. (2007)  Meervisage - A Community Based Annotation Tool. ‘Towards a Social Science of Web 2.0’ Conference at the University of York 5-6th September, 2007.

How SW can benefit from incorporation with existing 'Social Web'.
"...collaborative generation of metadata... using social networks as a user base..."

Uses fb groups created for sharing and organisation of research.  Suggests posting links to useful resources is comparable to annotating the resource.  Comments are more metadata.

Points out usual stuff of actually generating semantic data being a problem for SW.

System requirements:
  • Annotations must be shared in a community.
  • Annotations can be reviewed and edited (/audited) (by group)
  • Collaborative
  • Central repo.
  • Annotations contain semantic metadata.
  • Content of resource annotated, not URL.
  • Communication layer that doesn't interrupt annotation (uses external services).
Review of existing systems:
  • Annotea [9] [13]
    • J. Kahan, M.R. Koivunen, E. Prud Hommeaux, R.R. Swick. Annotea: an open RDF infrastructure for shared Web annotations. Computer Networks. 2002.
    • M Koivunen. Annotea and Semantic Web Supported Collaboration. Proc. Of 
    • the ESWC2005 Conference, 2005.
    • No communication layer (but has discussion threads, wat?). 
    • Can only be edited  by author, but can be reviewed by others.
    • Can be local, private or shared.  RDF.
  • Piggy Bank [10]
    • D Huynh, S Mazzocchi, D Karger. Piggy Bank: Experience the Semantic Web Inside Your Web Browser. Springer-Verlag GmbH. 2005. 
    • RDF. 
    • Auto and manual.  Bundled with scrapers; if they fail, manual.  Only of one type.
    • Share group or global, or save to local 'semantic bank' <-- find="" is="" out="" this="" what="">
    • Reviewed by all, edited by author.
    • Community of users, but no SNS integration.
  • KIM [14]
    • A Kiryakov, B Popov, D Ognyanoff, D Manov, A Kirilov, M Goranov. Semantic Annotation, Indexing and Retrieval. Journal of Web Semantics, Springer. 2004.
    • Automatic named entity recognition.  
    • Links to knowledgebase with ontology.
    • Creates new URIs for new entities or link swith entities it already knows about.
    • Global sharing.
    • Can be deleted but not edited.
    • No social involvement.
  • Magpie [11]
    • J Domingue, M Dzbor and E Motta. Semantic Layering with Magpie. Handbook on Ontologies. 2004.
    • Auto annotate webpage.
    • Similar to KIM, but does not hyperlink to knowledgebase; instead each item gets context menu (right click) with services depending on entity.
    • 'Multi-dimensional approach'. Uses ontology to trigger other services depending on concept.
    • Plugin for IE.  
    • Simply looks for entities that are in ontology (Dzbor 2004).
[1] Using existing information to derive semantics from folksonomies (delicious):
X Wu, L Zhang, Y Yu. Exploring social annotations for the Semantic Web. Proceedings of the 15th international conference on the World Wide Web, 2006.

[15] Social bookmarking tools and how semantic info aids resource discovery.  Probabalistic model of how resources are annotated:

A Plangprasopchok, K Lerman. Exploiting Social Annotation for Automatic Resource Discovery. Eprint arXiv, 2007.

[16] Distributed nature of folksonomies.  Improve search mechanisms.  Tags not great:

S Choy, A. Lui. Web Information Retrieval in Collaborative Tagging Systems. Proceedings of International Conference on Web Intelligence, 2006.
        (vs.)
[17] Rigid taxonomies not great:

C Shirky. Ontology is Overrated: Categories, Links, and Tags. Clay Shirky’s Writings About the Internet, 2005.

[18] Methodology for easier browsing of large scale social annotations:

Z Xu, Y Fu, J Mao, D Su. Towards the semantic web: Collaborative tag suggestions. Collaborative Web Tagging Workshop at WWW2006, 2006.

All use one annotation per resource, not annotation of content within, so only one lot of metadata about a page.

Meervisage

"To aid the process of collaborative annotation of web documents"

Allows sharing of annotations between subset of SNS users (eg. fb group).

Management of users and groups offloaded to third party.

Stored in central annotation store.

Annotations contain author, SNS, folksonomies and date.  Made from content within.

Meerkat is "responsible for generating semantic metadata by annotating external web resources."  Meervisage for management via social network.

Meerkat allows a user to edit another user's annotations if they are members of the same group on facebook.

Popularity rating of resources rises with fb discussion.
Meerkat informs browser users if they come across a resource that has been heavily discussed on fb, and by which group etc.

Meervisage also provides RSS feed.

Evaluate by comparing precision and recall metrics of annotations by one user in an allotted time, and those by a group.
-> Don't know how this helps to assess quality of annotations; maybe I'm dumb?  Find out.

Limited to private, says it's like that's a good think :s
Oh, because public access would be "laborious and resource intensive".

Annotations rated on usefulness and weighted.

[20] Attempt to describe folksonomies as part of formal ontology.  Meervisage doesn't; limited to users' viewpoint:

S Angeletou, M Sabou, L Specia, E Motta. Bridging the Gap Between Folksonomies and the Semantic Web: An Experience Report. Workshop: Bridging the Gap between Semantic Web and Web 2.0, European Semantic Web Conference, 2007.

[9] + [13] are most similar.  Have groups, but groups aren't already established networks.

Future work
Annotating multimedia.
Matching assigned tags with ontology terms mined from Web.
[19] Desktop app for annotating text with ontology:

A Chakravarthy, F Ciravegna, V Lanfranchi. AKTiveMedia: Cross-media Document Annotation and Enrichment. Poster Proceedings of the Fifteenth International Semantic Web Conference, 2006.


Sunday, March 03, 2013

Week in review: Reading and catching up

25th Feb - 3rd March

I spent most of this week in a remote village by the sea in the Scottish Highlands.

When I came back I read and made notes about the Semantic Web and social machines by TBL; and typed up some notes from a while ago about OntoMedia and the Semantic Web and communities by K. Faith Lawrence.

I also translated all of my written meeting notes into Evernote, which promptly glitched out and doubled the amount of typing I had to do.  (I considered switching back to Google Docs, but I need labels).  I sure love technology.

I refamiliarised myself with the structure OWL.  Awesome diagrams here.

I did some more thinking about how I need to work with amateur content creators to make an ontology that fits their workflow.  I should have finished the planning stage of this ages ago, but.. I blame the ILWhack.

I keep wondering about the best way to have a system of consistent URIs across a network where the information can move from server to server on the whim of a user.  During this wonderment I discovered that purl.org's login system is broken.  I joined the mailing list, and people complain about it and have it re-fixed fairly regularly, so I'll just wait..


Notes about the Semantic Web and social machines


J. Hendler, T. Berners-Lee, From the Semantic Web to social machines: A research challenge for AI on the World Wide Web, Artificial Intelligence (2009), doi:10.1016/j.artint.2009.11.010

Powerful human interactions enabled by futuristic high-speed infrastructure.
Empower Web of people via coupling of AI, social computing and new technologies.  "humanity in the loop".
Social machine: "...processes in which people do the creative work and the machine does the administration." (Weaving the Web, p172).
Struggling with social mechanisms to control predatory behaviour and threats to privacy.
-> Tech must be developed that allows user communities to construct / share / adapt social machines, so successful models evolve through trial, use and refinement.
Claims a new generation of Web Technologies needed to overcome barriers to this; cross-disciplinary approach needed.
  • creating tools
  • creating principles and guidelines
  • extending Web infrastructure re: information sharing and address privacy and user expectations of data use.

"...a revolutionarily more powerful platform for the individual, enabled by realizing that the individual is also a member of a community" (/ies)
"architecture of the future Web must be designed to allow the virtually unlimited interaction of the Web of people" (vs. documents now)

Giant Global Graph - dig.csail.mit.edu/breadcrumbs/node/215
SW deployment:
  • [3] T. Berners-Lee, J. Hendler, O. Lassila, The semantic web, Scientific American (May 2001) 28–37.
  • [10] J. Hendler, Web 3.0 emerging, IEEE Computer 42 (1) (January 2009).
  • [11] I. Jacobs, N. Walsh (Eds.), Architecture of the World Wide Web, Volume One, 2004, W3C Recommendation 15 December 2004, http://www.w3.org/TR/2004/REC-webarch-20041215/.


"disruptive potential" of SWt, "important paradigm shift"
"little work in understanding the impact of their new capability"
"the smaller we can make the individual steps of this transformation, the easier it will be to find humans who can be incentivized to perform those steps."
"need to develop mechanisms to enable [connections between people]"

Lack structure for formally computing qualities like:
  • trustworthiness
  • reliability
  • expectations about use of information
  • privacy
  • copyright
  • (etc)

"requires data structures... to treat social expectations and legal rules as first-class objects" ("declarative rule-based infrastructure that is appropriate for the Web").

"open and distributed nature of the Web requires that rule sets be linked together."Cross-context use, sometimes unanticipated.Inconsistency sure to arise.  No logics that control contradiction have been shown to scale well.
New approaches to problem of specifying contexts (need).
SMs must be able to apply different policies based on context.
Work in ontologies must extend to allow user communities to identify bias and share different interpretations.

Current security models / mechanisms insufficient.
[1] - formal models for privacy: L. Backstrom, C. Dwork, J. Kleinberg, Wherefore art thou r3579x?: Anonymized social networks, hidden patterns, and structural steganography, in: Proceedings of the 16th International World Wide Web Conference, Banff, 2007, pp. 181–190.
Provenance important in determining trustworthiness.
[20] information accountability, legal and public policy: D. Weitzner, H. Abelson, T. Berners-Lee, J. Feigenbaum, J. Hendler, G. Sussman, Information accountability, Communications of the ACM (June 2008).
policy-rule-based languages.Reasoners that can interpret policy and determine which uses of data are policy-compliant.-> How to tackle scaling?
RespectMyPrivacy dig.csail.mit.edu/2009/SocialWebPrivacy
[4] Lit review: T. Berners-Lee, W. Hall, J. Hendler, K. O’Hara, N. Shadbolt, D. Weitzner, A framework for web science, Foundations and Trends in Web Science 1 (1) (2006).