Amy Guy

Raw Blog

Showing posts with label semantics. Show all posts
Showing posts with label semantics. Show all posts

Thursday, July 11, 2013

#SSSW2013: Social semantics and serendipity

We started work on the serendipity project before breakfast today, although I didn't make it down as early as some of my teammates.

To start the day, Fabio Ciravenga talked about some really exciting practical applications of monitoring and analysing social media streams.  It's particularly interesting during emergencies, or large events where problems might occur.  The people on the ground make the perfect sensors if you can work out the differences between people who are saying something useful and who aren't; people who are really there, and people who are speculating or asking about the situation.  A main problem has been that people tweet crap.  They were trying to monitor a house fire, but so many people were tweeting lyrics from Adele's various singles at the time, which all apparently contain references to fire, it was almost impossible.

They also put (or tapped into existing) sensors in peoples' cars to monitor driving patterns with the aim of more fairly charging for car insurance.  I told my Mum about this the other day, and she was pretty alarmed by the idea.  Which made me wonder how they'll get mass adoption, if it's going to go anywhere.

Fabio did have some interesting things to say about using all this data ethically though, and never working for someone who is going to take that away from you.  But in case the 'bad guys' do find out about all this data you have about people, keep a magnet handy.

My notes are here.

This was followed by a hands-on session where we got to mess with a mini version of the twitter topic monitoring system that Fabio's team use at large events, to try to answer questions about the Tour de France only by manipulating the incoming social media streams and following only links which came through that.

Spanish omelette sandwiches were an amazing outdoor leisurely lunch.  We headed to the pool down the road and chilled out there for a couple of hours.  Us tough British folk found the water pleasantly tepid, whilst all those wimpy Europeans and Latin Americans shivered on the grass.  They'd made such a fuss in advance about how cold the pool was going to be.

We regrouped that afternoon to work on Project Cusack, creating a slide deck of pictures from Serendipity.  I don't like slides with too much to read on, so I enforced this.  The imagery from the movie will be lost on most people, but we have at least managed to choose pictures of John Cusack with appropriate expressions for each part of the presentation.  We worked outside in the forest, because Oscar's 3G was faster than the residence wifi.



We also brainstormed for the required short film, which we only just discovered doesn't have to be about our project.

We returned to the residence to find everyone eating ham and cheese, and attempted to get some shots for our film, but other people were unwilling to participate.

That evening we ate tasty vegetable soup, weird (in a bad way) pasta in a creamy onion sauce, and chocolatey ice cream cake.  The tutors spontaneously organised a game where students had to arrange the tutors by age, which was funny.  Someone suggested the tutors ought to play it with the students.  Obviously there were too many students, but they elected to find the youngest student, and that turned out to be me.

[Notes] Fabio Ciravenga at #SSSW2013


Make a model of what is happening.

WeSenseIt - citizen water observations.

River belongs to citizens, not authorities.
Physical sensors (hard layer) are expensive and brittle.
So use people instead (soft layer, social).

Give people small sensors.  Phones.
Then you just need software for information management.

  • capture.
  • integrate and correlate data.
  • share.

Can't rely on phones.
Old people in Doncaster.

Give them easy sensors instead.

  • camera
  • humidity
  • position GPS
  • water depth, velocity
  • rainfall via accelerometer
  • could coverage via luminocity

Costs about EUR 80.

Open Source & hackable.

Not expected to substitute professional sensors, but a way to crowdsource information you would never get.


In Delft

Give people flood preparation advice and record who ticks things off, to build a picture of who/how/when preparations take place.


The Floow Ltd

"Commercialises data solution for telematic insurance."

World divided 10x10m squares, sense things everywhere.
Traffic risks.


Sensors tell you people are going somewhere, not why.
That's what social media can tell you.



Monitoring development of a house fire via Twitter.
Seeing events through the eyes of the community.

Social streams:

  • High volume
  • Duplicated, incomplete, imprecise, incorrect
  • Time sensitive / short term
  • Informal
  • Only 140 characters
  • Spam

Large music festival.  Monitor geolocated messages, trends, topics and relations.

Most 'critical' events were management issues.
Developing system to warn you automatically about things to pay attention to.

Look/listen for event within 72 hours.  10 minutes to find out what it was.
- Simulation of station bombing.
Minute by minute description of event.
1.5 billion messages.

  • Linguistic issues
    • Alternative language
    • Negatives
    • Conditional statements
    • Hope/prayer statements
    • Irony/sarcasm
    • Ambiguity
    • Unreliable capitalisation
    • Data sparsity

Four things when monitoring:

  • What
    • Identify, classify, cluster
      • Events and sub-events
      • Involved entities
  • Who
    • Human or not?
    • Bots can be beneign, but many are a serious risk.
    • Bots that pretend to be humans.
  • When
  • Where


Big problem - people tweet crap!
People don't realise when people nearby are in danger.


Deception on social media

False crowdsourcing political support on social networks.
Smear campaigns using bots.
Bots to foster / prevent social unrest.


Identifying bots

23 behavioural features.
Feature set is open.
Recognise 90% of bots - more than humans can do.



Very small amount of tweets are geolocated, it's useless.
Have to use the text.

Timestamp is not necessarily correct.


Issues in events

No infrastructure (eg. at music festivals).
Phone signal issues, phone charging issues.

Most tweets from outside event.

Conclusions

Need to convince citizens that authorities are not spying on them.
Need to convince authorities that citizens are not all criminals.

Privacy and legality issues.

Creating a company on this research would be unethical.
Need to pass the right message.  Full disclosure.  Non-intrusive use of tweet content.

What happens when authorities demand this technology for privacy-invading stuff.

Have to be careful with what you publish.
Always assume the bad guys have thought of what you thought of.
Always be in a situation where you can destroy your data at short notice.
Bit legal barrage behind them.  Know what they are/aren't allowed, know what they do/don't have to do.
Start leading a blameless life.

Wednesday, July 10, 2013

#SSSW2013: Practical semantics and human nature

Harith Alani talked about using semantics to solve problems around evaluating the success of social media use in business.  The SIOC ontology is widely used to describe online community information.  It's not as simple as measuring someone's engagement with a brand's online presence - people are 'likeaholics' on Facebook, so you have to look at someone's whole behaviour profile to judge whether their like means anything or not.  It's no good just aggregating your data and spewing out numbers - you have to browse the data and try to understand where it came from.

He mentioned how little work has been done in classifying community types.  Most of the work that has been done seems to be with social networks internal to an organisation.  A bottom-up approach to community analysis can handle emergent behaviours and cope with role changes over time.  Looking at behaviour categories and roles can help an organisation to decide who to concentrate on supporting and how in order to sustain the community.  The results they have seen so far suggest that a stable mix of the different types of behaviours are needed to increase activities in forums - but they don't know what causes what.  They're reaching a point where they can use their behaviour analysis to guess what's going to happen to a community: how long it will last, how fast it will grow, how many replies a certain type of post is likely to get, etc.

Next they want to be able to classify community types, and be able to look at activities within a community over a period of time and automatically discover what kind of community it is; it might be something different than what it was set up for.

They created an alternative Maslow's Hierarchy of Needs to correspond with activities seen on forums, and found that most people are happy to stay at the lower levels of the hierarchy.  For example, join a community, lurk for a bit, ask one question and leave.  Not everyone wants or needs to be a power user.

Papers are being written that find patterns in individual datasets for a particular community in a particular context.  Harith and his team are getting tired of this; they want to generalise across communities.  So they took seven datasets and looked at how the analysis features differed as well as comparing the results across community types, randomness (vs. topicality) of datasets, and compared similar experiments.

Upcoming work includes the Reel Lives project, in which UoE is involved.  They're taking media fragments - photos, videos, audio clips, text recorded as audio - and creating automated compilations to tell a story.

Another is social methods to change energy consumption behaviour.  LiSC in Lincoln did something in this area back in the day.. an app that posted that you were listening to an embarrassing song on your facebook feed if you left your lights on.

Notes from Harith's talk are here.



From Tommaso Di Noia's talk, I learnt that recommender systems have a lot of maths behind them, especially for evaluating things, and reinforced something I already knew: I don't maths good enough to be taken seriously by most of the Informatics world.  I think I understand the principles behind the maths, but when something is descried in just maths, I have no idea what it relates to.  I'll work on this.

Real world recommender systems use a variety of approaches, including collaborative (based on similar users' profiles); knowledge-based (domain knowledge, no user history); item-based (similarities between items); content-based (combination of item descriptions and profile of user interests).  Linked Open Data is used to mitigate a lack of information about entities, and helps with recommending across multiple domains.  You do have to filter the LD you use before feeding it to your recommender system though, to avoid noise.  Notes here.

Tommaso's talk was followed up by a hands-on session, where we got to poke about with some of the tools he mentioned, including FRED (transforms natural language to RDF/OWL); Tipalo (gets entity types from natural language text); and using DBpedia to feed a recommender system.

Then we worked on our mini-projects for the afternoon.  We made some progress towards breaking down the concept of serendipity and working out what properties we might need to represent as linked data, and how we could observer a user and work out if/when/how they were having serendipitous experiences without intruding too much.

In the evening we took a coach to 'nearby' historical town Segovia.  Apparently an extremely motion-sickness-inducing two and a half hour coach journey around twisty mountain paths is 'nearby'.  Fortunately I was distracted from this horrible journey by a conversation with Lynda Hardman, which I wish I had recorded.  Lynda challenged various aspects of my PhD until I could explain/justify them reasonably, including:

  • Why digital creatives? (I'm used to that one now).
  • What is the outcome?
  • Why Semantic Web for this?

She also recommended a number of resources, including theses of her recent former students to help me with a structure for my own, and advice on maintaining a healthy balance between thinking and doing.

Plus she used to live in Edinburgh, more or less across the road from where I live now.  Cool.  Thanks Lynda!  You haven't heard the last of me :)

#travel

Once we got to Segovia, we had a guided tour of the ancient Roman architecture, interesting building façades and local legends.  It was a very good tour, but too hot to really focus.  Then they took us to a restaurant for a local speciality.  I was all set to write a whole individual blog post surveying the barbaric nature of human beings, but I didn't do it straight away and now the passion has faded slightly, so I'll leave it at a paragraph.  Some people watched the local 'ceremony' out of morbid curiosity I imagine, but it was the fact that so many people took so much pleasure in the idea of violently hacking up bodies of three-week-old piglets that really bothered me.  Fortunately the surging standing crowd allowed me (and only one other) to inconspicuously sit it out.  The veggie option was tasty, but it was difficult to really enjoy the rest of the evening whilst wondering vaguely about the states of minds of most of the people I was sharing a table with.

Monday, May 13, 2013

Week in review: annotating multimedia content

6th - 12th May

Discovering lots of things to write about semantically annotating multimedia content.  I decided there are three main ways to do this:
  • Technical / objective / statistical data: eg. media type; shutterspeed; framerate; duration; resolution; date created; number of times viewed at a particular source; number of time shared ...
  • Bibliographic: creators and contributors and their roles; methods/location of publication; methods/locations of creation ...
  • Content*: fictional characters; locations; camera movements; scene transitions; colours ...
These categories overlap somewhat really, and when I get round to it I'll type my Venn diagram up.

Technical is easy, and a lot of that is automatically captured by hardware or software used to produce and edit works.  It's also relatively easy to extract automatically.  Standards like MPEG-7 and MPEG-21 take care of formalising it, and Jane Hunter turned these standards into semantic ontologies in 2002.

Bibliographic can largely - but not entirely - be covered by vocabularies that have been around forever like Dublin Core, FOAF and various library-originated things.  Things that might be missing (or I just haven't found them yet) are associating roles with tasks involved in digital media production, since pieces are often a collaborative effort.   has some idea of participants and roles, but the purpose of  is digital rights management stuff, so it's more concerned with the distribution change, I think, than granular production of content.  I haven't read much about it yet.

Content is more interesting, and potentially more useful for ordinary human beings.  Imagine querying IMDB for "that film where John Goodman arrests an animated talking moose on a US highway" instead of scouring John Goodman's filmography or googling for pictures of animated meese until you see the right one.  Annotating characters, objects and events, and stringing them onto a timeline is possible with OntoMedia.  It's very focussed around narratives, which is great, but doesn't link back to technical so much.  So if you did find the answer to that query, it wouldn't be able to serve up the timestamp of that particular scene.

On top of what I've looked at already, I still have this list to (re)investigate: 
A thing I want to do is annotate some amateur content with OntoMedia and with ABC to see how they compare.  Maybe I'll do asdfmovie, because it has associated comics, and multiple people participating in production.  Then I'll do something live action as well, because I can't base all my research on non-sequitur lolrandom stick figure cartoons.

Now, back to work..

* I want a better name for this, since I'm referring to everything as 'content' anyway.  So some better way of saying 'content of content'.