Amy Guy

Raw Blog

Showing posts with label ontologies. Show all posts
Showing posts with label ontologies. Show all posts

Wednesday, July 10, 2013

#SSSW2013: Practical semantics and human nature

Harith Alani talked about using semantics to solve problems around evaluating the success of social media use in business.  The SIOC ontology is widely used to describe online community information.  It's not as simple as measuring someone's engagement with a brand's online presence - people are 'likeaholics' on Facebook, so you have to look at someone's whole behaviour profile to judge whether their like means anything or not.  It's no good just aggregating your data and spewing out numbers - you have to browse the data and try to understand where it came from.

He mentioned how little work has been done in classifying community types.  Most of the work that has been done seems to be with social networks internal to an organisation.  A bottom-up approach to community analysis can handle emergent behaviours and cope with role changes over time.  Looking at behaviour categories and roles can help an organisation to decide who to concentrate on supporting and how in order to sustain the community.  The results they have seen so far suggest that a stable mix of the different types of behaviours are needed to increase activities in forums - but they don't know what causes what.  They're reaching a point where they can use their behaviour analysis to guess what's going to happen to a community: how long it will last, how fast it will grow, how many replies a certain type of post is likely to get, etc.

Next they want to be able to classify community types, and be able to look at activities within a community over a period of time and automatically discover what kind of community it is; it might be something different than what it was set up for.

They created an alternative Maslow's Hierarchy of Needs to correspond with activities seen on forums, and found that most people are happy to stay at the lower levels of the hierarchy.  For example, join a community, lurk for a bit, ask one question and leave.  Not everyone wants or needs to be a power user.

Papers are being written that find patterns in individual datasets for a particular community in a particular context.  Harith and his team are getting tired of this; they want to generalise across communities.  So they took seven datasets and looked at how the analysis features differed as well as comparing the results across community types, randomness (vs. topicality) of datasets, and compared similar experiments.

Upcoming work includes the Reel Lives project, in which UoE is involved.  They're taking media fragments - photos, videos, audio clips, text recorded as audio - and creating automated compilations to tell a story.

Another is social methods to change energy consumption behaviour.  LiSC in Lincoln did something in this area back in the day.. an app that posted that you were listening to an embarrassing song on your facebook feed if you left your lights on.

Notes from Harith's talk are here.



From Tommaso Di Noia's talk, I learnt that recommender systems have a lot of maths behind them, especially for evaluating things, and reinforced something I already knew: I don't maths good enough to be taken seriously by most of the Informatics world.  I think I understand the principles behind the maths, but when something is descried in just maths, I have no idea what it relates to.  I'll work on this.

Real world recommender systems use a variety of approaches, including collaborative (based on similar users' profiles); knowledge-based (domain knowledge, no user history); item-based (similarities between items); content-based (combination of item descriptions and profile of user interests).  Linked Open Data is used to mitigate a lack of information about entities, and helps with recommending across multiple domains.  You do have to filter the LD you use before feeding it to your recommender system though, to avoid noise.  Notes here.

Tommaso's talk was followed up by a hands-on session, where we got to poke about with some of the tools he mentioned, including FRED (transforms natural language to RDF/OWL); Tipalo (gets entity types from natural language text); and using DBpedia to feed a recommender system.

Then we worked on our mini-projects for the afternoon.  We made some progress towards breaking down the concept of serendipity and working out what properties we might need to represent as linked data, and how we could observer a user and work out if/when/how they were having serendipitous experiences without intruding too much.

In the evening we took a coach to 'nearby' historical town Segovia.  Apparently an extremely motion-sickness-inducing two and a half hour coach journey around twisty mountain paths is 'nearby'.  Fortunately I was distracted from this horrible journey by a conversation with Lynda Hardman, which I wish I had recorded.  Lynda challenged various aspects of my PhD until I could explain/justify them reasonably, including:

  • Why digital creatives? (I'm used to that one now).
  • What is the outcome?
  • Why Semantic Web for this?

She also recommended a number of resources, including theses of her recent former students to help me with a structure for my own, and advice on maintaining a healthy balance between thinking and doing.

Plus she used to live in Edinburgh, more or less across the road from where I live now.  Cool.  Thanks Lynda!  You haven't heard the last of me :)

#travel

Once we got to Segovia, we had a guided tour of the ancient Roman architecture, interesting building façades and local legends.  It was a very good tour, but too hot to really focus.  Then they took us to a restaurant for a local speciality.  I was all set to write a whole individual blog post surveying the barbaric nature of human beings, but I didn't do it straight away and now the passion has faded slightly, so I'll leave it at a paragraph.  Some people watched the local 'ceremony' out of morbid curiosity I imagine, but it was the fact that so many people took so much pleasure in the idea of violently hacking up bodies of three-week-old piglets that really bothered me.  Fortunately the surging standing crowd allowed me (and only one other) to inconspicuously sit it out.  The veggie option was tasty, but it was difficult to really enjoy the rest of the evening whilst wondering vaguely about the states of minds of most of the people I was sharing a table with.

Tuesday, July 09, 2013

#SSSW2013: Collaborative ontology engineering and team formation

We were introduced to the various mini-projects on Tuesday morning, and encouraged to form teams with people who weren't from the same university.  I quickly shortlisted the five that sounded most interesting to me, but was disappointed that there weren't any about multimedia.  Because how to evaluate a very subjective system is a potential problem for me, the project proposed by Valentina Presutti was my first choice:

"Serendipity can be defined as the combination of relevance and unexpectedness: an information is considered serendipitous if it is at the same time very relevant and unexpected for a given user and in the context of a given task. In other words, a user would learn new relevant knowledge. To evaluate the performance of a tool (e.g., an exploratory search tool, a recommending system) in terms of its ability to provide users with serendipitous knowledge is a hard task because both relevance and unexpectedness are highly subjective. This miniproject focuses on two main research questions: what is the correct way of designing a user-study for evaluating an exploratory search tool performance in terms of serendipity? Is it possible to build a reusable set of resources (a benchmark) for evaluating ability to produce serendipity, allowing easier evaluation experiments and comparison among different tools?"

Nobody else seemed to be interested though, so I resigned myself to not being able to do it... until I explained the project and why it was interesting, to the best of my ability, to Andy, Oscar and Josef, and they were sold enough to mark it as our first choice.  Thus Team Anaconda Disappointed (a name of significant and mysterious origins) was born, and Project Cusack (because of the movie Serendipity, which nobody got) was underway.

Our first lecture today was from Lynda Hardman, about telling stories with multimedia objects.  It was super relevant to what I'm doing, to the point where I'm surprised I hadn't come across her work already.  My notes are here.  Lynda has done, for example, work with annotation of personal media objects like holiday photos in order to combine them into a media presentation.  She has considered similar things to me, in particular noting that there are many many aspects of data about multimedia - I had assembled my take on this into a Venn diagram for my poster..



One I hadn't considered is annotating an explict message of a piece of media, intended by the creator.  This isn't always relevant - sometimes the consumer's interpretation of the media is more important - and this in itself might be an interesting annotation problem.  Competing perspectives - something an ontology should be able to represent.

I need to check out COMM - Core Ontology for Multimedia.

She has an overview of the canonical processes they have consolidated the process of producing digital content into, and how annotation can be formed around these.

Lynda also told us about Vox Populi and and LinkedTV; practical applications of annotating multimedia.

I made lots and lots of notes.

Natasha Hoy gave us some insights from the biomedical world with regards to ontology development, particularly in relation to the International Classification of Diseases which, when last revised in the 80s, consisted of a lot of paper and a whoever-shouts-the-loudest algorithm for inclusion of terms.  But the next version, currently under creation, is being developed with a version of Web Protege, customised to be friendly for those who don't know or care about ontologies, and is a truly collaborative process (for those allowed to take part) with accountability for all changes.  It's open too though, so even those without modification rights can view and comment on the developments.  My notes are here.

Lunch was for the first time outside, under the shadows of the forest, and for me was a tray of tomatoey vegetables that were delicious but few.  A striking contrast to Monday's lunch.  Everyone else had some meat-potato combination, preceded by a salad with tuna, and followed by a peach.

The hands-on session followed on from Natasha's talk.  We teamed up (temporarily Anaconda Hopeful) and played with Web Protégé.  There were two magazines and two newspapers, each with four departments.  Anaconda Hopeful were randomly designated the Advertising Department of Iberia Travel (a food and travel magazine).  We got stuck in, on paper first to identify some classes and relations that were relevant to us, and then with Web Protégé, along with the other departments of Iberia Travel.  We didn't come into any conflicts, but ended up creating a few classes that we needed, but should really have been the remit of another department (I guess we just got there first).

Then it was announced that Iberia Travel had bought the other magazine (and one of the newspapers had bought the other), and we had to work together to merge ontologies with the other department.  It became apparent that the other magazine had never had an Advertising Department (no wonder they went under!) so we had no-one to attempt to merge ontologies with.  We attempted to sell our expertise to the Advertising Departments of the newspapers, but there were already too many people involved in the heated debate that came out of the ontology merging there, so we couldn't really get involved.

Later we got cracking with our mini-projects.  Valentina showed us aemoo, and the experiments her team had come up with to try to evaluate it.  We sat down by ourselves to brainstorm, describing a lot of concepts for ourselves, breaking down the notion of serendipity, figuring out what might be wrong with existing experiments to 'measure' serendipity, and collating literature in the area.  (Turns out there is a lot, and it's a very interdisciplinary issue; lots to read about from social sciences, anthropology etc, as well as philosophy of science.  In computing, it seems to be primarily discussed within the realms of recommender systems and exploratory search).

Serendipity seems to be mainly described as a combination of unexpectedness and relevance.  Problems include the sheer subjectivity of it.  Some people are going to get excited by all facts they find out, whether they're useful or not.  Some people are going to have hidden, inexplicit or subconscious goals that affect how 'relevant' something is to them.  People describe their different areas of expertise in different ways; some are more humble than others and would not call themselves an expert in a topic, for example.  So whether or not an event can be considered a serendipitous one is a complex question, which must take into account the person's background, goals and existing knowledge, the task they are trying to achieve (or lack thereof, as serendipity is particularly important - in my opinion - in undirected, loosely-motivated activities), the way they are able or encouraged to interact with a system, what they are doing before and after... all these things make up a context for someone's activities, and none of them seem to be particularly measurable.

Dinner was a vegetable and potato (yay!) starter, followed by spaghetti in tomato sauce (fish for everyone else, although Andy got a custom omelette, lah-de-dah).  Also an apple.  We learnt the hard way not to sit at a table directly underneath a light, as the bugs just raiiiin down.

After dinner we crowded around Enrico who had offered to provide advice about PhD-ing.  From this session, I have a signed diagram of the life of a PhD, because he borrowed my notebook to make it.  I tuned in and out of the discussion, and noticed some irregularities between my PhD and what seemed to be 'normal'.  For instance, most people didn't seem to have as much control over their topic, or what they were doing at any given moment in their first year.  I am really, really enjoying my freedom, but in order to justify that I deserve it I need to sort out my lack of direction and focus.  I need to believe in what I'm doing - not be told by someone else - which is one of the main reasons I am doing this particular PhD.  Perhaps I need to ask for more guidance to more quickly reach the necessary conclusions for myself.  (And, of course, perhaps I also need to stop taking big chunks of time out periodically for different reasons; that might speed up the process as well).

Later, the overriding sentiment was that the job of a PhD student was to answer a question, to produce a theory.  Not to create a system or solve a large problem; certainly not to worry about practical, real-world applications of theories.  Well, I've already explained that this is something I can't accept, and I still am not convinced that that is going to impact on my ability to do a PhD.  Theories develop during practice.  Coding and designing, like writing, are part of my thought processes, and I reach realisations or find new questions to ask through hacking and playing and making.  And why would I be hacking and playing and making, if not to try to produce something of real-world value?  If my motivation in making a system is explicitly to come up with new theories, then my approach and outcomes and realisations will be entirely different.  In trying to make something that works for real people, not researchers in a restricted domain or specific context; a clean and sterile laboratory, I figure out different things, that matter.

There was another discussion that I came into a bit late, but it sounded like a very harsh discussion about problems with research in industry (rather than academia) that seemed to be very overstated compared to what I have read and experienced myself.

By the end of the day, it felt like I'd been at Summer School for weeks, and had known everyone forever.

[Notes] Lynda Hardman at #SSSW2013

RELEVANT.

Users (consumers?):

  • Finding content
  • Media types * mostly text at the moment, little integration of different types
  • Specific tasks - not much connection of results with user tasks.

More data than just what you seen in the media (cue my Venn diagram).

Plus, eg. paintings - lots of 'cultural baggage'.

Care more about the story than the media.
Interpretation by end users.  Hopefully message that the author intended.

Meaning of combination of assets.
eg. Exhibition of artists work.

Interacting further with the media.

  • Search - serendipitous or focussed around a theme (or both).  Different search goals.
  • Sharing, passing it on.

(SW and multimedia community need to work together).

-> Raphael Troncy on Friday - attaching semantics to multimedia on the Web.

Need mechanisms:

  • to identify (parts of) media assets.
  • associate metadata with a fragment.
  • agree on meaning of metadata.
  • enable meaningful structures to be composed, identified and annotated.

Workflow for multimedia applications

  • Canonical processes of media production
    • Reduced to the simplest form possible without loss of generality.

Heard of MPEG-7? Don't bother.. very much from a media algorithms perspective.

Applications:

  • Feature extraction.
  • News production.
  • New media art.
    • An interactive exhibit that responded to audience present.
  • Hyper-video.
    • Linked video.
  • Photo book production (CeWe).
    • (Using this example for explaining processes).
  • Ambient multimedia systems with complex sensory networks.

Canonical processes overview...

There's a paper.

CeWe photobook - automatic selection, sorting and ordering of photos.
Context (timestamp, tags) analysis and content (colours, edges) analysis.

Things from these you want to represent your digital system (ie with LOD):

  • Premediate, eg.
    • remember to take your camera on holiday.
    • write scripts, plan shots.
    • place a security camera in the right location.
  • Construct Message (not really in the chain, appears all over the place); what to conveny with media? Intention? eg.
    • show people a great holiday.
    • sell a product.
  • inform/advise.
  • Create (method of creation might be important, so record in metadata), eg.
    • take photos.
    • make video.
  • Annotate, eg.
    • automatic or manual.  Stuff that is embedded by device vendors (but there's so much more...)
    • domain annotations: landscapes/portraits, timestamps, face recognition.
  • Publish, eg.
    • compose images into photobook.
  • Distribute, eg.
    • print photo book and post.
    • cyclic processes online.


COMM - Core Ontology for Multimedia.

Premediate and construct message - human parts, she doesn't expect them to be digitised any time soon.

Using Semantics to create stories with media

Can we link media assets to existing linked data and use this to improve presentation?

How can annotations help?

  • What can be expressed explicitly?
    • Message (somewhere between a html page and poetry).
    • Objects depicted.
    • Domain information. <--- li="">
    • Human communicaiton roles (discourse). <--- li="">

Vox Populi (PhD project)

Traditionally video documentary is a set of shots decided by director/editor.
vs.
Annotating video material and showing what the user asks to see.

interviewwithamerica.com

Annotations for these documentary clips:

  • Rhetorical statement; argumentation model (documentary techniques).
  • Descriptive (which questions asked, interviewee, filmic).
    • Filmic: continuity like camera movements, framing, direction of speaker, lighting, sound - rules that film directors know.
  • Statement encoding (eg. summary what the interviewee said):
    • subject - modifier - object statements.
    • Thesauri for terms.
    • Can make a statement graph, finding which statements contradict and which agree.
    • (He encoded this stuff by hand - automated techniques aren't good enough).
    • Argumentation model - claims, concessions, contradictions, support.


Automatically generated coherant story.

  • Are we more forgiving watching video? (Than reading these statements as text).  Peoples' own interpretations strongly affect understanding of the message.


Vox Populi has (not for human consumption) GUI for querying annotated video content.

User can determine subject and bias of presentation.
Documentary maker can just add in new videos and new annotations to easily generate new sequence options.


User informatio needs - Ana Carina Palumbo

Linked TV.  Enhancing experience of watching TV.  What users need to make decisions / inform opinions.

  • Expert interviews (governance, broadcast).
  • User interviews - what people thought they need (215 ppts).
  • User experiments - what people actually need.

Experiment - oil worth the risk?

  • eg. people wanted factual information from independent sources; what the benefits are; community scale information.


Published at EuroITV.

Conclusions

  • We can give useful annotations to media access, useful at different stages of interactive access (not just search).
  • Clarify intended message. Explicity with annotations.
  • Manual or automatic.
  • Media content and annotations can be passed among systems.
  • No community agreement in how to do this. <--- li="">
  • How to store?

Questions

Hand annotations are error prone - how to validate?
Media stuff - there can be uncertainty, people don't always care.

Motivating researchers to annotate...
Make a game.

Store whole video or segements?
W3C fragment identification standards - timestamps via URLs.

Monday, July 08, 2013

#SSSW2013: Research in theory and practice, and where on earth am I?

The 10th Summer School for Ontology Engineering and the Semantic Web

Sunday

Arriving by train into Cercedilla, north of Madrid, we immediately encountered other confused looking folk with poster tubes.  So we shared taxis (EUR 10) from Cercedilla station to the summer school residence further north, in the forest.

After getting keys for our pleasant, single, en-suite rooms, arrivals congregated in the shade by the building  to introduce ourselves.. Again, and again, and again, as new people continuously arrived over the space of a few hours.

A really broad mix of people are here in terms of nationalities and places and levels of study, but I still haven't quite got used to the fact that answering 'Semantic Web stuff' is not specific enough in this crowd, when someone asks you what your research is about.  Nobody needs convincing that these technologies are useful!

Later we received schedules, maps, ill-fitting t-shirts* and very helpful name badges, and headed for dinner at the bar down the road.

As is traditional when I write about my experiences in new places, I will describe the food every day.  It has become apparent, at this residence at least, that variety of ingredients is not ordinary, so in this respect meals are simple.  Dinner that first night started with a salad (lettuce, olives, tomato, onion, shredded beetroot and a single slice of hard boiled egg; no dressing), followed by - for the majority - slices of meat (beef? Pork? I dunno..) and fries.  Mine was a plate of mushy green vegetables with a little seasoning, that was pretty tasty.  Dessert was a single pear, delivered with ceremony, but otherwise unadorned.  Healthy, at least.

Yet we were all (those I sat with at least) were left feeling a little unsatisfied.

I shared a table with a French, Spanish, Italian and Irish guy.  Conforming appropriately to stereotypes, and setting up reputations for the rest of the week, the French and the Italian shared the bottle of wine on the table; the rest of us went without.

I returned to bed after a couple of hours of socialising and enjoying the cool air in and around the bar.

* For next year, they could ask for t-shirt sizes when they ask for dietary preferences?

Monday

The day started early, and with no hot water or wifi for anyone.  Breakfast was combinations of sweet pastries, coffee, tea, juice and bread.

Punctuated variously by coffee breaks, the learning began in earnest.

During the introduction by Mathieu D'Aquin, I found out that I am one of 53 students selected out of 96 applicants to attend this year's Summer School of the Semantic Web!  I had no idea it was that selective, or that there had been that much competition.

The first keynote was by Frank van Harmelen, about all the Semantic Web questions we couldn't ask ten years ago.

Slides:



Frank started by saying that the early Semantic Web vision has morphed into the more manageable vision of a Web of Data, or a Giant Global Graph, and outlined the principles of the Semantic Web as they appear to stand at present:

1. Give everything a name (entities).
2. Relations form graph between things.
3. Names are addresses on the Web (so we inherit properties of Web like AAA).
4. Add semantics.

Frank pointed out the advantages of the fact the Linked Data crowd, grown naturally and not designed, is now so big we don't know how many triples it contains, nor how fast it is growing.  Companies and organisations (like Google, NXP, BBC, DataGov) are using Semantic Web technologies to achieve their own ends, for a variety of different use cases, without caring much about the Semantic Web, and this is contributing to the growth.

This growth has given rise to a number of research areas that were impossible to realisitically ask questions about ten years ago, including self-organisation, distribution of data, provenance, dynamics and change, errors and noise (how to deal with disagreements).

Frank asserted that rules and structures, algorithms and patterns in data, exist whether we are looking at them or not.  He used the analogy that OWL is our microscope, and it may be the tool that distorts our vision of the information universe rather than properties of what we are looking at (for example, structures in data presenting themselves well in some domains but not others).

He went on to promote the roll of the Informatician to be to test theories, hypothesis and falsify, as scientists rather than engineers.  To discover, rather than build.

I struggle with this view of the world, and feel instinctively that theory and practice are intrinsically linked; one can't exist without the other, not just in the grand scheme of things, but in day to day work and research.  This is one of the main points of contention with my own PhD, and I've no doubt there will be many more blog posts about this issue in the near future as I reconcile my need to create something immediately useful with the necessity of producing a contribution to knowledge at large.

See my raw notes here.

We had an Introduction to Linked Data by Mathieu D'Aquin (raw notes here), followed by a workshop.  We wrote SPARQL queries to populate a pre-written web page with information about Open University courses, sub-courses and locations thereof.

Lunch, similar to the previous night's dinner, was a starter salad, an entire half chicken (or something) plus fries for the carnivores and the most unappealing risotto of my life for (not that I'm ungrateful, but I have never been unable to finish a meal due to boredom before).  I went for a walk with some others to grab some fresh air before the afternoon's work, and missed out on watermelon.

Manfred Hauswirth presented some really exciting stuff about annotating and using streams of data.  Particularly challenging is how to integrate this with static data and make inferences over the lot.  Streams include sensor data, as well as ever-flowing social media streams for example; anything that changes over time.

They've built some systems to process this kind of data, and one of them is available as middleware.

My raw notes are here.

In the afternoon we had a poster session, where all participants pinned up posters about their work, and discussed at length with anyone who was interested.  Here's evidence that I participated.


And here's Paolo's:



I wrote a few notes about things from other peoples' posters that I need to look up.

The main feedback I received was about making sure I focus, narrow down my topic, and concentrate on some evaluatable deliverables that are PhD-worthy.

Questions like (paraphrasing) "why should we care about digital creatives?" threw me, because I thought the obvious answer - that they are people too, Web users, technology users, contributors to culture and an ecosystem of digital content and data - was apparently not enough from an academic standpoint.

I was simultaneously told to focus more, and to explain why the problem I'm trying to solve is applicable to all domains, not just digital creatives.  But some of the problems I'm looking at have been (or are being) solved in other domains (like e-health, biological research, education) and the reason what I'm doing is interesting is because none of these solutions quite work for digital creatives, and I want to find solutions that do, and try to figure out why.

I'm still stuck in some sort of struggle between theory and practice; thinking and doing.  And the long-standing problem of how to decide which doing actually worked.

I've started scribbling notes about the narrowing down problem.  I'll need to have this figured out before my first year review in August anyway, so stay tuned for another post all about it.

Then I sneaked off for a nap.

Dinner at the bar again; the usual salad, plus some eggy fish thing for most.  I got a plate of artichoke.  Artichoke is great, I love it, and I'm all for simple meals.  But I remain unconvinced that a plate of only artichoke constitutes an acceptable level of effort on the part of caterers.  And the sheer quantity made it start to taste a bit funny after a while.  But not to worry; we rounded off with a solitary peach apiece.

Further socialising, and appreciation of the night sky, before returning to bed write blog posts.

I'm super excited and inspired by the talks, work I've heard about so far, and the atomsphere of the place.  I'm excited to learn a helluva lot, and remind myself that I'm not facing impossible problems, and am not facing many problems alone.  I remember that I am instinctively passionate about the Web and the possibilities it holds (and indeed has already realised) for the empowerment of individuals.  I remember how lucky I am to be able to sustain myself through studying something I love so much, and to have the potential to make a change, and through my work maybe even facilitate others to be able to make a living doing what they love, as well.

Monday, May 13, 2013

Week in review: annotating multimedia content

6th - 12th May

Discovering lots of things to write about semantically annotating multimedia content.  I decided there are three main ways to do this:
  • Technical / objective / statistical data: eg. media type; shutterspeed; framerate; duration; resolution; date created; number of times viewed at a particular source; number of time shared ...
  • Bibliographic: creators and contributors and their roles; methods/location of publication; methods/locations of creation ...
  • Content*: fictional characters; locations; camera movements; scene transitions; colours ...
These categories overlap somewhat really, and when I get round to it I'll type my Venn diagram up.

Technical is easy, and a lot of that is automatically captured by hardware or software used to produce and edit works.  It's also relatively easy to extract automatically.  Standards like MPEG-7 and MPEG-21 take care of formalising it, and Jane Hunter turned these standards into semantic ontologies in 2002.

Bibliographic can largely - but not entirely - be covered by vocabularies that have been around forever like Dublin Core, FOAF and various library-originated things.  Things that might be missing (or I just haven't found them yet) are associating roles with tasks involved in digital media production, since pieces are often a collaborative effort.   has some idea of participants and roles, but the purpose of  is digital rights management stuff, so it's more concerned with the distribution change, I think, than granular production of content.  I haven't read much about it yet.

Content is more interesting, and potentially more useful for ordinary human beings.  Imagine querying IMDB for "that film where John Goodman arrests an animated talking moose on a US highway" instead of scouring John Goodman's filmography or googling for pictures of animated meese until you see the right one.  Annotating characters, objects and events, and stringing them onto a timeline is possible with OntoMedia.  It's very focussed around narratives, which is great, but doesn't link back to technical so much.  So if you did find the answer to that query, it wouldn't be able to serve up the timestamp of that particular scene.

On top of what I've looked at already, I still have this list to (re)investigate: 
A thing I want to do is annotate some amateur content with OntoMedia and with ABC to see how they compare.  Maybe I'll do asdfmovie, because it has associated comics, and multiple people participating in production.  Then I'll do something live action as well, because I can't base all my research on non-sequitur lolrandom stick figure cartoons.

Now, back to work..

* I want a better name for this, since I'm referring to everything as 'content' anyway.  So some better way of saying 'content of content'.

Thursday, April 11, 2013

2nd UK Ontology Networks Workshop

The UK Ontology Networks Workshop took place over one day in the Informatics Forum.

There was a mix of people there; some talks were way over my head and very technical, and some talks were by people who confessed they had had to look up "ontology" that morning.  And things in between.

Lazy writeup, but following are notes as I scribbled them:



John Callahan

US navy research.
Focused information integration.
Human intervention to keep predictive part on track. Tweaking.

Alan Bundy

Interaction of representation and reasoning.
Changing world so agents must evolve. How to automate? What would trigger a need for change:
Inconsistency
Incompleteness
Inefficiency
how to diagnose which?
Interested in language and perception change.
Unsorted first order logic algorithm called Reformation. Based on standard unification algorithm.
Allows blocking and unblocking unification.


Phil Barker

Schema.org
Cetis (JISC funded)
learning resource metadata initiative.
Big names behind schema.org.
= ontology + syntax
Big and growing ontology.
Dumbed down for people.
LRMI adds to it. W3C go through it. It's creeping, how much do the big names actually care about stuff that's added?
don't know how Google uses it.
People should consider using it for more sophisticated search and disambiguation.

Gill Hamilton

Doing more with library metadata. Learnt from OKFN. Had to convince people in charge.
Dublin core, didn't like; not specific enough. Instead RDF > OWL. "We know best how to structure our data"

Hardest was convincing marketing people that there was no commercial value. Metadata is advert to actual resource.

Enrico Motta

Traditionally top down approach. So now so many people interacting with semantic structures, so should involve users.
Recognise there isn't a unique or best way of doing things.
Initial study included modeling task with binary relations.

Patterns that are more or less intuitive. 4D least, 3D+1 most.
N-ary most widely used by experts.

Relationship between reasoning power and intuitiveness of writing? More creativity needed for simpler ones. (Not really sure what he's saying)

Email him for copy of study.

Chris Mellish

Ontology authoring is hard. Better ways to do it.

Controlled language input (mature tech); responsive reasoning (also mature, information as you're editing); understanding the process (beginning to understand more).

Hypotheses:
users don't know what they're doing. What if questions.  Many answers, what is relevant? Depends on context.

Authoring as dialogue.
Todo list.

Useable in the same ways as protégé.

Peter Winstanley

UN classification schemes.
Various vocabularies.
Allow development of cross mapping between government administrations.

Mostly internal currently. Moves to bring externalizing data into the 21st century.

Peter Murray-Rust

Fight for your Ontologies.
Ontologies in physical sciences. Chemists don't want ontologies. They'll sue you.
Crystallography uses 'dictionary'. Written in CIF. 20 years to build CIF.

Compare physical sciences to government.

Every program author writes dictionaries that work for them. When different parties agree, promote to communal dictionary. Provide conventions to help disagreements.

Show a company can do it as opposed to a rabbiting academic ..

Jeff Pan

Tractable ontological stream reasoning.
Need to be more efficient, scaleable, as things change. Inputs from web.

Dealing with complexities: approximate owl2.
Dealing with frequent updates: to-add stream and to-do delete stream. Truth maintenance. Evaluation criteria.

Trowl.EU can use with protégé, also supports jena.

Edoardo Pignotti

Semantic web tech to support Interdisciplinary research.
ourSpaces VRE
Provenance crucial.
OPM prov ontology.

Deployed since 2009, 180 users. Comprehensive ontologies but people unwilling to provide metadata.
paper! Edwards et al. ourSpaces.

Tom Grahame (BBC) @tfgrahame

Content arrangement on BBC sport by tagging, automatic to free up editors to write.
LD API so systems don't need to know about each other.
Growing from simple rdfxml to more complex ontology.
Can ask much more general and much more detailed questions about sport.

Mapping incoming data is outsourced.
Lots of errors, sometimes system alerts, sometimes manual.

Working on opening the data. Maybe a dump, but licensing issues.

Ewan Klein

Mining old texts for commodities, adding place and time and putting in structured database.
Transcriptions of customs import records.

Skos for synonyms.
Dbp concepts.

Why? Want to query.
Visualisations.

Tools? Python script.

Janice Watson

Harnessing clinical terminologies and classifications for healthcare improvements.

Bob Barr

Geographical addressing.
Addressing and address geocoding is important and broad. Not always postal, but this not addressed (punlol) in ontologies.
Different contexts change meaning of address (for delivering, you only care about postbox; property sale whole building).
Loads of things to address. Loads of reasons why.
Work held up as national address file is owned by royal mail and might be sold!

Fiona McNeill

Run time extraction of data. Failure driven. Looking at extraction of specific information.
Emergency response. Lots of data, timely sharing of data required.
From domestic level to humanitarian disasters.
How can it be automated?
Multilayered incompatibility.
Format
Terminology
Structure
...

Richard Gunn

Towards an intelligent information industry.

Elena Simperl (Soton, sociam)

Crowdsourcing ontology engineering.

CSrc: Brabham 2008.

Distribute task into smaller atomic units.

Humans validating results that are automatically detected as not accurate.
What are the costs? What resources?

Games with a purpose. Like quizzes.
Micropayments or vouchers.
MTurk. CrowdFlower.
Paper about useage of microtask crowdsourcing.  ISWC 2012.

Claudia Paglieri

Ontologies in ehealth.

Enrico Motta - Rexplore
Klink algorithm mines relations between research topics.
Use this!  Nope, it's not public.   Uees MS Academic research.

Peter Murray-Rust

Content mining expands regular text mining.
Focus on academic stuff.
Chemical Tagger. Takes chemistry jargon and annotated it, knows actions, conditions, molecules etc.. NLP. Uses ontologies and contributes to ontologies.
In chemistry,  no need to put everything in rdf because there are already lots of formalisms.
Proper cool PDF to sensible format conversion. Amy the kangaroo. Looking for collaborators.

Yuan Ren

Ontology authoring in whatif project.

Reasoning with protégé and trowl .

Tractable reasoning. Trowl v fast.


Notes from conversations / breakout discussions:

BBC use owlm triplestore  .
Store all their datasets in svn. But they have reads and writes to the live triplestore all the time.

Lots of people saying minimise owl use because of unpredictable output.

Versioning ontologies (available in owl2) in case third parties change stuff you use. You're dependent on their software engineering practices. Only good if they're ahead of the game.

IRIs, Arabic characters in ontologies!
Semantic heavy, maybe make a decision to abstract away to ids and make heavier use of labels.

Difference between importing and using someone else's.

There's no (practically useful) software that lets you reason over stuff you haven't imported? (over HTTP?)

Build ontology from reality (data), don't start with no data.

Lode.

Problems with dbpedia URIs changing or disappearing.

Hard to visualize massive graphs. Relational, tabular much easier to understand.

Friday, March 15, 2013

Notes about Meervisage - A Community Based Annotation Tool (for the Semantic Web)

Rowe, M. (2007)  Meervisage - A Community Based Annotation Tool. ‘Towards a Social Science of Web 2.0’ Conference at the University of York 5-6th September, 2007.

How SW can benefit from incorporation with existing 'Social Web'.
"...collaborative generation of metadata... using social networks as a user base..."

Uses fb groups created for sharing and organisation of research.  Suggests posting links to useful resources is comparable to annotating the resource.  Comments are more metadata.

Points out usual stuff of actually generating semantic data being a problem for SW.

System requirements:
  • Annotations must be shared in a community.
  • Annotations can be reviewed and edited (/audited) (by group)
  • Collaborative
  • Central repo.
  • Annotations contain semantic metadata.
  • Content of resource annotated, not URL.
  • Communication layer that doesn't interrupt annotation (uses external services).
Review of existing systems:
  • Annotea [9] [13]
    • J. Kahan, M.R. Koivunen, E. Prud Hommeaux, R.R. Swick. Annotea: an open RDF infrastructure for shared Web annotations. Computer Networks. 2002.
    • M Koivunen. Annotea and Semantic Web Supported Collaboration. Proc. Of 
    • the ESWC2005 Conference, 2005.
    • No communication layer (but has discussion threads, wat?). 
    • Can only be edited  by author, but can be reviewed by others.
    • Can be local, private or shared.  RDF.
  • Piggy Bank [10]
    • D Huynh, S Mazzocchi, D Karger. Piggy Bank: Experience the Semantic Web Inside Your Web Browser. Springer-Verlag GmbH. 2005. 
    • RDF. 
    • Auto and manual.  Bundled with scrapers; if they fail, manual.  Only of one type.
    • Share group or global, or save to local 'semantic bank' <-- find="" is="" out="" this="" what="">
    • Reviewed by all, edited by author.
    • Community of users, but no SNS integration.
  • KIM [14]
    • A Kiryakov, B Popov, D Ognyanoff, D Manov, A Kirilov, M Goranov. Semantic Annotation, Indexing and Retrieval. Journal of Web Semantics, Springer. 2004.
    • Automatic named entity recognition.  
    • Links to knowledgebase with ontology.
    • Creates new URIs for new entities or link swith entities it already knows about.
    • Global sharing.
    • Can be deleted but not edited.
    • No social involvement.
  • Magpie [11]
    • J Domingue, M Dzbor and E Motta. Semantic Layering with Magpie. Handbook on Ontologies. 2004.
    • Auto annotate webpage.
    • Similar to KIM, but does not hyperlink to knowledgebase; instead each item gets context menu (right click) with services depending on entity.
    • 'Multi-dimensional approach'. Uses ontology to trigger other services depending on concept.
    • Plugin for IE.  
    • Simply looks for entities that are in ontology (Dzbor 2004).
[1] Using existing information to derive semantics from folksonomies (delicious):
X Wu, L Zhang, Y Yu. Exploring social annotations for the Semantic Web. Proceedings of the 15th international conference on the World Wide Web, 2006.

[15] Social bookmarking tools and how semantic info aids resource discovery.  Probabalistic model of how resources are annotated:

A Plangprasopchok, K Lerman. Exploiting Social Annotation for Automatic Resource Discovery. Eprint arXiv, 2007.

[16] Distributed nature of folksonomies.  Improve search mechanisms.  Tags not great:

S Choy, A. Lui. Web Information Retrieval in Collaborative Tagging Systems. Proceedings of International Conference on Web Intelligence, 2006.
        (vs.)
[17] Rigid taxonomies not great:

C Shirky. Ontology is Overrated: Categories, Links, and Tags. Clay Shirky’s Writings About the Internet, 2005.

[18] Methodology for easier browsing of large scale social annotations:

Z Xu, Y Fu, J Mao, D Su. Towards the semantic web: Collaborative tag suggestions. Collaborative Web Tagging Workshop at WWW2006, 2006.

All use one annotation per resource, not annotation of content within, so only one lot of metadata about a page.

Meervisage

"To aid the process of collaborative annotation of web documents"

Allows sharing of annotations between subset of SNS users (eg. fb group).

Management of users and groups offloaded to third party.

Stored in central annotation store.

Annotations contain author, SNS, folksonomies and date.  Made from content within.

Meerkat is "responsible for generating semantic metadata by annotating external web resources."  Meervisage for management via social network.

Meerkat allows a user to edit another user's annotations if they are members of the same group on facebook.

Popularity rating of resources rises with fb discussion.
Meerkat informs browser users if they come across a resource that has been heavily discussed on fb, and by which group etc.

Meervisage also provides RSS feed.

Evaluate by comparing precision and recall metrics of annotations by one user in an allotted time, and those by a group.
-> Don't know how this helps to assess quality of annotations; maybe I'm dumb?  Find out.

Limited to private, says it's like that's a good think :s
Oh, because public access would be "laborious and resource intensive".

Annotations rated on usefulness and weighted.

[20] Attempt to describe folksonomies as part of formal ontology.  Meervisage doesn't; limited to users' viewpoint:

S Angeletou, M Sabou, L Specia, E Motta. Bridging the Gap Between Folksonomies and the Semantic Web: An Experience Report. Workshop: Bridging the Gap between Semantic Web and Web 2.0, European Semantic Web Conference, 2007.

[9] + [13] are most similar.  Have groups, but groups aren't already established networks.

Future work
Annotating multimedia.
Matching assigned tags with ontology terms mined from Web.
[19] Desktop app for annotating text with ontology:

A Chakravarthy, F Ciravegna, V Lanfranchi. AKTiveMedia: Cross-media Document Annotation and Enrichment. Poster Proceedings of the Fifteenth International Semantic Web Conference, 2006.


Sunday, March 03, 2013

Week in review: Reading and catching up

25th Feb - 3rd March

I spent most of this week in a remote village by the sea in the Scottish Highlands.

When I came back I read and made notes about the Semantic Web and social machines by TBL; and typed up some notes from a while ago about OntoMedia and the Semantic Web and communities by K. Faith Lawrence.

I also translated all of my written meeting notes into Evernote, which promptly glitched out and doubled the amount of typing I had to do.  (I considered switching back to Google Docs, but I need labels).  I sure love technology.

I refamiliarised myself with the structure OWL.  Awesome diagrams here.

I did some more thinking about how I need to work with amateur content creators to make an ontology that fits their workflow.  I should have finished the planning stage of this ages ago, but.. I blame the ILWhack.

I keep wondering about the best way to have a system of consistent URIs across a network where the information can move from server to server on the whim of a user.  During this wonderment I discovered that purl.org's login system is broken.  I joined the mailing list, and people complain about it and have it re-fixed fairly regularly, so I'll just wait..


Notes on SW and Communities

Lawrence, K.F., schraefel, m.c.: Bringing communities to the semantic web and the semantic
web to communities. In: Proceedings of WWW2006. (2006)



Research into SW communities:

  • Communities of practice
  • Social networks, eg. FOAF

Compare with other definitions of communities outside of SW.
Concept: Internet Based Community Network, has properties of COP and SN.
Case study: Amateur Fiction Online.

Early community definitions, Howard Rheingold: "..webs of personal relationships in cyberspace."

1996 CSCW Conference defined prototypical attributes of communities (Whittacker):

  • Shared goal / interest / need
  • Repeated active participation, emotional ties and shared activities.
  • Shared resources and access policies
  • Information, support and services reciprocated between members ( overlap with ^ ?)
  • Shared context (culture, language)
  • Can be applied to virtual and offline communities.

-> More attributes = clearer example of community.

Preece:

  • Social interaction
  • Shared purpose
  • Common set of expected behaviours
  • Computer system that facilitates and mediates communication


^^ Things in common.  Whittacker's is more inclusive/broad.

So for a SW SN:

  • Accessible via browser
  • Explicit links between users
  • System supports creation of these links
  • Links are visible and browseable

COP or SN may describe a community, not necessarily.  IBCN will do, and could be a COP or SN too.

Problem of Amateur Fic. is fluctuation of archive.  Personal sites go down etc.  How to find a story you remember a bit of?

IBCN is also combination of WBSN and virtual community.
Lack of incentive to use FOAF (eg. on LiveJournal etc) (Plus ignorance).
Doesn't offer anything they don't already have.
They don't use much metadata, just tons of human-readable stuff.

SW would allow:

  • "better integration of distributed systems"
  • "improved searching and filtering"
  • "more personalised services"
    •   experienced users
      • expand options
      •   new ways to interact
    •  new users
      •   ease introduction re: unwritten rules, expectations, terminologies


FOP extension to FOAF for anonymous identities.
('Fan Online Persona' - why not just 'Online Persona'?)
Consistency likely in community-based system because of advantages of reputation etc.  Identity cost.
Shared set of behaviour values, or risk losing rep.
Reputation gained by taking part.  (definitive part of community).
Additionally by creating works.

foaf:document and foaf:groups allow users to give details about their own creations and review work of others.

OntoMedia to describe content complements FOP.
Options in FOP gathered from study of metadata of works in mailing lists, websites and groups.
Recommender system -> notification system.

Allow SNS of writers to be studied at friend level and collaboration level.

Application to allow users to create FOP under development...

Notes about Annotating Multimedia with OntoMediaLawrence, K.F., schraefel, m.c.: Bringing communities to the semantic web and the semantic web to communities. In: Proceedings of WWW2006. (2006)

Michael O. Jewell, K. Faith Lawrence, Adam Prugel-Bennett, and m. c. schraefel (200?) Annotation of Multimedia Using OntoMedia

Check out a bit of discussion about this paper on Ontologies with a View.

OntoMedia for representing "diverse range of media".

Others for media:

  • CIDOC Conceptual Reference Model (museums)
  • ABC Ontology (multimedia in libraries and digital archives)
  • Functional Requirements for Bibliographic Records (attribute and relationshops for task performed when consulting bibliographic records)
  • FictionFinder - FRBR to Online Computer Library Centre
  • WorldCat db - Metadata about characters and fictional places
    • Describe contents of films & comics etc for tracking things down to share you forgot?

None quite did what was needed.
So created to map to current models but specifically describe media content.
Hierarchical approach.

1. Overview
Entity / Event system
Entity: object, concept
Event: interaction between one or more entities
0 or more Entities are modified OR new Entity created
Entities not destroyed, but may have not-exists attribute


Decompose to sub-ontologies.

ontomedia.ecs.soton.ac.uk/ontologies

Mediate? Graphical interface.  interaction.ecs.soton.ac.uk/ir/projects/ontomedia/ontomedia

Sub-ontologies

  • Core
    • Expression (primarily elements and subclasses; Entity, Event)
    • Media (binding between media and Expression objects)
    • Space (extension of Signage Location Ontology, buildings, and regions of structures)
  • Extensions (more detailed subclasses to Core)
    • Being (people)
    • Trait (attributes of Entities)
  • Events (extends Core->Event)
    • Action
    • Gain
    • Loss
    • Travel
    • & properties thereof
  • Fiction
    • Character (on Being)
    • spoiler info. and accuracy
  • Media
    • More detailed than the one in Core; includes audio, image, photo, text and video subclasses
  • Misc - classes used by any or all of other classes, eg. colour, geometry

Specified in OWL
Developed in Protege and SWOOP.

2. Case Study
Scene from Total Recall annotated.  Represent script and characters, and characters from related book, and links between two forms.

Screenplay annotation - SiX - Screenplays in XML.
Wraps around existing content.
Transition (cuts, fades, blackouts), location, dialogue and direction (action taking place in the script)

SiX allows for DC, for creators, date, descr, title.
Custom XSL to conform marked up scripts to Oscar requirements for readability.

Script Item extends Media Item to link script representation to OntoMedia.
Use has-expression to tie to OntoMedia:Expression.

Describing places and access etc, like lift, like IF.
- For describing character continuity, eg 'can character really see x' etc.

Must describe events that don't occur.  Characters want to occur, etc.  Multiple timelines, dreams.

Declare events.
Create timeline and add occurances.
- events can be reused, and coincide.

3. Testing / querying
Imported into Sesame triplestore.
RDQL queries (subset of SPARQL, simpler, only ever has 1 graph pattern, doesn't use RDF data typing)

4. Conclusion
No examples of pictures - how to annotate comics?

Combining OM with other apps.
Stuff integrated into Mediate.

Thursday, February 14, 2013

Computer Mediated Social Sense-Making

I was fortunate enough to attend the Computer Mediated Social Sense-Making workshop, conveniently situated on the ground floor of the building I work in, on the 14th of February.

Whilst more technical than the Digital Methods conference I went to in December, the talks and panel sessions served to build upon things I started to think about then.  Namely, beginning to situate my research interests amongst many concepts from the currently quite alien fields of sociology and anthropology.

The talks were varied, and key themes that emerged were the collection/use of data for social improvement (health and wellbeing, teaching and learning, disaster recovery), and the importance of context in making collected data genuinely useful.  A notable challenge is that one piece of data might have a thousand different contexts from the perspectives of a thousand different human beings.  So how to communicate these variations to software that processes this data, and perhaps makes decisions using it?

Perhaps not to worry too much about that at all.  Process things locally instead of globally, using local contexts and understandings, but make sure everything is annotated such that information can still be exchanged across the whole network, and differences in understanding can be accounted for or reasoned out if a need occurs.

For the record, I'm looking at how Semantic Web technologies could be used to better connect human and machine in the context of amateur digital content creation (movies, comics, music, art), including how semantically annotating creative (often collaborative) processes as well as the end products of these processes and the engagement of an audience with these products, could improve the overall experience of creating content (along a number of dimensions).  A massive part of this will be creating tools that actually collect the necessary data from users.  Ultimately, these tools will need to be invisible, ie. easily integrated into existing online routines, with no effort required to use them for the non-technically minded so that a network effect can take place.

Incentives for crowdsourcing came up during CMSSM, and someone pointed out that by gamifying data collection for research projects, incentives become the same as ones offered by gambling companies; something competitive and potentially addictive.  I think things like global systems of reputation and trust are useful on a network where people are to share data about their own work (or opinions of the work of others) and may be nurturing a desire for popularity or exposure on the network (a network where the people are central, because the data could not exist without them, but where the users and the data are simultaneously co-dependant).

Anyway, I'm still brainstorming.

Monday, February 04, 2013

Week in review: more brainstorming


28th January - 3rd February

I read Annotation of Multimedia using OntoMedia by K. Faith Lawrence et al., and we discussed it during Ontologies with a View.  OntoMedia might well prove useful in that describing the content of digital media can improve searching, sorting and sharing.

I started reading a couple of other papers by Faith, but haven't finished them yet, so expect summaries in the near future.

I brainstormed about decentralised networks, with thinking of ways of individuals sharing linked data about themselves and their projects without surrendering all that data to a server in mind.

Monday, January 28, 2013

Week in review: brainstorm

21st - 27th January

Brainstormed about how to work with amateur digital media creators, mostly with the aim of starting to put together a vocabulary for representing various digital media creation processes, collaboration dynamics and audience engagement.  More coherent thoughts on that next week, I should think.

Lightning-talked about the Berlin Open Data Dialogue at an OKFN meetup in the National Library of Scotland on Thursday 24th.

Talked about ontologies for sensor data at Ontologies with a View.

Monday, November 26, 2012

Notes about Semantic Web tools for online communities


K. Faith Lawrence & Dr. Monica Schrafel (2007)  Amateur Fiction Online - The Web of Community Trust: A Case Study in Community Focused Design for the Semantic Web. Intelligence, Agents, Multimedia (IAM) Group, School of Electronics and Computer Science, University of Southampton.

NB. Need to read her full thesis, of the same name.  Will probably clear up some of the questions I scribbled whilst reading the paper.

Finding out if Semantic Web tools can be brought to hobbyist groups on the Web.
  • Uses the online fiction community; suggests they could benefit from:
    • improved searching
    • improved meta data
    • automatic recommendations
    • trust webs
    • personalisation.
  • A HCI project, so usability tests and comparisons with current systems are key.
Related work
  • Community centered design
    • to determine user needs - through continual interactions and user studies.
    • to consider how reader-facing apps present themselves and particular community
      • responsibility of being a portal - need clear affordences and points of failure.
  • Trust and Semantic Communities
    • The semantic web “provides a common framework that allows data to be shared and reused across application, enterprise, and community boundaries.” - Tim Berners-Lee, James Hendler, and Ora Lassila. The semantic web. Scientific American, May 2001.
    • Do we trust:
      • metadata
      • data
      • mechanism by which data is returned
      • person requesting data?
    • Many definitions of trust.
    • Jennifer Golbeck's trust onotology to go with FOAF (Jennifer Golbeck, Bijan Parsia, and James Hendler. Trust networks on the semantic web. In Proceedings of Cooperative Intelligent Agents 2003, 2003.): ratings of 1 - 9 for trust of associates.  Extended for FicNet (this paper) 
    • Here, trust: "the expectations that arise that an individual will not act in a way that is detrimental to another individual or community."
      • I might need to expand that for my stuff... maybe... maybe this will do.
Case study
  • Community predates the Internet. (duh)
  • Necessary to get opinions from people outside of the amateur writing community, because it is broad. Such as parents/guardians of members.
  • Questionnaire
    • General information
    • Reading habits
    • Community involvement
    • Access and distribution of materials
  • Questionnaire distributed by:
    • requests to archives to pass on to members
    • LiveJournal
    • emails to specific interested parties
    • mailing lists / bulletin boards of relevant special interest groups.
  • In two weeks, 1116 responses, from 30 countries.
  • Used to inform ontology design for FicNet, and OntoMedia.
Fan Online Persona (FOP)
  • Extension of FOAF, tailored for needs of online readers and writers.
  • foaf:person -> fop:persona
  • Separates environments for on and offline
  • fop:NomDe - context for name
  • Illusion of anonymity is fundamental to fanfic community (who are a large part of online amateur writers)
  • Most authors have one or more pseudonyms.
  • 80% said email address is the most personal information they should be asked for.  ('of the 80%, 15% said no personal information should be requested from anyone - does this make sense? Do they mean other than email address?  but that's the 80... If they're in the 80, they can't be part of that 15...)
  • Privacy is the main thing holding back FOAF (Joseph Smarr. Technical and privacy challenges for integrating foaf into existing applications. Presented at 1st Workshop on Friend of a Friend, Social Networking and the Semantic Web, September 2004.)
  • Personas aren't meaningless, because people become very attached to them, and only create new ones for specific reasons (says who? No citation..)
  • Expands foaf:document and foaf:groups
  • Creation, exchange and review of works is the point of these communities.
  • FOP dismisses FOAF info like work and school as irrelevant or potentially dangerous.
    • [Me]  I think the on/offline divide won't be so extreme for many amateur film makers (another story for consumers) because often their faces are in their movies... Also anecdotal evidence from my own experiences that I'm open to having proven to be a minority.  Actors vs characters is an interesting distinction too.  One amateur film maker can have many personas, even across one channel of output.
  • Options for FOP determined through long term study of metadata commonly attached to works. (Something I can do, too).
Trust
  • FilmTrust by Golbeck (just joined, it was closed last time I looked).
    • Could be prettier... but 2169 members!
    • Visualisations of the network - I need to get good at this.
  • Reader has to trust info from author falls within a certain level of accuracy.  In amateur writing, it is more acceptable to be over cautious than lenient.  Differing standards of acceptable content.
    • Less trust is lost if a story is underrated than overrated (resulting in disappointment)
    • A minority mislabeling work has a big effect on reputation of an archive/community (HelpingHands community members. A place to pitch in and help - a website creation resource and project. LiveJournal Community, 2005.)
  • Writer has to trust reader to make the right decision.
  • FicNet has a more specialised trust system than Golbeck's.
    • Largest contention in this field is adult material and younger readers (debated because this contrasts with IRL - no restricted areas in book stores, or suitability rating scheme for books).
    • Initially focussed on age.
    • Personas could vouch for each other.  Creating fake personae to validate another wasn't worth payoff?  Non malicious statements of distrust?
    • How to integrate trust and distrust webs?
Future
  • Ontologies developed, ready to be used by applications!
    • Ontologies will be continually refined.
    • Now designing applications.
      • Using info already gathered via quesitonnaire, re: UI, functionality.
  • Integrate with OntoMedia to describe works, and link works with people.

=> How did they get to talk to the parents of younger users?  Did they ask the members to put them in touch?  That doesn't seem like a realistic expectation to have, to me..

Week in review: Pancakes & Project management

19th November - 25th November

I read two papers about ontology development methodologies.

I read two articles by Bennett Haselton about decentralized social networking, which happened to pretty much sum up and beautifully articulate everything about that that has been floating in my subconscious for a couple of weeks.  I saw links to them in the latest Circumventor email, which I've been subscribed to since High School for bypassing the internal blacklist, and remain subscribed to because the jokes at the end are always laugh-out-loud funny.


I attended an all day course entitled 'practical project management for research students'.

  • It was attended by a diverse bunch of seemingly really lovely people.
  • The two ladies running it, from the IT Project Management department in the University, were lovely too.
  • The stuff covered was all obvious, common sense stuff (and pleasantly the organisers didn't try to claim otherwise) - but sometimes it's helpful to have it all written down and waved in your face.  And structured, in particular.  Made me actually focus on thinking about organising my project.  The main thing I hadn't much considered, even subconsciously, was formally identifying stakeholders for a project and their relative interest/power in the project.
  • There are a bunch of tools at projects.ed.ac.uk to aid in project management.
  • It prompted me to do these week-in-review posts, as I realised I haven't been recording properly everything I've been doing (an overview of my time goes on my calendar, but no detail).
  • The sandwiches weren't great, but fortunately when I got back to the Forum there were massive slabs of chocolate cake left over from some event.  I love the Forum.
I booked a place at the 1st International Open Data Dialogue in Berlin, and necessary flights.
  • Despite the short notice, it worked out logistically because I need to be in London on the 7th anyway, so I can simply go to London on the 4th, fly to Berlin from there for the 5th and 6th, and back to London for the 7th.
  • I'm particularly looking forward to "Open Statecraft: Openness as a Means (not an End)" by Philipp Müller, "The Open Data Movement vs. Business Models - is this a Contradiction?" by Dr. Peter A. Hecker, "Linked Open Data @ W3C-Vocabularies, Working Groups, Usage Scenarios" by Prof. Felix Sasaki, "The potential of Open Data for improving urban sustainability" by Dr. Marianne Linde and "Towards Trustworthiness: Establishing Transparency with Open Information Flows" by Dr. Edzard Höfig.
  • I'm also looking forward being in Berlin again, even if it is just for one evening, and I'll probably be too exhausted to appreciate it.
Ontologies with a View took place at a different place and time to usual.
I started preparing for presenting at Digital Methods as a Mainstream Methodology in London in a couple of weeks.
  • I scribbled lots of notes.
  • I skimmed a few papers by organisers/speakers but didn't read any in detail yet.  Mostly stuff about analysing data gathered from comments, tweets etc. 
  • There will be more about both of those things next week, I imagine.
I made a plan for the two weeks following the 20th.
  • It mostly consists of finishing my Digital Methods preparation.  I have a lot of non-PhD related things to do as well, plus lots of travelling.  Also graduation from my MSc, and subsequent parental visitation will get in the way.

Wednesday, November 21, 2012

Notes about ontology creation methodologies (2 papers)

Yesterday I unexpectedly read two whole papers about ontology development methodologies.  They were open in tabs I don't remember opening, but presumably did so during our weekly Ontologies With A View meeting last Friday.  There are still a bunch more tabs open with papers or articles about the same thing, so maybe I'll read those later..

The notes are here more or less as I scribbled them down whilst reading, and I haven't expanded with any analysis or discussion as of yet.

Notes in purple are things I intend/need to investigate further; colour-coding is just for me, really.


Jean Vincent Fonou-Dombeu & Magda Husiman (2011)  Combining Ontology Development Methodologies and Semantic Web Platforms for E-government Domain Ontology Development.  International Journal of Web & Semantic Technology (IJWesT) Vol.2, No.2, April 2011 

  • Start by describing ontology in a human-readable way, then turn to RDF (etc) to be machine readable.
  • Says there's not sufficient practical research around existing technologies or ontology development guidelines that would allow non-experts in e-government domain to make ontologies.
  • Uses framework from Uschold & King (see later in this post) to describe ontology - technique used here should be platform independent.
  • Then uses UML to semi-formally represent ontology
  • Uses Protege and Jena to convert to OWL and RDF
  • Paper's goal is to produce guidelines for e-government developers to create semantic content AND strengthen adoption of Semantic Web technologies in governments (particularly developing countries).
  • Outlines RDF, OWL, Protege and Jena (described as leading platforms; mentions other platforms: WebODE, OntoEdit, KAON1, Sesame).
  • Very critical of other literature; either ontologies have been produced but no practical information given; they've been developed with proprietary platforms; or they're only conceptual and don't say how they could actually be constructed with existing technologies.  other studies have not focused on a methodological approach, which means nothing is easily repeatable.
  • Detailed comparative studies of methodologies in:
    • M. Fernandez-Lopez, “Overview of Methodologies for Building Ontologies, ” In Proceedings of the IJCAI-99 workshop on Ontologies and Problem-Solving Methods (KRR5), Stockholm, Sweden, 2 August, 1999. 
    • H. Beck and H.S Pinto, “Overview of Approach, Methodologies, Standards, and Tools for Ontologies,” Agricultural Ontology Service (UNFAO), 2003. 
    • C. Calero, F. Ruiz and M. Piattini, “Ontologies for Software Engineering and Software Technology, ” Calero.Ruiz.Piattini (Eds.), Springer-Verlag Berlin Heidelberg, 2006.
  • Case study: Ontology for monitoring development projects in developing countries (OntoDPM)
    1. Create with Protege:
      • class heirarchies
      • slots
      • domain and range of slots
      • Based on the UML
      • Saved as OWL
    2. Then put content in RDF with Jena

Mike Uschold and Martin King (1995*) Towards a Methodology for Building Ontologies. Workshop on Basic Ontological Issues in Knowledge Sharing, IJCAI-95.
  • Steps:
    1. identify purpose
    2. build ontology
      • capture
      • coding
      • integrating existing ontologies
    3. evaluation
    4. documentation
  • Purpose
    • Many ontologies are intended for reuse
    • Should survey purposes to clarify options for future projects
  • Building
    • Capture
      • Identify key concepts and relationships in domain of interest
      • Produce unambiguous text definitions for these
      • Identify terms for these
      • Agree on all of the above
    • Coding
      • Explicit representation of conceptualisation in a formal language (choose a language)
    • When can capture and coding stages be merged?
    • Differences between building ontology and creating a general knowledge base (thinking about methodology will help with this)
    • Integrating (during either or both of above)
      • Work must be done in agreement between communities
      • Make explicit all assumptions underlying an ontology
  • Evaluation
    • Judge against requirements specification (and/or)
    • Judge against competency questions (and/or)
    • Judge against real life
    • This paper looks at knowledge base systems, and adapts for ontologies.
  • Documentation
    • Desirable to have established guidelines for documenting
    • Main barrier to effective knowledge sharing is inadequate documentation
    • ALL important assumptions should be documented
  • Case Study
    • Main emphasis is on capture phase
    • Initially:
      • define ontology (Gruber)
      • identify users and usage (initially abstract, then clarify with real life)
      • choose language (Ontolingua was chosen)
      • choose method for capture - BDSM (IBM) supported by others:
        • KADS
        • IDEF5
        • OO Analysis and Design techinques
        • Gruber's principles for ontology design
    • Categorisation is fundamental to the human condition (Lakoff)
      • Not heirarchical, but:
        • GENERAL
                 ^
             BASIC  ->  primary with respect to knowledge organisation
                 v
          SPECIFIC
        • eg.
                SUPER:   Animal    /   Furniture
                BASIC:    Dog        /   Chair
                SUB:        Retriever /   Rocker
        • Certain concepts used subconsciously, rather than understood intellectually.
          • These have a more important psychological status.
        • Therefore paper uses middle-out approach to capture terms
          • (bottom-up = too much detail unnecessarily,
          • top-down = risks imprecision)
        • BASIC concepts first because:
          • most important
          • used to define non-BASIC terms
          • increase clarity, especially for non-technical use
          • backed by BSDM experience of paper author
  • Scoping
    • Brainstorming
    • Consult corpora if there aren't enough domain experts to brainstorm
    • Grouping
      • structure terms into naturally arising sub-groups
      • collate synonyms
      • consider things that might refer to each other
  • Meta-ontology
    • Don't commit too early, can restrict thinking.  Let concepts and relationships themselves determine requirements.
    • Be consistent.
    • Use technologically neutral language ('thing' vs 'entity').
    • Start with areas where there's most overlap.
    • Work from basic terms to more abstract ones within an area.
  • Producing definitions
    • Agreeing on definitions (varying degrees of problems)
    • Handling ambiguous terms
      • clarify ideas without technical terms
      • use a dictionary!
      • label definitions, eg. x1, x2
      • determine most important concept
      • choose a term, avoiding original ambigious one
    • Avoid new terms
    • Terms get in the way (peoples' preconceived ideas) - concentrate on underlying meaning and concepts.