Amy Guy

Raw Blog

Showing posts with label pst-events. Show all posts
Showing posts with label pst-events. Show all posts

Wednesday, November 20, 2013

OKFN Glasgow #2

I ventured to Glasgow for the second Open Knowledge Foundation meetup on Monday 18th. It was well attended, and there were six short talks:

Lorna Campbell from Cetis talked about Open Scotland. I understood this to be a collaboration between Cetis, the SQA, JISC and the ALT Scotland, to do with the opening up of education, and influencing policy and practicein this area. Here's a blog.

Grianne Hamilton from JISC talked about Mozilla's Open Badges. You can use them to reward learning, skills and achievements in all sorts of areas, and any organisaiton can create and issue badge packs to people who have earned them. Recievers can then show them off anywhere they can put HTML.

Graeme Arnott talked about a collaboration between Glasgow Womens' Library and Wikimedia, which resulted in the Scottish Women on Wikipedia event. This was a group of Scottish women getting together to edit Wikipedia articles about Scottish Women, and there was very positive feedback. They have more events planned. Graeme also reminded us about Wikimania, which is taking place next August in London.

Jennifer Jones told us about the Digital Common Wealth project. She pointed out that with media-saturated global events like the Olympics, the official story is already decided before the event even starts. An alternative to relying on what is broadcast by the mainstream media is to turn the camera on the crowd, and get the 'real' version of what is going on. The Digital Common Wealth project will encourage citizen journalists to work together to craft the story of the Glasgow Commonwealth Games from their perspective. Jennifer also raised the point that although free tools like YouTube, AudioBoo and Twitter are great for spreading stories, the data is still held by third parties - what happens if they disappear? How should initiatives like this safely archive their stories, and keep them in context?

Pippa Gardner talked about Glasgow's Future Cities project, for which they have £24 million to develop. It's about "people and data", but she was here to talk about data. There's the Data Innovation Engagement (which apparently needs a better acronym) and Glasgow's data portal which has already launched. Not all of the data on their is 'properly' open, but it's more open than it was before. There's a maps portal coming soon. Follow @openglasgow to keep up to date. Someone asked how they can avoid inadvertantly widening the digital divide by making all this data available - as it will only improve things for people who already have understanding and access. Pippa said there's a dedicate group in the Council working on widening digital participation, so they're involved.

Duncan Bain, and MPhil student at the University of Edinburgh, talked about Open Architecture. He says it's hard to define 'knowledge' and 'data' in architecture; architects create drawings/representations, not buildings. There are efforts towards opening certain aspects of this, like wikihouse.cc and the Open Architecture Network, but the culture of the architecture world, and where the money is, seems to be preventing things from going in the same direction as software development any time soon.

Here are livestreams of the talks by Jennifer Jones: one and two and a twitter timeline by Sheila Macneill.

Thursday, November 14, 2013

Prewired 3

Our third Prewired event went smoothly, with 20 young people (about 4 new) and 7 or so parents attending, plus 8 mentors. So lower signups than usual (a few cancellations due to school commitments), but we decided not to do a big publicity push and see how it ran with a smaller group. I didn't notice much difference, since they organise themselves into smaller groups anyway to work on different things. I think next time we'll try to reach our capacity of 40.

We had a big group working on a variety of Python projects (games, basics, algorithms, I'm not sure what else..), a small group doing front-end web, and quite a few doing amazing things with Scratch.

Every week I discover new things these super-talented young people are doing with their time, and it won't be long before many of them are spending a lot of time mentoring their peers as well as working on their own projects.

Nantas came by to talk about what he does with the University's Robotics lab, including the challenges of making humanoid robots play football, and the state of the art, two-million-pounds, full sized humanoid robot that is moving to Edinburgh in the near future. Definitely stuff to get young people excited about learning to code.

We've been trying to encourage them to code between Prewired sessions, too, and about half of them said they had. I hope by the next time all of them have, and I'm really excited to see what they're capable of making in a few months time!

But...

Some of the young people attending are disadvantaged by not being able to bring their own laptop, or having only really old laptops which can't support modern browsers and therefore have trouble even executing the JavaScript their writing (true story).

We'd love to be able to pay for a set of simple but up-to-date laptops that we could lend to the attendees who don't have their own during sessions.  This at least will put them on a level playing field with the others during the sessions, and I suspect that many of them have adequate desktop machines or family laptops at home.

Prewired runs on a budget of volunteer blood, sweat and tears, and zero pounds.  We're lucky enough to be able to use space in the University Informatics building for free, and there are no shortage of keen mentors and helpers willing to chip in their time (and in some cases cash for snacks).

So if you work for a company who might be able to support the purchase of resources for our young coders, or know someone who does, then please get in touch!

Friday, October 18, 2013

The Launch of Prewired

Several weeks of debating and planning following Young Rewired State finally came to fruition on the 16th of October, with our first Prewired event.

Thirty eight kidsyoung people arrived between 9:30 and 10 on that Wednesday morning (it was half term week in Scotland, so we weren't pulling them out of school), grabbed some kindly donated Google swag, made name badges with stickers and felt-tipped pens, and sat down for two and a half hours of lightly guided learning.



They were between the ages of three and eighteen, although the three to six year olds were more there to be tagging along with older siblings or University staff. It's obviously impossible to divide attendees up by age and decide what to work with them on, as older definitely does not mean more experienced. We had decided on no lower bound for the age limit, and no lower bound for experience either, figuring that the only real requirement is enthusiasm about programming. There was a huge mix of interests and abilities, and we let them decide for themselves which topics would be worth listening to.

We also had about fifteen students, University staff or industry professionals along as mentors.

After a few minutes of welcomes, where most of the room were willing to introduce themselves and tell us what they wanted to learn ("Python", "Scratch" and "more about programming in general" were popular ones) we kicked off with three five minute introductions: to HTML and CSS (beginner), to HTML5 Geolocation (intermediate) and to Python's Natural Language Toolkit (advanced). They then had the chance to spend 40 minutes in a hands-on session for whichever of these they chose. The groups were very evenly spread, and despite a few hiccups with Python installations on Windows and Chrome not playing nice with geolocation (worked through thanks largely to the mentors) most people got some code up and running and appropriately hacked about with by the end.



We took a break for juice, crisps, chocolate and fruit, plus a bit of hardware tinkering. We'd borrowed a Nodecopter, but hadn't managed to get it charged in time so it wasn't in the air, but there were still plenty of people interested in looking at the code to control it. We also had a demo of a robot arm, which could be controlled by an Android app connected to a Python server, which had been written over the summer by one of our mentors.




Next up were three more lightning talks: introduction to Scratch (beginner), doing cool things with Redstone Circuits in Minecraft (intermediate) and introduction to PyGame (intermediate-advanced). The following hands-on session for Scratch was under-attended, possibly ousted by the allure of Minecraft, but the PyGame session had over a third of the group and made some great progress, which was awesome.

We finished a little late, but still managed to have time for a quick demo of a football playing robot from the nearby robotics lab, and a few attendees who took their time dragging themselves away from their screens.



I'm told that overall it was a success. I was concerned because I was generally called upon when something was going wrong, so my perspectively was weighted towards the negative. But it wasn't too chaotic, none of the
kidsattendees played up, and as far as we could tell they were doing something in some way productive at all times.

A lot of them had had little to no programming experience before that morning, and I really hope they were able to take away something positive and, most importantly, feel encouraged to try things out by themselves at home. Plenty, too, had enough experience that they were calling out to correct the speakers, and helping their peers to get things working. It's a huge challenge to find enough activities to engage so many different levels of experience and interest, and I don't think we did a bad job.

Our next Prewired event will be on the 30th of October, and we're running them bi-weekly on Wednesday evenings from now on. They will be henceforth less structured. Our primary aim is to help young people to realise that with programming (and related areas) they can create anything, express themselves, and change the world. We don't wish to enforce a curriculum, but encourage them to explore areas they are interested in, learn how to teach themselves and figure out how to make what they want, and most of all to persuade them not to be afraid to experiment - to hack - and to just keep trying if it doesn't work first time. To get them excited before they become jaded and before this society's stereotypes have a chance to impact on them.

You can find out more about Prewired at prewired.org, and join the mailing list there too.

Photos and feedback


Here are some of the photos from the day:



If you took some that you'd like us to add, then please send them to hello@prewired.org!

Similarly, send any feedback you have about the event to us that way, as well.

Resources


I'll update this post (as well as the website) with resources from the speakers and mentors as I get hold of them.

Beginning HTML and CSS:

HTML5 Geolocation:

Building a chatbot with Python's Natural Language Toolkit:

Intro to Scratch:

Minecraft Redstone Circuits:

Intro to PyGame:
  • Coming soon...

Sunday, August 11, 2013

Young Rewired State in Edinburgh #yrs2013

Young Rewired State is a week-long hack event for under 19s.  There are centres all over the UK, and the week finishes with a giant sleepover in the Custard Factory in Birmingham, presentations and prizes.

I was helping out with running the Edinburgh centre this year, between the 5th and 11th of August.  We had 15 young people taking part, and a few parents popping in and out as well.  Not to mention several fantastic mentors.

Every day we gathered in one of the University of Edinburgh Informatics computer labs.  On the first day we did some brainstorming, introduced the young people to Open Data, and they sorted themselves into teams.

We had a diverse range of projects by the end of the week.

The Weatherproof app was written in Scala with a Web frontend, and as well as telling you the weather forecast, gives you practical advice on what to wear and what to take with you.

Stuff Index was a Python Web app that lets people photograph and upload stuff they've left out on the street that they want to get rid of, so anyone browsing the site can opt to take it away if they fancy it.  Helping to keep stuff out of landfill, and without the dreaded social interactions that come with Freegle.

Tag is a game by a one-man team, with a Python game server and a JavaScript front end that lets you chase your friends around the real world, and automatically tags them when you're in range.

PokeGame is a real-world Pokemon simulator that lets you roam IRL and capture virtual Pokemon.

Great stuff!

On Friday we crammed into a coach along with the participants from Aberdeen, Dundee and Glasgow, and set off on a seven hour road trip to Birmingham for the finale.

The Edinburgh teams didn't win anything, but the presentations were fantastic and everyone had an amazing time.  The young people made new friends, learnt tons of new stuff, and hopefully remain enthused about coding.

Next year we're going to do more to walk through the creation process of some example apps to get them started off, and maybe do a better job of introducing Open Data and the possibilities it holds.

We're also thinking about starting a regular under 19s code club in Edinburgh - weekly or bi-weekly - so stay tuned for more info about that.  (And if you want to help or participate, get in touch!)

Tuesday, July 09, 2013

#SSSW2013: Collaborative ontology engineering and team formation

We were introduced to the various mini-projects on Tuesday morning, and encouraged to form teams with people who weren't from the same university.  I quickly shortlisted the five that sounded most interesting to me, but was disappointed that there weren't any about multimedia.  Because how to evaluate a very subjective system is a potential problem for me, the project proposed by Valentina Presutti was my first choice:

"Serendipity can be defined as the combination of relevance and unexpectedness: an information is considered serendipitous if it is at the same time very relevant and unexpected for a given user and in the context of a given task. In other words, a user would learn new relevant knowledge. To evaluate the performance of a tool (e.g., an exploratory search tool, a recommending system) in terms of its ability to provide users with serendipitous knowledge is a hard task because both relevance and unexpectedness are highly subjective. This miniproject focuses on two main research questions: what is the correct way of designing a user-study for evaluating an exploratory search tool performance in terms of serendipity? Is it possible to build a reusable set of resources (a benchmark) for evaluating ability to produce serendipity, allowing easier evaluation experiments and comparison among different tools?"

Nobody else seemed to be interested though, so I resigned myself to not being able to do it... until I explained the project and why it was interesting, to the best of my ability, to Andy, Oscar and Josef, and they were sold enough to mark it as our first choice.  Thus Team Anaconda Disappointed (a name of significant and mysterious origins) was born, and Project Cusack (because of the movie Serendipity, which nobody got) was underway.

Our first lecture today was from Lynda Hardman, about telling stories with multimedia objects.  It was super relevant to what I'm doing, to the point where I'm surprised I hadn't come across her work already.  My notes are here.  Lynda has done, for example, work with annotation of personal media objects like holiday photos in order to combine them into a media presentation.  She has considered similar things to me, in particular noting that there are many many aspects of data about multimedia - I had assembled my take on this into a Venn diagram for my poster..



One I hadn't considered is annotating an explict message of a piece of media, intended by the creator.  This isn't always relevant - sometimes the consumer's interpretation of the media is more important - and this in itself might be an interesting annotation problem.  Competing perspectives - something an ontology should be able to represent.

I need to check out COMM - Core Ontology for Multimedia.

She has an overview of the canonical processes they have consolidated the process of producing digital content into, and how annotation can be formed around these.

Lynda also told us about Vox Populi and and LinkedTV; practical applications of annotating multimedia.

I made lots and lots of notes.

Natasha Hoy gave us some insights from the biomedical world with regards to ontology development, particularly in relation to the International Classification of Diseases which, when last revised in the 80s, consisted of a lot of paper and a whoever-shouts-the-loudest algorithm for inclusion of terms.  But the next version, currently under creation, is being developed with a version of Web Protege, customised to be friendly for those who don't know or care about ontologies, and is a truly collaborative process (for those allowed to take part) with accountability for all changes.  It's open too though, so even those without modification rights can view and comment on the developments.  My notes are here.

Lunch was for the first time outside, under the shadows of the forest, and for me was a tray of tomatoey vegetables that were delicious but few.  A striking contrast to Monday's lunch.  Everyone else had some meat-potato combination, preceded by a salad with tuna, and followed by a peach.

The hands-on session followed on from Natasha's talk.  We teamed up (temporarily Anaconda Hopeful) and played with Web Protégé.  There were two magazines and two newspapers, each with four departments.  Anaconda Hopeful were randomly designated the Advertising Department of Iberia Travel (a food and travel magazine).  We got stuck in, on paper first to identify some classes and relations that were relevant to us, and then with Web Protégé, along with the other departments of Iberia Travel.  We didn't come into any conflicts, but ended up creating a few classes that we needed, but should really have been the remit of another department (I guess we just got there first).

Then it was announced that Iberia Travel had bought the other magazine (and one of the newspapers had bought the other), and we had to work together to merge ontologies with the other department.  It became apparent that the other magazine had never had an Advertising Department (no wonder they went under!) so we had no-one to attempt to merge ontologies with.  We attempted to sell our expertise to the Advertising Departments of the newspapers, but there were already too many people involved in the heated debate that came out of the ontology merging there, so we couldn't really get involved.

Later we got cracking with our mini-projects.  Valentina showed us aemoo, and the experiments her team had come up with to try to evaluate it.  We sat down by ourselves to brainstorm, describing a lot of concepts for ourselves, breaking down the notion of serendipity, figuring out what might be wrong with existing experiments to 'measure' serendipity, and collating literature in the area.  (Turns out there is a lot, and it's a very interdisciplinary issue; lots to read about from social sciences, anthropology etc, as well as philosophy of science.  In computing, it seems to be primarily discussed within the realms of recommender systems and exploratory search).

Serendipity seems to be mainly described as a combination of unexpectedness and relevance.  Problems include the sheer subjectivity of it.  Some people are going to get excited by all facts they find out, whether they're useful or not.  Some people are going to have hidden, inexplicit or subconscious goals that affect how 'relevant' something is to them.  People describe their different areas of expertise in different ways; some are more humble than others and would not call themselves an expert in a topic, for example.  So whether or not an event can be considered a serendipitous one is a complex question, which must take into account the person's background, goals and existing knowledge, the task they are trying to achieve (or lack thereof, as serendipity is particularly important - in my opinion - in undirected, loosely-motivated activities), the way they are able or encouraged to interact with a system, what they are doing before and after... all these things make up a context for someone's activities, and none of them seem to be particularly measurable.

Dinner was a vegetable and potato (yay!) starter, followed by spaghetti in tomato sauce (fish for everyone else, although Andy got a custom omelette, lah-de-dah).  Also an apple.  We learnt the hard way not to sit at a table directly underneath a light, as the bugs just raiiiin down.

After dinner we crowded around Enrico who had offered to provide advice about PhD-ing.  From this session, I have a signed diagram of the life of a PhD, because he borrowed my notebook to make it.  I tuned in and out of the discussion, and noticed some irregularities between my PhD and what seemed to be 'normal'.  For instance, most people didn't seem to have as much control over their topic, or what they were doing at any given moment in their first year.  I am really, really enjoying my freedom, but in order to justify that I deserve it I need to sort out my lack of direction and focus.  I need to believe in what I'm doing - not be told by someone else - which is one of the main reasons I am doing this particular PhD.  Perhaps I need to ask for more guidance to more quickly reach the necessary conclusions for myself.  (And, of course, perhaps I also need to stop taking big chunks of time out periodically for different reasons; that might speed up the process as well).

Later, the overriding sentiment was that the job of a PhD student was to answer a question, to produce a theory.  Not to create a system or solve a large problem; certainly not to worry about practical, real-world applications of theories.  Well, I've already explained that this is something I can't accept, and I still am not convinced that that is going to impact on my ability to do a PhD.  Theories develop during practice.  Coding and designing, like writing, are part of my thought processes, and I reach realisations or find new questions to ask through hacking and playing and making.  And why would I be hacking and playing and making, if not to try to produce something of real-world value?  If my motivation in making a system is explicitly to come up with new theories, then my approach and outcomes and realisations will be entirely different.  In trying to make something that works for real people, not researchers in a restricted domain or specific context; a clean and sterile laboratory, I figure out different things, that matter.

There was another discussion that I came into a bit late, but it sounded like a very harsh discussion about problems with research in industry (rather than academia) that seemed to be very overstated compared to what I have read and experienced myself.

By the end of the day, it felt like I'd been at Summer School for weeks, and had known everyone forever.

[Notes] Natasha Noy at #SSSW2013



Stanford, Protégé.

In past 10-15 years, through collaboration with scientists (particularly biomed), ontologies have become essential.

Don't need to sell ontologies to scientists, they believe in it.

Focus on science because that's where she has experience etc.

We're not so bad at versioning ontologies, more versioning data is the problem.

Experts add stuff, curator checks quality, and publishes upcoming tasks.

Similar to open source developments, but no research to compare the two.
- Different because biomed people are paid (well).

ICD - International Classification of Diseases.
  • Started 17th century.
  • Causes of death, medical bills, policy making.
  • Revised in 80s over 8 annual conferences.
    • 17-58 countries, 1-5 person delegations, mainly health statisticians.
    • Manual, on paper.
    • Whoever shouted loudest..
    • Paper copies, only English, pdf.
  • ICD-11 - OWL ontology!
    • Open, Protégé (a customised, Web version), links to others.

Conflict resolution:
  • People naturally don't step on each others' toes.
  • Users expect stuff like Web 2.0 interactions, Web interface.
Web Protégé:
  • No consistency checking - coming but currently must go offline.
  • Ontologies are solution to everything - versioning, roles, social interactions.
  • Also plugins are the solutions to everything - visualisations.

Monday, July 08, 2013

#SSSW2013: Research in theory and practice, and where on earth am I?

The 10th Summer School for Ontology Engineering and the Semantic Web

Sunday

Arriving by train into Cercedilla, north of Madrid, we immediately encountered other confused looking folk with poster tubes.  So we shared taxis (EUR 10) from Cercedilla station to the summer school residence further north, in the forest.

After getting keys for our pleasant, single, en-suite rooms, arrivals congregated in the shade by the building  to introduce ourselves.. Again, and again, and again, as new people continuously arrived over the space of a few hours.

A really broad mix of people are here in terms of nationalities and places and levels of study, but I still haven't quite got used to the fact that answering 'Semantic Web stuff' is not specific enough in this crowd, when someone asks you what your research is about.  Nobody needs convincing that these technologies are useful!

Later we received schedules, maps, ill-fitting t-shirts* and very helpful name badges, and headed for dinner at the bar down the road.

As is traditional when I write about my experiences in new places, I will describe the food every day.  It has become apparent, at this residence at least, that variety of ingredients is not ordinary, so in this respect meals are simple.  Dinner that first night started with a salad (lettuce, olives, tomato, onion, shredded beetroot and a single slice of hard boiled egg; no dressing), followed by - for the majority - slices of meat (beef? Pork? I dunno..) and fries.  Mine was a plate of mushy green vegetables with a little seasoning, that was pretty tasty.  Dessert was a single pear, delivered with ceremony, but otherwise unadorned.  Healthy, at least.

Yet we were all (those I sat with at least) were left feeling a little unsatisfied.

I shared a table with a French, Spanish, Italian and Irish guy.  Conforming appropriately to stereotypes, and setting up reputations for the rest of the week, the French and the Italian shared the bottle of wine on the table; the rest of us went without.

I returned to bed after a couple of hours of socialising and enjoying the cool air in and around the bar.

* For next year, they could ask for t-shirt sizes when they ask for dietary preferences?

Monday

The day started early, and with no hot water or wifi for anyone.  Breakfast was combinations of sweet pastries, coffee, tea, juice and bread.

Punctuated variously by coffee breaks, the learning began in earnest.

During the introduction by Mathieu D'Aquin, I found out that I am one of 53 students selected out of 96 applicants to attend this year's Summer School of the Semantic Web!  I had no idea it was that selective, or that there had been that much competition.

The first keynote was by Frank van Harmelen, about all the Semantic Web questions we couldn't ask ten years ago.

Slides:



Frank started by saying that the early Semantic Web vision has morphed into the more manageable vision of a Web of Data, or a Giant Global Graph, and outlined the principles of the Semantic Web as they appear to stand at present:

1. Give everything a name (entities).
2. Relations form graph between things.
3. Names are addresses on the Web (so we inherit properties of Web like AAA).
4. Add semantics.

Frank pointed out the advantages of the fact the Linked Data crowd, grown naturally and not designed, is now so big we don't know how many triples it contains, nor how fast it is growing.  Companies and organisations (like Google, NXP, BBC, DataGov) are using Semantic Web technologies to achieve their own ends, for a variety of different use cases, without caring much about the Semantic Web, and this is contributing to the growth.

This growth has given rise to a number of research areas that were impossible to realisitically ask questions about ten years ago, including self-organisation, distribution of data, provenance, dynamics and change, errors and noise (how to deal with disagreements).

Frank asserted that rules and structures, algorithms and patterns in data, exist whether we are looking at them or not.  He used the analogy that OWL is our microscope, and it may be the tool that distorts our vision of the information universe rather than properties of what we are looking at (for example, structures in data presenting themselves well in some domains but not others).

He went on to promote the roll of the Informatician to be to test theories, hypothesis and falsify, as scientists rather than engineers.  To discover, rather than build.

I struggle with this view of the world, and feel instinctively that theory and practice are intrinsically linked; one can't exist without the other, not just in the grand scheme of things, but in day to day work and research.  This is one of the main points of contention with my own PhD, and I've no doubt there will be many more blog posts about this issue in the near future as I reconcile my need to create something immediately useful with the necessity of producing a contribution to knowledge at large.

See my raw notes here.

We had an Introduction to Linked Data by Mathieu D'Aquin (raw notes here), followed by a workshop.  We wrote SPARQL queries to populate a pre-written web page with information about Open University courses, sub-courses and locations thereof.

Lunch, similar to the previous night's dinner, was a starter salad, an entire half chicken (or something) plus fries for the carnivores and the most unappealing risotto of my life for (not that I'm ungrateful, but I have never been unable to finish a meal due to boredom before).  I went for a walk with some others to grab some fresh air before the afternoon's work, and missed out on watermelon.

Manfred Hauswirth presented some really exciting stuff about annotating and using streams of data.  Particularly challenging is how to integrate this with static data and make inferences over the lot.  Streams include sensor data, as well as ever-flowing social media streams for example; anything that changes over time.

They've built some systems to process this kind of data, and one of them is available as middleware.

My raw notes are here.

In the afternoon we had a poster session, where all participants pinned up posters about their work, and discussed at length with anyone who was interested.  Here's evidence that I participated.


And here's Paolo's:



I wrote a few notes about things from other peoples' posters that I need to look up.

The main feedback I received was about making sure I focus, narrow down my topic, and concentrate on some evaluatable deliverables that are PhD-worthy.

Questions like (paraphrasing) "why should we care about digital creatives?" threw me, because I thought the obvious answer - that they are people too, Web users, technology users, contributors to culture and an ecosystem of digital content and data - was apparently not enough from an academic standpoint.

I was simultaneously told to focus more, and to explain why the problem I'm trying to solve is applicable to all domains, not just digital creatives.  But some of the problems I'm looking at have been (or are being) solved in other domains (like e-health, biological research, education) and the reason what I'm doing is interesting is because none of these solutions quite work for digital creatives, and I want to find solutions that do, and try to figure out why.

I'm still stuck in some sort of struggle between theory and practice; thinking and doing.  And the long-standing problem of how to decide which doing actually worked.

I've started scribbling notes about the narrowing down problem.  I'll need to have this figured out before my first year review in August anyway, so stay tuned for another post all about it.

Then I sneaked off for a nap.

Dinner at the bar again; the usual salad, plus some eggy fish thing for most.  I got a plate of artichoke.  Artichoke is great, I love it, and I'm all for simple meals.  But I remain unconvinced that a plate of only artichoke constitutes an acceptable level of effort on the part of caterers.  And the sheer quantity made it start to taste a bit funny after a while.  But not to worry; we rounded off with a solitary peach apiece.

Further socialising, and appreciation of the night sky, before returning to bed write blog posts.

I'm super excited and inspired by the talks, work I've heard about so far, and the atomsphere of the place.  I'm excited to learn a helluva lot, and remind myself that I'm not facing impossible problems, and am not facing many problems alone.  I remember that I am instinctively passionate about the Web and the possibilities it holds (and indeed has already realised) for the empowerment of individuals.  I remember how lucky I am to be able to sustain myself through studying something I love so much, and to have the potential to make a change, and through my work maybe even facilitate others to be able to make a living doing what they love, as well.

[Notes] Poster session at #SSSW2013

Things to investigate further!

LibRDF - linkeddata-perl for Debian by Kjetil Kjernsmo.

Rakebul Hasan, Fabien Geandon (? not sure about names, can't read my handwriting..) - Trustworthiness of inferences.

Taldea - fostering spontaneous communities.
Ghada Ben Nejma.

NERD ontology for spotting entities.
nerd.eurecom.fr
See photo:


[Notes] Manfred Hauswirth at #SSSW2013




Streams: Any time dependant data / changes over time.

Has done a paper about P2P stuff.

Data silo - "natural enemy of SW scientists"

Massive exponential growth of global data.

Still have to integrate dynamic data with static data.
Multiway joins are domintion operator.  Need to be efficient.

Everything/body is a sensor.

Various research challenges:

  • Query framework.
  • Efficient evaluation algorithm.
  • Optimise queries.
  • Organisation of data.

CoAP ~= http for sensors.

Stuff about sensor networks and context - useful for Michael.

  • Common abstraction levels for understanding.
  • SSN-XG ontology
    • Application: SPITFIRE
  • You can buy a sensor off the shelf that runs a binary RDF store and can be queried.  So possible to use SW tech with resource constrained devices.
  • RESTful sensor interfaces stuff being standardised - CoRE, CoAP.
  • Linked Stream Model
  • CQELS-QL (extension to SPARQL 1.1; already legacy)

Rewrite query to spit out static and dynamic - lots of overhead.
But need to optimise between these.
Neither existing stream processing systems nor existing databases could be efficient enough.
So the built own LD stream processing system.  (Optimised and adopted existing database stuff).

HyperWave - didn't succeed.  Didn't listen to customers and wasn't open source (license fees).
But better than hypertext was back in the day.
Performance important for success/uptake.

Just putting it on cloud infrastructure doesn't mean it scales.

  • Need to parallelize algorithm.
  • Took it to a point where adding more hardware did help.
  • Problems!  Inconsistent results, engines don't support all query patterns.. very early, don't fully understand yet.
  • Long way to go.  How to prove what is a correct result?
  • Needs to be easy to use - dumb it down.
    • Linked Stream Middleware (available):
      • Flights, _live trains_ - SPARQL endpoint!, traffic cams.
      • SuperStreamCollider.org
      • Current Tomcat problem with twitter streams.

To do?

  • Scaleability
  • Stream reasoning (only processing, pattern matching, so far.  Want to infer conclusions).

World is:
... uncertain, fuzzy, contradictory.
So combine statistics and logics.
Hard to scale logical reasoning, so use statistics to shoot in the right direction.

Privacy?

  • Build systems! Can't do thought experiments about the Web.

Don't get hung up on approaches / labels.

[Notes] Introduction to Linked Data at #SSSW2013

(by Mathieu D'Aquin).



Linked Data = universal connections, like Lego.

Universal is why it's important.

Workshop instructions.

The only problem we had during the workshop was disagreement about how to read the 'broader' and 'narrower' relations between courses.  It instinctively ready contrary to what (my) common sense suggested (eg. that 'arts and humanities' is broader than 'history', which some people disagreed with).  A quick reference to the ontology documentation resolved that.

[Notes] Frank van Harmelen at #SSSW2013


Semantic Web & Web of data = a more manageable mission.
Metaweb movie - got bought by Google and incorporated into Knowledge Graph.

SW Principles:
1. Give everything a name (entities).
2. Relations form graph between things.
3. Names are addresses on the Web (so we inherit properties of Web like AAA).

This becomes Giant Global Graph.  (Maybe SW should be called Giant Global Graph?)

4. Add semantics.

  • Types of things, relationships.
  • Hierarchy, constraints:
    • Inferences.  Bounding shared beliefs by sharing ontological information.  Space for confusion gets smaller and we begin to agree on interpretation of information.
Semantics = predictable inference.

Google: from just links to results, to information boxes (last May).  Can't directly address Google Knowledge Graph.
NXP (microprocessors): 26,000 products. Integrated all databases into triplestore.  Exposing subset of triplestore to customers.
BBC: 125 million triples.  Many data sources.  APIs to website.  Own ontologies.

All have the same triple-layer architecture:

Raw data
   |
SW layer
   |
Output / API / UI etc

DataGov: eg. air quality in cities, campaign money, if policies work.

Companies don't care about SW, but are using these technologies for their own IRL purposes.

These are all different types of use cases of SW technologies:

  • search;
  • data integration;
  • content re-use;
  • SEO;
  • data publishing.

It's important that the SW graph is so big.

  • More questions to ask.
  • Good that we no longer know how big, or how fast it is growing... Tens of billions of facts.
    • How many are really permanent?
    • Some are stable, some will disappear - just like the 'regular' Web.
      • "...it being a mess is the only reason why it scales."
We need to get used to the idea of SW being a mess - aka "a system so large you can no longer enforce central control" (complex system).

The LD cloud is still poorly interconnected, but good graph properties.

SameAs.org

Heterogeneity is unavoidable.
Socio-economic, first to market - why certain systems/ontologies get used, eg. schema.org, dbpedia.

Self-organisation.
LD cloud grew, nobody designed it.
Knowledge follows power curve.  This has an impact on mapping and reasoning, storage and indexing.

Distribution.
Web not geared for distributed SPARQL queries.  Everyone pulls in all data and queries local copy.  Not very 'webby', disadvantageous.  So subgraphs?  Query planning?  Caching?  Payload priority?

Provenance.
Representation, (re)construction.  Metametadata (knowledge about knowledge; uncertainty; problems with vocabs for this).
How to get from provenance to trust.

Dynamics (change).
Cool Web in 60 seconds graphic.
SW not changing this fast, but soon..

Errors and noise.
Sometimes we disagree.
Deal with by: avoid, repair or contain.  Or just deal with it - allow argumentation.
Fuzzy, rough semantics - almost, maybe.

Lots of research questions.  But not ones we could ask 10 years ago.

Information universe - "algorithms exist without us looking at them".

We should ask if things work in theory.
Scientists vs. engineers.
Discovering vs. building.
  • Is this incidental or universal?
OWL is our microscope.
We can see structure well in some domains, but not so well in others.  Maybe it's our tool that distorts, rather than a property of the domain.

Says we should change our mindset from building stuff to hypothesising and falsifying.

Sunday, May 26, 2013

Week in review: VidFest

20th - 26th May

Continued to work on literature review.  Nothing much to report.

Went to MCM Expo in London and managed to find time (around non-stop merch selling for TomSka and Eddsworld) to ask between 30 and 40 content creators - a wide variety of ages, experience, types of content - about their process and collaborative practices.  The thing they all had in common (I randomly picked people as they were waiting in the two hour long queue to get autographs from Tom) was that they all do what they do because the love it, want to entertain people, and if the could earn a living from it too that would be amazing; but that's not why they do it.  For many it's the dream, but not one they expect realistically to achieve.

That is why this is important to me.  Because everybody should be able to make a living from doing what they love*, and the technology exists to allow it.  How exciting.

* Unless they're really bad at it.  There's only so much technology can do.  But they should definitely have the chance to get good before caving in to a ninetofive that they're not totally passionate about.

Saturday, May 18, 2013

OKFN Meetup #6

Thursday 16th of May was the 6th Open Knowledge Foundation meetup in Edinburgh.  We had a great room in Techcube, more speakers than usual and loads of attendees.  Here's an account.

Bill Roberts

Bill, founder of Swirrl ("the linked data company") talked about tools and user interfaces they are developing to make handling data easier for communities; particularly for the less technical.  They've encountered a spectrum of users with different levels of technical abilities and needs, so they have to account for this in the tools they build.  Technical complexity for accessing data ranges from SPARQL endpoints, JSON APIs, downloadable spreadsheets to visualisations, maps and charts.

They're focussing on providing the data in accessible ways rather than building visualisations though.  They'd struggle to meet or even understand everyone's needs; instead, it's important to concentrate in empowering the communities to use the data themselves.

Kim Taylor

An undergraduate Informatics student and participant of the Smart Data Hack, Kim showed us placED, the project her team had worked on, and are continuing to develop.

This is a place finder for people who are looking to maintain or improve their personal wellbeing.  They used datasets from the City of Edinburgh Council (and presumably ALISS?) to create an Android app. They stored their data in a Google AppEngine datastore, but I'm not sure if it has a web frontend as well.

Some of the problems they encountered include copyright issues with ordnance survey place data, and royal mail postcode data, which made up part of the Council's data but wasn't available for anyone to use due to licensing restrictions.  They worked around this by recomputing location data from the parts of addresses the did have access to with Google's geocoding API.

When they open their database up for user input, which they inevitably will if they want their app to stay current and useful, they'll have to think about how to maintain the content.

Gavin Crosby

Gavin works for the Council, with a title I've forgotten, but it's to do with youth work.  Youth work has a very specific definition to do with people aged between 11 and 25, meeting in organised groups with a volunteering adult present.  There's loads of this going on in Edinburgh, organised by Scouts/Guides, schools, churches or maybe even self-organised.  There's no central database about what is going on where, which is one of the Council's biggest issues in this area.  Word of mouth is usually how this kind of information is spread amongst young people, and Gavin suggested that a lot of youths may be unwilling to attend something they'd heard about without a direct invitation from someone they know.

In an attempt to reign some of this information in, they've created the Youth Work Map.

It's not an ideal system, as they have to update it manually when a youth group or activity organiser decide to inform the Council that they exist.  Not everybody opts in, so there is data missing.  Manual updating also means the map is not 'live'; things might go out of date and not be removed straight away.

Gavin said it is the constraints of the Council's web system that has caused a lot of the problems, and points out that they haven't considered accessibility issues (for example, access for people with vision problems), and it's not interactive.  He'd love to see the ability for kids to chat to each other through the map, or leave reviews for particular events.  There are issues with child protection here, of course.

He would also like to see better tagging and organisation of the content on the map, links to other data repositories (there are parallel similar projects), and the ability to connect events to areas or routes rather than single points.

Gavin pointed out that a lot of the audience for this map is likely to be adults looking for youth projects, rather than young people themselves.

Leah Lockhart

Leah made a quick announcement about the new Local Government Open Data Working Group.  They're organising open data surgeries (similar to her social media surgeries that you've definitely heard of by now if you're floating around the OD scene in Edinburgh).  They're also hoping to fill in the OKFN Open Data Census for Scotland, and meet regularly in the pub.

Tweet Leah if you're interested!


Fiona McNeill

Fiona works in Informatics at the University of Edinburgh, and she told us about Open Data in climate change science, or the lack thereof.  A team she has put together have got some funding to carry out a small investigation about Open Data use in climate change science, and to try to build a network around this.  They'll be looking at trends and patterns of the past decade to see if research has been any more successful when existing datasets were used, or if papers are more well-cited when they make their data open at the end of it (for example).

She thinks the lack of Open Data in this area could be due to the expensive nature of making data good enough quality to share, and of course the fact that when people have worked hard to gather data they feel that they own it; why should they share?

They're hoping that their report might go some way to persuading funding bodies to have sharing of data as a criteria for applications.

Contact Fiona if you're interested in this kind of thing.

John Kellas

John said, brilliantly, that talking about information visualisation usually means graphs.  But normal people "don't think in graphs".

He works in community education, and a couple of years ago he started working in "volumetric and comparative" visualisations, which can be much more useful and empowering to people.  He showed us a visualisation of one trillion dollars (which I can't find a link to, so let me know if anyone has one).

He's not had much support with creating tools and visualisations, because he's not interested in making money from it, so it's hard to attract funding.  What he's doing looks really useful though, so hopefully we'll see more!

Ben Jeffery

Ben is another undergraduate Informatics student who took part in the Smart Data Hack and whose team is still working on the project they started at the hack.  They're re-imagining the University's student information portal by pulling in lots of different data sources, and presenting the information more sensibly.  They've been doing a fantastic job, but of course are all busy with exams and general learning, so the haven't been able to spend as much time on this as they'd like.

They're also struggling to get raw data out of the University, and point to (my alma mater) the University of Lincoln's open data portal as an example of what could and should be done about this.  So they're turning their project into a pilot to demonstrate what they could do if they had the data they wanted.  They're also conscious of similar-but-different projects, like projects.ed.ac.uk, and don't want to duplicate effort.

Ben said they've found that a lot of the University of Edinburgh's data is held by middleware vendors, so it's particularly hard to access.  But this is information that is funded by students, so it should be available to them!  He said the "University should be a breeding ground for knowledge" so data shouldn't be silo'd up.

He also said that there are a lot of politics in the way with this sort of thing.  They, as any level-headed software developer, just want to build stuff.  They're still in various talks though, so this is a space to watch...

Susan Pettie and Marc Horne

These guys are from So Say Scotland and aim to change culture to make Scotland better.  Open Data is important for democratic movements, so they told us about some of their events.  They're building a network of activists and campaigners, and hold large scale assemblies themed around 'thinking together', which is a kind of en masse guided brainstorming.  They're trying to spark a movement, and are aiming for 25,000 people.  They're investigating ways to make their assemblies more efficient, as currently collating all of the ideas that are generated is a manual process.  This would be nigh on impossible when they reach their participation goal.

There will be a report about their progress on the 27th of May.

Devon Walshe

Devon was our Techcube host, and he told us about Sync Geeks, Geeks in Residence.  This is a program funded by Creative Scotland that puts the technologically minded into arts organisations.  Previous efforts by arts organisations to employ 'geeks' to solve a technical problem or produce a digital solution for something have been problematic due to the 'black box' approach.  The developers produce an outcome, get paid and leave, often aiming to do the minimum amount of work.  Geeks in Residence promotes developers and the organisations working together more closely, to allow for sustainable solutions.

Part of the project is to analyse the relationships of people who know about technology, and those who don't, with each other.

In my notes I've scribbled "convert fear into technology", and I can't remember what that originated from, but it sounds awesome.

Devon did some work with Stills photography centre.  Nobody knew what they needed, so after some collaboration they developed an interactive floor plan (because the Stills building is way confusing) and some kind of interactive timeline because Stills has an interesting history.

He also plugged the Culture Hack Scotland in Glasgow in July (12th-14th), which I'm terribly disappointed I won't be in the country for.

Next OKFN meetup

Will be on the 22nd of August, in Informatics.  Here's a link to the Meetup so you can RSVP.  I'll sure be there, if I'm not somewhere else.. (depends if/when/where my Mum books an obligatory family holiday).


Thursday, April 25, 2013

Starting up in IT panel discussion


I had an amazing evening at the Starting up in IT panel discussion, followed by Innis & Gunn beer tasting on Thursday evening.  It was held in the shiny MMS Quartermile One offices.  (When I'm rich, I want a flat on Quartermile.  A turret-y one, not a glass one.  Or maybe both).

I felt chronically under-dressed when I arrived - a majority were suited - but everyone was really friendly and forthcoming with advice.

Anyway, speaking of being rich.  There were lots of interesting business-wise people to talk to at this event, including CEO of Skyscanner Gareth Williams, and Craig Anderson of Pentech Ventures.  Plus lawyers specialising in things like IP, employment, company formation, from MMS.  The panel discussion was enlightening; I'll go through some highlights raw notes...

Funding


  • Skyscanner - 2 mil from Scottish Equity Partners 2007.
  • Getting funding isn't a goal or validation.
  • Best way to get funding is not to need it.
  • Scottish Enterprise: match funding.
  • Give as much as you get. Confide in investor.

Getting wise

  • Don't pitch too early. Build traction first.
  • Prove potential marketshare one way or another.
  • Preparing business plan is productive.  Converting to a vision to a plan when you get funding.
  • Subscribe to investment bloggers.
  • Networkiiiing. Find someone to champion you to an investor.
  • Gareth: As many people are delusional as have a key insight. How to know which you are yourself?


Employees

  • Do you need employees or contractors? Casual employees in between.
  • Consultant / contractors own IP for work they do. Unless contract says otherwise.  Employees don't, employer owns it.



I heard about some really interesting ventures, too, like Identity Artworks which looks like they're making a huge difference to young people, and have really inspiring stories to tell.  Plus ShareIn, soon launching an equity crowdfunding platform. Veeerrry interesting...

The panel was followed by beer tasting hosted by Innis & Gunn.  I don't drink, but I would have sipped along to be sociable.  However, it turned out the beer wasn't vegetarian (filtered through isinglass).  This, at least, meant more for everyone else on my table.  MMS had come up with a written seating plan, by the way, that separated people who had arrived together.  Forced networking!  Excellent.

This served as great chance for Steve and I to independently practice our GeoLit elevator pitching, and I think we'd got it down to perfection by the end of the evening.  Extremely encouragingly, we were consistently met with enthusiasm and responses like "that's an amazing idea!".  We left pretty buzzing.


Monday, April 22, 2013

[Notes] How to write a literature review workshop


Just notes!

Workshop by Dr Mimo Caenepeel on Monday 22nd April.

'Critical' does not mean you have to pass judgement, or say why it's good or bad.
Not taking things at face value.

Started with freewriting about what has particularly influenced / inspired our own research.  Five minutes, not allowed to stop or edit, don't worry about quality of writing, not for anyone else to read.  A good way to get ideas out of your head and start to organise your thoughts without censoring or constraining yourself.

How many pages will a review usually take up in a thesis?  My policy is to write what needs to be written and stop when you're done.  But apparently 20 to 30, sometimes more, is normal in sciences.

There's no consistent / right answer to 'how many publications to review'.  For some people it's in the tens, for some the hundreds.

Think about how to integrate literature review into the thesis.  You're unlikely to have a chapter that is just 'literature review' and no mention of the background reading elsewhere.

Good qualities for a lit review?
- Coherence (avoid fragmentation)
- Structure, clarity.
- Proof of novelty - purposeful.

A review can often be considered as an indicator of the quality of the rest of the research - demonstrating scholarship.

A good place to start:
1. Write your research question, formulated as a question.
2. Write up to five research areas that are relevant to your research question.
3. Note some related issues/areas that will not be considered in your review.

Think about balance of content.
1. Three studies influential in your field (I couldn't answer this, I clearly need to read more).
2. Two significan older contributions.
3. Five recent sources.
4. Two sources that have strongly influenced your thinking.

You don't need to consider all papers in the same level of detail.  Decide which papers are more important / useful than others.

For some papers (important ones) you should work through these questions in the same way every time you read something (this is 'SQ3R'):
1. Survey: What is the gist of the article? Skim the title, abstract, introduction, conclusion and section headings. What stands out?
2. Question: Which aspects of the research are particularly relevant for your review? Articulate some relevant questions the article might address.
3. Read: Read through the text more slowly and in more detail and highlight key points / key words.  Identify connections with other material you have read.
4. Recall: Divide the text into manageable chunks and summarise each chunk in a sentence.
5. Review: To what extent has the text answered the questions you formulated earlier?

Critical reading (these seem like really useful questions to work through whilst reading papers):
1. What is the author's central argument or main point, ie. what does the author want you, the reader, to accept?
2. What conclusions does the author reach?
3. What evidence does the author put foward in support of his or her conclusions?
4. Do you think the evidence is strong enough to support the arguments and conclusions, ie. is the evidence relevant and far-reaching enough?
5. Does the author make any unstated assumptions about shared beliefs with readers?
6. Can these assumptions be challenged?
7. Could the text's scientific, cultural or historical context have an effect on the author's assumptions, the content and the way it has been presented?

See Ridley, D. The Literature Review: A step-by-step guide for students.  Sage Study Skills Series. Sage Publications, 2011 (2008).

Thursday, April 18, 2013

[Notes] 'How to write a thesis' workshop

Just notes from a three-hour workshop about how to write an Informatics thesis, on the 16th of April.


State contributions (to knowledge) explicitly.  Intro, conclusions; each chapter should have some (probably not all) contributions discussed.  Be obvious; use headings.

Knowledge - background:

  • justify choices
  • explain methods
  • acknowledge alternatives
  • evaluate

Evidence, well-reasoned arguments, acknowledge limitations.

Clear openings for future work.  Be clear where they are.

Make it reproduceable.

Short / concise.  Examiners like short theses.

Introduce what's interesting and important.

When outline thesis, look at structure of main argument, not of document.

Background material must have point.  Only include as much detail as you need to make point.
Points, eg:

  • Explain method you use.
  • Novelty of your approach. Similarities with existing work.
  • Justify choices (evaluate other work).
  • Don't tear down others' work. 'Build on'.
  • Cite examiners, they've probably published something relevant.. (but not for the sake of it).


Then we had five minutes to write down what our PhDs are about and what we have already found out.  I wrote:

How do the futures of the Semantic Web and amateur digital content creation fit together?
Can Semantic Web tools and technologies be used to enhance collaborative creative partnerships and encourage fruitful outputs?

There are knowledge sharing systems and collaborative tools for scientific fields and in education, but nothing for creative artsy things.

Attitudes towards data sharing and privacy amongst content creators are in flux.  There are lots of projects and energy around open data and decentralised social networks that allow data to become portable and not tied to one platform.  One of TBL's visions for the Semantic Web is the dissolution of data silos and 'walled' applications that disadvantage the user, and as such the promotion of the 'ownership' of a user's data by the user themselves, rather than the software or organisation that uses the data.

There are lots of reasons people make content.  There are lots of reasons people don't make content (who could / would like to).

[Notes resume]
Use backreferences; don't repeat yourself.

Info / advice
...homepages.../sgwater/resources.html
..homepages.../imurray2/teaching/writing
Style: Toward Clarity & Grace (book)
The Craft of Research (book)

When to start writing thesis?

  • Do you already have papers?  Slot them into a thesis template asap.
  • Maybe a year beforehand.  Slower pace is better.

Don't assume appendices will be read.  More for extra info if needed by people trying to reproduce your work (not your examiners).

Too many direct quotes look like you don't understand and are avoiding explaining yourself.

Keep copies of web resources and cite access dates in case they change / disappear.
Figures might be copyright if you just copy them from papers, even if you cite them.  Remake them, and put 'adapted from' as citation.

Examiners?

  • Depends on your supervisor.  Discuss.  Student might be able to suggest someone to examine.
  • Maybe a balance between internal and external knowledge.
  • Won't be someone junior, even if they're considered an expert in the field.
  • Helpful if supervisor knows how that person will behave in viva.  Might be a good reason to avoid someone you think would be perfect from their background.
  • Conflict of interest regulations.  You can know them personally though.  External can't have been affiliated with UoE in the last three years, or substantially involved in your research (like co-authoring a paper).  No ex-supervisors, from any university.

No grading system (ie no different levels of passed PhD).  Might be external prizes if you want extra recognition.

Thursday, April 11, 2013

2nd UK Ontology Networks Workshop

The UK Ontology Networks Workshop took place over one day in the Informatics Forum.

There was a mix of people there; some talks were way over my head and very technical, and some talks were by people who confessed they had had to look up "ontology" that morning.  And things in between.

Lazy writeup, but following are notes as I scribbled them:



John Callahan

US navy research.
Focused information integration.
Human intervention to keep predictive part on track. Tweaking.

Alan Bundy

Interaction of representation and reasoning.
Changing world so agents must evolve. How to automate? What would trigger a need for change:
Inconsistency
Incompleteness
Inefficiency
how to diagnose which?
Interested in language and perception change.
Unsorted first order logic algorithm called Reformation. Based on standard unification algorithm.
Allows blocking and unblocking unification.


Phil Barker

Schema.org
Cetis (JISC funded)
learning resource metadata initiative.
Big names behind schema.org.
= ontology + syntax
Big and growing ontology.
Dumbed down for people.
LRMI adds to it. W3C go through it. It's creeping, how much do the big names actually care about stuff that's added?
don't know how Google uses it.
People should consider using it for more sophisticated search and disambiguation.

Gill Hamilton

Doing more with library metadata. Learnt from OKFN. Had to convince people in charge.
Dublin core, didn't like; not specific enough. Instead RDF > OWL. "We know best how to structure our data"

Hardest was convincing marketing people that there was no commercial value. Metadata is advert to actual resource.

Enrico Motta

Traditionally top down approach. So now so many people interacting with semantic structures, so should involve users.
Recognise there isn't a unique or best way of doing things.
Initial study included modeling task with binary relations.

Patterns that are more or less intuitive. 4D least, 3D+1 most.
N-ary most widely used by experts.

Relationship between reasoning power and intuitiveness of writing? More creativity needed for simpler ones. (Not really sure what he's saying)

Email him for copy of study.

Chris Mellish

Ontology authoring is hard. Better ways to do it.

Controlled language input (mature tech); responsive reasoning (also mature, information as you're editing); understanding the process (beginning to understand more).

Hypotheses:
users don't know what they're doing. What if questions.  Many answers, what is relevant? Depends on context.

Authoring as dialogue.
Todo list.

Useable in the same ways as protégé.

Peter Winstanley

UN classification schemes.
Various vocabularies.
Allow development of cross mapping between government administrations.

Mostly internal currently. Moves to bring externalizing data into the 21st century.

Peter Murray-Rust

Fight for your Ontologies.
Ontologies in physical sciences. Chemists don't want ontologies. They'll sue you.
Crystallography uses 'dictionary'. Written in CIF. 20 years to build CIF.

Compare physical sciences to government.

Every program author writes dictionaries that work for them. When different parties agree, promote to communal dictionary. Provide conventions to help disagreements.

Show a company can do it as opposed to a rabbiting academic ..

Jeff Pan

Tractable ontological stream reasoning.
Need to be more efficient, scaleable, as things change. Inputs from web.

Dealing with complexities: approximate owl2.
Dealing with frequent updates: to-add stream and to-do delete stream. Truth maintenance. Evaluation criteria.

Trowl.EU can use with protégé, also supports jena.

Edoardo Pignotti

Semantic web tech to support Interdisciplinary research.
ourSpaces VRE
Provenance crucial.
OPM prov ontology.

Deployed since 2009, 180 users. Comprehensive ontologies but people unwilling to provide metadata.
paper! Edwards et al. ourSpaces.

Tom Grahame (BBC) @tfgrahame

Content arrangement on BBC sport by tagging, automatic to free up editors to write.
LD API so systems don't need to know about each other.
Growing from simple rdfxml to more complex ontology.
Can ask much more general and much more detailed questions about sport.

Mapping incoming data is outsourced.
Lots of errors, sometimes system alerts, sometimes manual.

Working on opening the data. Maybe a dump, but licensing issues.

Ewan Klein

Mining old texts for commodities, adding place and time and putting in structured database.
Transcriptions of customs import records.

Skos for synonyms.
Dbp concepts.

Why? Want to query.
Visualisations.

Tools? Python script.

Janice Watson

Harnessing clinical terminologies and classifications for healthcare improvements.

Bob Barr

Geographical addressing.
Addressing and address geocoding is important and broad. Not always postal, but this not addressed (punlol) in ontologies.
Different contexts change meaning of address (for delivering, you only care about postbox; property sale whole building).
Loads of things to address. Loads of reasons why.
Work held up as national address file is owned by royal mail and might be sold!

Fiona McNeill

Run time extraction of data. Failure driven. Looking at extraction of specific information.
Emergency response. Lots of data, timely sharing of data required.
From domestic level to humanitarian disasters.
How can it be automated?
Multilayered incompatibility.
Format
Terminology
Structure
...

Richard Gunn

Towards an intelligent information industry.

Elena Simperl (Soton, sociam)

Crowdsourcing ontology engineering.

CSrc: Brabham 2008.

Distribute task into smaller atomic units.

Humans validating results that are automatically detected as not accurate.
What are the costs? What resources?

Games with a purpose. Like quizzes.
Micropayments or vouchers.
MTurk. CrowdFlower.
Paper about useage of microtask crowdsourcing.  ISWC 2012.

Claudia Paglieri

Ontologies in ehealth.

Enrico Motta - Rexplore
Klink algorithm mines relations between research topics.
Use this!  Nope, it's not public.   Uees MS Academic research.

Peter Murray-Rust

Content mining expands regular text mining.
Focus on academic stuff.
Chemical Tagger. Takes chemistry jargon and annotated it, knows actions, conditions, molecules etc.. NLP. Uses ontologies and contributes to ontologies.
In chemistry,  no need to put everything in rdf because there are already lots of formalisms.
Proper cool PDF to sensible format conversion. Amy the kangaroo. Looking for collaborators.

Yuan Ren

Ontology authoring in whatif project.

Reasoning with protégé and trowl .

Tractable reasoning. Trowl v fast.


Notes from conversations / breakout discussions:

BBC use owlm triplestore  .
Store all their datasets in svn. But they have reads and writes to the live triplestore all the time.

Lots of people saying minimise owl use because of unpredictable output.

Versioning ontologies (available in owl2) in case third parties change stuff you use. You're dependent on their software engineering practices. Only good if they're ahead of the game.

IRIs, Arabic characters in ontologies!
Semantic heavy, maybe make a decision to abstract away to ids and make heavier use of labels.

Difference between importing and using someone else's.

There's no (practically useful) software that lets you reason over stuff you haven't imported? (over HTTP?)

Build ontology from reality (data), don't start with no data.

Lode.

Problems with dbpedia URIs changing or disappearing.

Hard to visualize massive graphs. Relational, tabular much easier to understand.

Thursday, April 04, 2013

Inspiring and empowering: The Lovelace Colloquium, Nottingham 2013

In 2008 I was in my first year of university, and the second ever Lovelace Colloquium was held in Leeds. I was encouraged to attend by Professor Cornelia Boldyreff and then-PhD student, now-Dr, Beth Massey. Doing so may have changed my life.

At my first Lovelace, I was introduced to the very concepts of conferences, mentors and (importantly) networking. The event was, and has been ever since, a forum for thought-provoking technical talks, inspiring motivational speeches and stimulating discussions about technology-related disciplines, careers, and womens' role within this world. To attend Lovelace is to be surrounded by extraordinary and excited minds; undergraduates at the top of their game, and successful academics and industry professionals to advise and mentor. Having now been along as an attendee, a poster competition entrant and for the past two years as a judge, the conference has provided perfect annual milestones to mark my own academic progression and personal development. I have met so many wonderful people and made so many important connections thanks to this event that I genuinely think I would be in a different place today, perhaps as a different person, had I never been. I can trace back directly or indirectly to one or other Lovelace Colloquium many of the opportunities I have had to develop academically (poster presenting, inspiring conversations), professionally (networked my way to a Google internship) and personally (overcoming low self-confidence, understanding imposter syndrome and conquering public speaking).

This year's, hosted by the School of Computer Science at Nottingham University, has been no different.

For the first time ever I arrived with time to spare before registration, and got to know some of the other helpers and attendees. I was put in charge of organising posters, directed towards a room containing lots of large fuzzy blue boards, divided up the space based on the number expected in each category (First Year, Second Year, Final Year, and taught Masters) and cheerfully handed out drawing pins to entrants as they arrived.

At 10 the crowd who had gathered in a lecture theatre were welcomed by the superhuman Dr Hannah Dee, and the first round of talks began.

Instantly relevant (to me), Natasha Alechina discussed work on logic in ontologies. The use of logic can help with debugging when creating new ontologies by detecting inconsistencies (eg. fallasies, contradictions) or incoherance (eg. empty sets). The method they use is to compute a minimal set from a big graph in which nodes are statements, and they keep track of where all the statements are derived from. It was "surprisingly fast" when tested with 1600 large random ontologies, compared to state of the art methods to compute minimal sets.

Logic is also useful in ontology matching, for example Ordnance Survey vocabularies versus Open Street Map. Logic helps the process by finding what might need to be changed or removed, but human intervention is needed to make the final call.

Next up, Jemma Chambers turned out to be a brilliant speaker and surely inspired everyone in the room by telling us how she'd made the most of a career in technology over the past decade. She was in her last week as a CISCO business development manager, about to move to a similar role at Virgin Media.

She started with some statistics:

  • 51% of gamers are girls, but only 6% of those who make games are female.
  • 21% of jobs in technology overall are held by women.
  • Companies with women in their management report a 34% return on investment over companies with only males.
  • 20% of C-level (CEO, CTO, CIO, etc) leaders worldswide are female.

(Disclaimer: I may have botched the context of those stats slightly, my notes aren't very clear. But you get the idea. Also she didn't say where these stats are from).

Jemma did a year-in-industry during her degree, programming for Oracle. She was bored out of her mind coding (I'm sure some people in the audience sympathised, but probably a minority) and thus learnt what job she didn't want to do when she graduated. Instead, she joined an accounts management graduate program at CISCO, had some doubts but stuck it out, rocked hard in sales and climbed the ladder through hard work and force of will, despite various sexist or ageist behaviour directed her way. A key point here is whatever you end up doing, do it well; being successful wherever you end up opens doors to what you really want to do, if you're not already there. Especially in the big tech companies like CISCO, where moving between jobs internally is facilitated and even encouraged.

On a related note, Jemma talked a bit about the flexibility of CISCO (and other similar companies). Working hours, for example, are yours to choose so long as you get the job done. Similarly she's had no problem negotiating maternity leave, and eighteen months after the birth of her son she's working three days a week (and still feels guilty about dropping him with the babysitter).

Naturally she mentioned a few (legitimate) generalisations about women in the workplace (nothing I haven't heard before, but this is my fifth Lovelace) and followed them up with some solid advice. Women seem to attribute success to outside forces like luck, or kindness of others, where men attribute success to themselves. It's much easier to move forward if you remind yourself that you worked hard for this and deserve it.

Successful men are more likely to be percieved as likeable than successful women, who are often construed as bitches. Ignore what other people think, and don't let yourself get walked on to try and make friends. At the same time, don't let this stereotype go to your head; remember to support other women in the workplace rather than being competitive.

Women and men have different leadership styles (generally) as well as other strengths and weaknesses of their own, and it's a combination of the two that really make a successful team, not more of one than the other.

Jemma recommends reading Lean In by Sheryl Sandberg.

She also discussed the various merits of networking (of which I am happy to attest there are many!) and how to source mentors in the tech community.

This talk was a fantastic one to start the day with, especially to prompt any in the audience who might otherwise have not done so, to talk to everybody. Jemma's enthusiastic speaking style will have kept everyone engaged, too, even those still waking up.

Dr Julie Greensmith filled us in on her journey from a pharmacy undergraduate through to her current work on artificial immune systems. These are algorithms inspired by human immune systems; robust, decentralized, adaptive and tolerant. They work by knowing what is normal instead of what isn't, which is particularly useful if you don't know what attack is going to come next. Their early work, though excellent, was based on a rudimentary computer scientist understanding of how immune systems work; these days they have a more interdisciplinary team with biologists to improve things even further.

Gillian Arnold, who is exceptionally well known and officially recognised as An Inspiring Woman, was filling in for a speaker who couldn't make it. She talked through the best career moments of various people she knew, which ranged from getting software into the hands of the public to promotions and financial incentives. She also talked through a few of the stereotypical problems women have in the a male-dominated workplace, but most of what she could have said had been covered by Jemma. A pro tip for getting attention at meetings if you're being talked over is to bang the table.

Dr Hannah Dee gave us a technical talk about her current research, as well as a little background on how she got where she is. She is much happier as a lecturer as opposed to a post doc, as she gets to direct her own research areas, and isn't constrained within fields she's not totally comfortable (like surveillance). So now she's interested in time and change in nature, doing things like laser scanning and time lapsing plants to find out new things that are particularly hard to find out. Some really interesting stuff about camera hacking with the Canon development kit, which lets you write programs in Lua or BASIC, and provides the sorts of menu options you'd usually only find on a really expensive camera.

Milena Nikolic is an engineer at Google London who has worked on Google's mobile sites, integrating results from mobile app stores into search results and the Android Market / Play Store. She says she has undergone a "journey of scale", and loves shipping projects that make a real difference and are used by real people. She answered lots of questions about working at Google. As with Jemma's experience at CISCO, hours are flexible at Google, and there are no strict iterative phases for development, but projects have their own cycles. She doesn't spend as much time coding as she'd like, but this varies depending on the stage a project is in, too.

Then someone asked "why are girls scared of coding?" and a lively discussion ensued. For some reason I didn't take notes, but things I can remember that were suggested include:

  • Girls are more hesitant about diving in, or scared of breaking things. To progress with programming, you've gotta just keep trying and failing.
  • Girls are more emotionally affected if their code does fail. Guys just shrug it off and try something out. (I personally have never felt like this).
  • Girls are less likely to be exposed to programming or programming-like activities at an early age, so by the time they come across computer science they may see it as boring, too mathsy or not creative. I suspect that had I not got interested in making websites aged ten, it might have passed me by during high school, and I would have ended up doing chemistry or French at university.

There were more; I'll add them if I remember.

In between these fantastic talks were coffee, lunch and networking breaks and of course, poster judging. I teamed up with Milena Radenkovic to assess the second years, and after three quarters of an hour of lunch, plus a 'last minute' extra half-hour before the decision had to be made (thus I missed the panel discussion), we had narrowed it down to five... It was hard. Seriously. We discussed the poster content, presentation, practicality of the ideas, whether the student was showing a project they were personally involved with or intending to do (this holds weight with me) and how well the student explained their ideas in person. They were all brilliant on all counts. We negotiated splitting the second place prize in two, but still had to choose three out of our final five.

Eventually we settled on Carys Williams (quantum cryptography; University of Bath) for the first prize, and Heidi Howard (routers that pay their way; University of Cambridge) and Jo Dowdall (smart tickets; University of Dundee) for joint second place.

I only wish I'd had time to look at the rest of the posters!

I finished off the day by joining other attendees for dinner, which was all round brilliant, and resulted in a late night.

See other fantastic blog posts...