Amy Guy

Raw Blog

Wednesday, January 16, 2013

1st International Open Data Dialogue, Berlin, 5-6 December


Read my complete notes from day one, and complete notes from day two.

The 1st International Open Data Dialogue in Berlin in December was broadly a discussion about real-world applications of Open Data.  Lots of practice, less theory.  Despite this (or perhaps because of this, now I think about it) it wasn't as technical as I expected.  Felix Sasaki [1] talked about some basic technicalities of Linked Data and the Semantic Web, kind of the first things you'd learn if you were studying it in a structured way, and I heard a lot of people afterwards complaining that that had been too technical.

Importantly, there was a real message of getting things done at this event, and plenty of evidence that a world built on Open Data is not an idealistic pipe dream, but a reality right now.  Challenges are being articulated, and solutions are being created, and problems are being overcome.

I stress this particularly because a couple of sceptics who weren't at the conference tweeted things along the lines of "Sounds like your conference is a bunch of idealist hippies preaching to the choir…"  A genuine concern, but what's really exciting is that this definitely wasn't the case.  It was instead a bunch of realist technologists with the expertise and influence to actively overcome barriers to improving the world.

Open Data is about social change and empowerment.  It is about accountability of organisations with massive influence over the lives of ordinary people.  It is not about an abandonment of personal privacy, or everybody knowing everything about everyone else.

It should go without saying (yet it still needs to be said) that it is not appropriate to blindly make all data available to everyone about every aspect of everybody's life.  But what if you had access to all of the data anyone had ever collected about your life?  Think about purchase history (shop loyalty cards, travel tickets), online activities (searches, browsing history, social networking).  All this stuff is being stored anyway, all over the place.  Often by organisations who fully intend to profit from it, presumably with your unwitting consent.  They went to the trouble of collecting it, but you went to the trouble of providing it.  It's your data too.  What could you do with it (or hire a software developer to do with it)?  Then imagine you had access to the same data from everyone in your town, aggregated and anonymised, and visualised in a nice way.  Maybe you could team up with your neighbours for cheaper bulk food purchases?  Maybe you'd realise that others had similar hobbies or problems nearby, and could form special interest or support groups?  Reduce costs by sharing transport to similar destinations (or just have some company on the journey)?

There's so much potential within data that's already held.

The UK government's Midata initiative is a massive step in the right direction [3] toward compelling commercial enterprises to hand over machine-readable datasets to consumers upon request.

In Slovakia and Kenya (and possibly others, but these were the ones that came up), there is a constitutional right to data held by the government.  Not without loopholes and other problems, of course [5, 2].

One of the obvious problems is convincing large organisations that hold lots of data (like commercial enterprise and governments) of the circumstances in which it would be in everybody's best interest to release (some of) it.  Reasons they don't include a lack of understanding of the benefits; disproportionate assessment of risks; aversion to change; a lack of technical expertise and infrastructure; "data hugging syndrome" [2]; licencing issues; outdated business models.

Nigel Shadbolt's experience says that large organisations who open data always see benefits.  It's always worth the effort.  When the data is there, suddenly developers start doing things with it; applications appear, many unexpected, and usually free.  He stressed that it's important to have a stockpile of success stories in case you need to convince someone in charge of the value of Open Data, and his favourite one was the publication of MRSA rates in hospitals (resulting in sharing of good practice, and an 85% reduction in MRSA over two years).  See a list at the end of this post for all of the success stories I came across over the course of the two days.

There were lots of discussions about the users or audiences of Open Data, and the various different roles people can have.  Most consumers of Open Data are developers, and 'ordinary people' see the data via an application.  Many won't know (or care) about the source of the data that powers the app, even if it about them.  Many will, and trust must be built for people see the value that such apps could bring to their day to day lives.  Ideally, releasing a dataset would be part of an ecosystem, rather than a one-time thing.  Data providers should value consumer feedback, and commit to good quality, up-to-date data.  Rufus Pollock wonders why every dataset doesn't have a public issue tracker, and notes that poor quality data creates wasted time, especially at hack events [4].

A successful Open Data world needs partnership between the public, media and organisations.  All of these parties need educating on appropriate combinations of the realistic potential of Open Data, and the technicalities of releasing and using it.  Michael Hörz [6] discussed the journalist perspective on Open Data; they're desperate for data about everything, and often manage to get hold of it.  But they find themselves begging for spreadsheets or CSV files, because what they get given are PDFs.  Eugh!  Yet they're not asking for Linked Data formats?  Which means, presumably, that after they've been through the trouble of extracting data from PDFs, they're putting it in a spreadsheet or something, and there's still a whole level of usefulness missing.  And I assume that's because they don't know otherwise, or perhaps don't have the resources to learn even if they're aware of the possibilities.  Similar sorts of reasons that they're being given PDFs by organisations in the first place.

So awareness, and easily digestable educational resources (how about SchoolOfData.org) need to be promoted.

Now then, about those success stories...  This list includes data publishing projects, groups and apps that have been built on Open Data.

That'll do for now.  Lots of the portals and competitions have links to app examples etc.  There's lots to explore.


Finally, I highlighted in my notes quite a lot of things that I need to find out more about.  A lot of them are technology or platforms for publishing or sharing Open Data, and various standards or studies I need to read in more detail.

I have a couple of questions to ponder on, too:

There's a massive focus around hacks (more often than not one off events) as a way of using and promoting Open Data.  What other ways are there?  What will the path to a deeper integration of Open Data in society look like?

There are lots of datasets and vocabularies about public services and society, as well as science and education.  What arts, culture and media datasets are out there?  (And what has been done with them?)  Ooh, or online social interactions?  Maybe I'll do a survey.

[1] Prof. Dr. Felix Sasaki, keynote: "Linked Open Data @ W3C-Vocabularies, Working Groups, Usage Scenarios."
[2] Prof. Dr. Simon M. Onywere, talk: "The Kenya Open Data Incubator Project – Outreach to Research Community."
[3] https://www.gov.uk/government/consultations/midata-2012-review-and-consultation via Nigel Shadbolt
[4] Dr. Rufus Pollock, keynote: "Open Data, Building the Ecosystem"
[5] Peter Hanečák, talk: "Open Data and Open Government Partnership in Slovakia."
[6] Michael Hörz, talk: "Open Data in Local Journalism: An Excel file?"

[Notes] 1st International Open Data Dialogue (Day two)

December 6th, 2012.

Notes as I scrawled them.  Read a proper review.  Purple is calls to action for myself.

Rufus Pollock - Open Data, Building the Ecosystem


Open API is a contradiction.  
An API is not open (though still valuable) but not the same as bulk open data.
Should we open governments?  So much is hidden.
Ultimately a lot of it is about in/justice (debts paid off by younger generations).

Open spending
- only 30% of departments are up to date - can keep checking who's updating their stuff.
- but sometimes more up to date that government's own records - they use it themselves.

Getting the data is only the beginning.  Need a platform.

Getting citizens, journalists and government to work together.

Why doesn't every dataset have a public issue tracker?  This is what he'd like to see most.

The process is important.

Digital doesn't run out.  That's why open makes so much sense.

FourSquare uses OSM.

Average consumers won't see raw open data - but via products and services.

Next?  Building communities.  Takes time to nurture.  publicdata.eu.

We'll go from hackdays to deep integration.  Toy vs. core datasets.

Geodata is mature.

SchoolOfData.org.  Data analysis, programming, training.
- Teachathons better than hackathons.

Should have access to all our own data - travel (Oyster), shopping (loyalty cards), social.
- Would like to control who uses it.
- Compare with others (aggregate).  Only companies can do that.
- Some people will choose to publish/open their own data.  Even if it's a few people, still lots of data.

Open Data = Platform; !Commodity
- Build on it rather than sell it (and let others build on it)
- Already seeing people building companies/making money on this principle.

Ecosystem.  Need feedback.  Poor quality data creates wasted time, especially at hacks.

Datagov census dashboard - who has what out.
There's no gov. data that's come back cleaned up from the community (even though people are almost certainly cleaning this data up).

Be patient.  Needed to wait for pressure to build up (why it's happening now).  Low investment.

Way to fund production of gov. data (or combination of):
- Charge users
- Taxpayer revenue
- Charge creators (most attractive)
  eg. Companies Register data; more efficient to charge people registering companies a bit extra, people won't be deterred from registering a company because of this.  Basically, add to existing fees for things people will pay anyway.

Open data saves lives!  Heart surgery data caused improvements.

MiData (more from Nigel)
Supermarkets not very good, but some big companies involved.

Prof. Esteve Almirall - Reinveting Cities - Open Innovation in the Public Sector

Competition now is about innovation, not money.

Adopt a Hydrant (Boston)

Peter Hanecak - Open Data in Slovakia

2012 govt signed open government partnership. 

Laws already align with Open Data principles. 

Everything must be public, unless stated otherwise (eg. according to a specific law, like army secrets. Must be clearly defined, no rubbish excuses for not publishing something). 

Open licences (GPL, CC) are not recognised, because there's no signed paper. Currently they're campaigning to fix this. 

Existing laws do not fully apply to regional governments, only to state government. 

Data that is available isn't in great formats. Lots of things are ignored or misunderstood.

Achievements so far: - Datastores and apps. Datanest.fair-play.sk, znasichdani.sk, cenastau.sme.sk, otvorenezmluvy.sk, data.gov.sk. - app competitions, conferences. Restart Slovensko.

Future
- Spread the word 
- Advance principles on regional levels 
- Major release of OD by gov 
- develop OS publication platform for OD 
- incorporate results into other projects

They have a standard published for state government.

Ivonne Jansen-Dings - Code4EU and Apps for Amsterdam

Technology is a key component, but impact on citizens is central. 
Taking linked OD out of academia and applying it to 'real life'.
Tourist one. FairPhone. 

Collaboration between government, coders and citizens. Can't do this top down. Creating something together, that has to evolve naturally. There isn't a formula.

apps for Amsterdam is a platform for people to talk about OD; creating it and working with it. 
  • Lots of reasons why coders participate. 
    • Hobbies
    • Solve local problems
    • Network, get new business
    • Ideas start during hack events/workshops. 
  • Past two years gone from 22 - 130+ datasets. 
  • Quality of data is essential for good apps. 
  • Mostly about getting the government to see the necessity of open data. 
  • All sorts of working groups arising with municipalities to solve specific regional issues using OD. 
  • Helping people who have made/are making apps, to help them move forward. 
  • Lots of people are stuck because they need certain data. 
  • Creating ideas for apps is a good place to start with deciding which data to open up. 
  • Essential for app developers is also getting paid. Necessary to help people evolve apps. So a4A helps to connect developers to companies, NGOs, etc. for whom the app is relevant, so they can work out a way to work together.
  • They don't have much government data about spending and stuff, working on that.

Apps for Democracy (Apps Voor Democratie) 
  • Can see which parties and people collaborated with each other on getting motions put in and passed etc. Now integrated into actual parliament website. 
  • PolitFutures. Stock exchange around politicians.
Code4EU
  • Currently hiring developers to create solutions for problems they see within municipalities. 
  • Changing government from within. Civic innovators.
  • Lots of people come with problems, not ideas for apps.
Not about the apps, about the social change.
Local issues can have solutions that have global value.

Lena-Sophie Mueller - Open Government is more than Open Data

Stuttgart local people had issues with train station and there was a huge protest. Planning happened in a black box, people didn't understand why decisions had been made. 
(Ed trams)

Same as ACTA. People negotiated about it for 7 years, but it was not transparent.

Need transparency to face obstacles of 21st century. 

IT-Planungsrat say open govt needs to be a focus point, but will focus on open data first. 
  • Open goverment though, is more than just open data. 
  • It's hard to keep information secret now. Things get leaked. To counter leaks, they put it online too, so it's a trusted source. Can publish accurate new versions. 
  • What's needed is a Diff, so information needs to be in machine-readable formats so documents can be compared. 
  • Then politicians can be asked specific questions about changes.
Open gov means taking contribution of people as valuable. 
  • US crowdsourcing patent information for making decisions about patents. 
  • Open Budgets - people participate in deciding how budget is spent.
  • Governments and administrations collaborate. 
  • barnet.gov.uk. PledgeBank ("a site to get things done")
Lots of administrative workers aren't used to working with data. 
  • Need help with extracting data and organisation and processes and technology. 
  • Change management. 
  • It's going to take some time, but it's worth making the journey.
Convince organisations that it's okay to open data (sans personal information). 
  • Lots is digitalised already, and expensive to digitse paper. 
  • Smaller municipalities are mostly paper. 
  • eGovernment projects are important for open data.
Status of open government in a country? The thing Rufus mentioned. 
Initatives that measure freedom of information laws, but it's very difficult to measure openness. Need to develop a measurement that could be used internationally (big complicated project). 
Web Foundation published an index a few months ago. 

What does she use?  Usually use CKAN.

Christoph Lutz - Open Data and Social Media

Social medai readiness in Hamburg. 
Many insights from social media can be transferred to open data. 
Social media adoption and readiness can take place on several levels 
- organisational level 
- individual level 
- societal level

This project looks mainly at organisational and individual. 

Organisational: 
  • different agencies, like culture and financial services; regional agencies. 
  • structure, leadership and culture

Individual: 
  • drivers and barriers to social media adoption: cognitive (know how) or affective (acceptingness of technology, concerns)
Some agencies in Hamburg have started intiatives (like fb, twitter). Some were recalled, but in general there is political support and Hamburg is more advanced than other cities. 

Their research project: 
  • conduct interviews with employees who had contact with social media. 
    • 8 people 
    • analaysed successful vs unsuccessful social media use 
    • different levels of responsibility, and different agencies 
  • case studies 
  • soon a big quantative survey
Organisational factors: 
  • political support (crucial) 
  • leadership support (some people high up in hierarchy aren't familiar with technologies) 
  • autonomy and trust (people can experiment, be proactive) 
  • structures of organisations (complicated, different motives and experiences in different departments, hard to coordinate) 
  • processes (hierarchy and bureaucracy; you need fast feedback for social media, which contrasts with usual way of work) 
  • resource (stressed most in interviews; employees don't have time at work to administer social media accounts, sometimes IT resources)
Individual factors: 
  • age (younger people more interested) 
  • affinity (how much people enjoy working with IT; intrinsic motivation) 
  • experience 
  • social capital 
  • concerns (privacy, security, technostress)
Identifiable strategies: 
  • avoid resistance (make projects appear small, non-invasive, simple) 
  • externalise project (work with other organisations, avoid bureaucracy)
Engage people in open data via social media.

Objective of Hamburg project? Mainly about representing administration to the citizens, and providing feedback to people. Later about engaging people in conversations and participation.

How to make use of social media in administration? How to get public offices to produce open data? The quantitative questionnaire will help show how open people are to social media.

Simon M. Onywere - Outcome of the Kenya Open Data Consultative Forum - Kenya’s Strategy to Make Government Data available to Communities

Increase transparency and accountability of government. 
Help people make decisions. 
Support economic development in the country. 
Supported by World Bank.

In Kenyan Bill of Rights, citizens have a right to the data held by the government.

opendata.gov.ke 

Media plays a very important role.

KODI - Code4Kenya:
  • Lack of certain data
  • lack of metadata
  • limited search
  • data duplication. 
  • Need for better analysis and visualisation tools. 
These observations allowed holding a stackeholders forum.
  • People interested in solving problems that face the country.
  • (MANY issues; things Europe was facing 100+ years ago, but no data about what's going wrong, what the impacts are likely to be). 
  • Interested in building a platform, but the content is important, and what is it supposed to help us do? 
  • Issues about data collection. 
  • In the forum, ended up talking about issues that face the country, rather than the data.
Challenges: 
  • data hugging syndrome: 'this data is mine' 
  • Lack of Freedom of Information Act 
  • Slow digitization 
  • Lack of trust and low culture of openness
Demand for open repository of all Kenyan PhD and Masters theses. Needs metadata, and needs to be widely accessible. Will help with research, and help people understand research that has been done that might be able to solve problems. Help avoid double research.

Not really free information, because most public sectors spend money to get information.

Project requires various degrees of collaboration. Countries in Europe that are one or two steps ahead. Need training for socio-economic transformation.

Are the licensing issues being sorted out? There's a comprehensive statement on the website (so a custom license?). 
Citizens will hold government accountable for information provided, so they (gov) are worried about data being inaccurate. 
Letting research community use the data can help get useful feedback about what is wrong - you don't know if things are wrong until you try to use the data. 
Up to people to give the government the correct situation on the ground, because gov't is not always right.

Nigel Shadbolt - Finding the Value in Open Data

Local data matters. 

Open Data Institute - build economic value in a serious way. 
2005 AKTive PSI was early beginnings with getting various bodies imagining opening their data. Nobody would really give them the data at that point, but were curious about what could happen. 
Reported to parliament in 2007 - said it was exciting and had great promise, but nothing else. 
Activism, top-level political will and committed individuals needed to move things on.

Two years is not a long time for really disruptive movements like open data. Natural lifecycle to processes. 
Early enthusiasm, but wall of people who question impact, value, if it's worth the effort. We can see now that it's always worth the effort, and doesn't cost much.

Government is becoming more comfortable with this stuff. 
"Can we put a government website up that says 'beta' in the right hand corner, and not be ridiculed?" 
Gov't got used to the idea of agile web development; of not knowing how the system would be finished when it's started.

Suddenly, applications appear. Many unexpected. 
Gov't efforts would be more expensive and less effective.

Virtuous cycle: open licences - open standards - open source - open data - open participation.

CC isn't for everyone... companies worried about giving rights away. In the UK there's an open government license, developed by government lawyers (to make officials feel more comfortable, and understand conditions).

Open Data market needs a steady stream of successes. Always have a story the person in charge can understand. 
OD is abstract principle. 
He uses MRSA. They began to publish infection rates in hospitals. Two years later, it's down by 85%. Worst hospitals can look at ones that are doing better. People started to ask questions; simple procedures implemented (sunlight, disinfectant).
We have to understand that companies are in the business of making money, and public services are in the business of providing efficient services.

Transport... Companies think it's more valuable to keep hold of the data than to have people on the transport - lolwat?

Visualising data highlights issues.

Quality over quantity. 
Routine to publish certain sorts of information, but there's other stuff that's really important. 
Needs to be found easily. 
Data portal needs lots of metadata (quality of content, what kind of links, how much). 
Every public data should rate itself on the 5* score card.

Open Data business models: 
It isn't enough just to publish. 
Need to build demand for data that you're supplying. 
If the data is poor or gets turned off, people will let you know. 
Data Marketplace. 
OD Apps (people think this naturally). 

Innovate - economic benefits for host, sponsor and developers. 
Developers innovate on behalf of companies. 
Build and maintain trust. 
Prove that you're doing good things, eg. where materials come from, or effects on environment. Various different kinds of open APIs.

Open Data needs a balanced and broad ecosystem. 
Complex. 
Not just gov'ts. Businesses are beginning to, citizens might eventually (think about social networks). 
Lots of varieties of open data AND closed data (some just cannot be released). Or personal data that only the owner can/should have access to. 
It's much richer than just "everything's open and we need to work out a way to monetise it".

midata: 
We will increasingly become aware that we can collect our own data. 
So why can't we get the data that other people collect about us? Why don't you have access electronically to every receipt - and what would that world look like? Switching suppliers, teaming up with neighbors for shopping. 
Energy providers in the UK (three of them, other three will do soon) give access to all their raw data.
It's hard to get data out of companies, but those who do see real benefits. Telephone companies in the UK are seeing increasing data exchange between them and consumers.
Products and processes that we'll see emerge from this are exciting.

What's the mix between open data and personal data (midata)? 
Government midata. Most people don't claim cold weather allowance. Costs a lot to get credits moved around. Open data meeting midata would benefit this. Same in health area.

Open Data Institude - theodi.org 
Leading the creation of the open data ecosystem. 
Trying to improve public supply. 
Training people to produce and publish open data. 
Incubating companies. 
Work with public bodies, big corporates, small startups, trying to find values in datasets. 
After 8 weeks, 4 companies working in their space. 
Locatable
placr.mobi is a transport API provider. High quality access to all open transport data. 
Mastodon - green cloud computing options. 

It isn't all sweetness and light. Have to demonstrate tangiable benefits to keep progress going. Have to give company/politicians good reasons why this is better.

Huge amount of capital value is based on "I know something you don't know". 
As information becomes abundant, the landscape will be changed. You can't rely on knowing something any more, you have to provide something extra quality. 
Drive innovation and improvement in service delivery. 
With heavy investment in acquiring certain information, why should they share it? 
Evolutionary arms race means that someone else will find a way to collect that data more cheaply. How long can they sit on their monopoly?

Governments are not here to become revenue generating businesses, but to provide public services.
In the UK, the office of national statistics gross value added figures are from 2010.
Hasn't seen any examples yet that cost more than the benefit that's gained. 
Hospitals, traffic data. These studies need doing more carefully than they have.


[Notes] 1st International Open Data Dialogue (Day one)

December 5th, 2012.

Notes as I scrawled them.  See here for a proper review.  Purple is calls to action for myself.

Dr. Philipp Mueller - Openness as a means, not an end
I missed the opening keynote.  His slides are here.

Dr. Wolfgang Both - One Year Open Data Portal Berlin

Open cities EU - Nov '10 - Apr '13.
Amsterdam, Barcelona, Berlin, Helsinki, Paris, Rome.
Berlin responsible for open data working group; several working groups (OD is just one).

EuroCities

Knowledge Society Open Data working group
Guidebook for cities, 30 pages so far.

Open Data Berlin
Portal Sept '11
Publication
Press conference Feb '12
Short term: Political agenda, budget, working group.
Mid term: Harmonize data formats.
Long term: Legal framework (Berlin can't decide laws by itself, for whole EU).

Open Data Day May '11, '12 and '13 in prep.

WG open traffic hack (29th Nov '12)
- 150 programmers with transport data.

Portal stats
Are users interested?  Peak at start.  Other peaks for hacks, Apps4D contest (Nov '11)
Possibility for feedback - questions, advice, ideas.
100 datasets; daten.berlin.de

WG 2012
- formats and metadata
- licensing and user rules
- education for staff (lectures, this is new for many working in public sector)
- organisting and processing
WG 2013
- evaluation of OD studies
- Recommendations
- Exchange with other cities.

(Q&A)
Datasets were volunteered, not selected
- But are looked at for quality, machine readable, looking for wide range of topics of interest to public.
- Want open, transparent process for publishing.
- Communicate with media as well as community.

Licenses
- Heuristic... no legal advice available because it hasn't been done before.  Many possibilities; for opening data for individuals, CC was familiar to Internet community, includes origin data (CC-BY).  Some smaller datasets are licensed for non-commercial usage.  Discussion still ongoing.
Knows other cities will follow / copy there example whatever they do!

Jan Schallabok - Right to Freedom of Information on Enterprises


Call for open enterprise data.

Why?
Scenarios, set a timeline:
2015 - personal search
2016 - data disasters (identity hack)
2017 - pictures omnipresent, know all about everyone becomes normal
2019 - Google Glass on market

If there's no data on you, things don't work (eg. personalised advertising)
Society down the drain if it didn't open data (Switzerland in the story didn't open data, so Swiss woman moving to German couldn't settle in easily).

Moving away from clear facts towards probabalistic.
eg.
Google Translate fed by open (input) data, but algorithms aren't open.
Siri - enriches dataset from Web (he said Google search?)
OpenStreetMap (counter example)
If there was more data, everyone would use OSM paradigm (eg. government).

Harm businesses?
..maybe.  But more damage in the long run?

Need to make businesses move away from using peoples' data.   Like, we work for facebook.  our data, not theirs.

Data protection:
by law, data subject has rights to know logic involved
but
as long as it doesn't affect trade secrets

Privacy implications
'Anonymous' data can be used to identify people.  See AOL search database fail.
When can datasets go public?
Weather data can be personal data (no time for example..)

Michael Hörz - Open Data in Local Journalism

Journalists expect everything
- spending
- political decisions
  (district levels, searchable)
- quality (schools, pollution, food)
- real time sensors (air quality, traffic, energy)

- Open Data Paris (loads, on a map)
- Locrating (school performance on map, UK)
- Chicago bike crash reports (map) (sort by injury, date, day; data all open from Chicago Transport Authority, in a nice format).
- LA Times LAFD (fire dept.) response times.

- Airplane noise map, taz.de (Journalists had a PDF, eugh).  One of the first Berlin interactive visualisations.
- Berlin election.  Was real time.  Down to the polling station.
- Berlin bicycle accidents '11 - came from massive PDF (3,800+ cases)

- Wishes for xls or csv... wants directly processable.  WHY NOT RDF?!
- ..or APIS.  WHY NOT LD?!
Reality = PDFs, requests ignored, data incomplete or hidden.  Hard to get for journalists.
All datasets are interesting and should be out there.  In Berlin often only one or two districts are available, which is no good.

(Q&A)

It's not always straightforward just to release data - need priorities; raw data/API documentation isn't always available straight away.

Why PDFs?  They don't know any better.  Need to make people aware.

Consequences of public seeing data they're not used to?  Panic?  Or activism?  Pressure politicians for change.  Empowers people.

Is there are resistance to making data available (eg. Italy - data there but useless).  Maybe, or maybe they just don't realise [it's useless].

Prof. Felix Sasaki - Linked Open Data @ W3C-Vocabularies, Working Groups, Usage Scenarios

== first half of MASWS. 

New work on LD http://www.w3.org/2012/ldp/wiki/Main_Page 

Media fragments - spec finalised www.w3c.org/TR/media-frags/ 

Ontology for Media Resources (DC for video and audio?) http://www.w3.org/TR/mediaont-10/ 

Internationalization Tag Set 2.0

SW core is stable, so work with vocabs now. Need interoperability. Decide: - Syntax - - Microdata not necessarily for SEO - - Schema.org; w3.org/wiki/WebSchemas; very basic schemas with increasing numbers of more specialised extensions. Discussion at lists.w3c.org/Archives/Public/public-schemas - 

Application scenarios.

Organisation ontology - Membership and reporting structure, location information, organisational history - Interoperable organisations - 'Final call' stage - nearly done. Need feedback.

DCAT (interoperability between data catalogues) - Uses FOAF, DC, SKOS
w3.org/TR/publishing-linking draft Namespace neutrality - xmlns.com

Language graph of the Web is cool.

Tomáš Knap - Tracking Data Provenance of the Published (Linked) Open Data

xrg.cz
opendata.cz
watch film

Defines provenance and agents, artifacts, processes. Provenance useful for data integration. Which is right/recent etc. 

How to cite. Vocabularies: PROV-O (almost w3c final, w3/ns/prov), VoiD (datasets w3/TR/void), FOAF, DC ODCleanStore - http://iswc2012.semanticweb.org/sites/default/files/paper_37.pdf (prov aware storage, processing, querying) - Write rules/queries using web front end. Certain automation from inserting ontologies. Still manual work.

LOD2 WP9a - EU project LD tools

Maria Magdalena Theisen - Open Data and Big Data

Big Data - have to ask questions to understand what questions to ask. Consists of volume, velocity, variety. Can't say that all open data is big data, and vice versa.

Some BD from external social media, disaster information, sensors, smart meter. Lots of things bringing data to process.

Facebook has largest data collection by 2010.

Open and Big - Eye on Earth - air watch (over 1k stations in Europe providing live data), noise watch, water watch (static historical data) - can rate quality of data and give attributes

Cloud computing is an enabler for B and OD - Don't have to manage servers to provide data. - Flexibility and scalability - Interoperability with existing infrastructure - Easy access to data - Development platform (Azure) - can enter an app with open data into marketplace. Lots of examples.

Is Azure marketplace integrated with ckan? - No.

Evanela Lapi - Building Sustainable Open Data Platforms

Understand stakeholders

* Consumers
* Developers
* Citizens
* Less technical, can use open data to help with life

* Journalists, scientists, researchers
* First two more critical
* disseminate data
* Need open, standards-based, non-proprietary formats.  Easy to download/browse/search/redistribute/share.


* Publishers
* Provide transparency
* Want a cost-effective, easy solution platform
* Public sector has lots of data not online - because it's hard to publish?
* lots of friction, fragmentation




Socrata - end to end, custom solution.  Many implementations in US, Kenya.

vs.

Integrated, loosely-coupled - existing SW, eg. CKAN + Drupal (data.gov.uk)
Faunhofer OD platform is Java (Amsterdam uses)

Open Cities
- open innovation (see last time this was mentioned)
- on Github - get feedback from use.

Virtuoso triplestore + Liferay CMS + CKAN catalogue
(Java wrappers for REST APIs)

User roles:
  • Data owner
    • Publish
    • Maintain
    • Bulk upload
  • Platform user
    • Query
    • Discuss
    • Search
    • Browse
    • Download
    • Propose new
amsterdamopendata.nl - 137 datasets in 18 categories
flevoland.....nl(?) - 22 datasets

It's a good start, but still not enough - why?
  • Too much manual work, redundancy across different platforms.
    • Modernise environment - by modular, high level stuff?  (I think that's what she said)
"Germany isn't much into OD yet.."

Oliver Adamczak - Big Data for Smarter Cities

Leaders must innovate to exceed citizen expectations.
Functionality of BD - use variety and volume to innovate.
Vision - do things you haven't thought of before.

I should do a survey of open data about arts/media?  It's all about gov/science.

IBM BD platform

Hadoop to store
- low cost (open source)
- scaleable
- easy to load data - don't have to care about structure until afterwards

Text analytics to read PDFs etc. and extract data with context.

Streaming data is important
- Not for repo, just use/analyse and discard.


Monday, January 14, 2013

Dynamic Web Design

For the second year, I'm tutoring PHP, MySQL, HTML, CSS and JavaScript to MSc Design and Digital Media students for the Dynamic Web Design course in Edinburgh College of Art.

Sometimes this makes me feel like a wizard.

Digital Media Studio Project (DMSP)

I'm supervising (created the brief, overseeing and guiding a group of 6 MSc students, grading) a Digital Media Studio Project in the Edinburgh College of Art.

I challenged my lot to use responsive web design techniques in a unique and inventive way.  You can see what they're up to here.

Wednesday, January 09, 2013

Morrissquirrel (crochet)

As a farewell present to Beth, who was vacating sunny Scotland for the harsh and unforgiving shores of the US, I made a squirrel.  But not just any squirrel.

A Morrissquirrel.

That's Morrissey-squirrel.  Don't question it.

I largely followed this pattern for the normal squirrel parts, then improvised to Moz it up.












Sunday, December 16, 2012

Week in review: Not much again

10th December - 16th December

Boy, this Week in Review thing really highlights when I haven't done a lot, doesn't it?  I guess that's the point.

I spent most of the week not getting round to writing up and analysing my notes from the conferences.

We had an Ontologies with a View meeting.

I made a few connections and sent a few emails.

I did lots and lots of useful, but non-PhD-related things.  Honest.

Saturday, December 08, 2012

Week in review: Conferences

3rd December - 9th December

I went to the 1st International Open Data Dialogue (#odd12) in Berlin on Wednesday and Thursday, and gave a lightning talk at Digital Methods as Mainstream Methodology (#dmmm2) in London on Friday.

Both were brilliant; I met a ton of interesting people, learned loads and gained much inspiration.  My extensive notes are in the process of being typed up and thought about, and will be published on here as soon as humanly possible!

Now I'm going to catch up on sleep that crack-of-dawn flights and many hours of train journeys have denied me recently.

Monday, December 03, 2012

Week in review: Not much

26th November - 2nd December

I read a paper about Semantic Web tools for online communities.  Large parts of this project are comparable to my own, I think, and I will probably keep going back to it for inspiration.

I did a lot of non-PhD related things, too.

Wednesday, November 28, 2012

Advice: willful misunderstanding, and audience as individuals

Picked up on some great advice during an afternoon-long course entitled 'How to do an Informatics PhD' a couple of weeks ago.

All through my undergraduate, and probably High School as well, I was told that when writing assignments I should treat the person marking it as if they don't know anything at all about the subject.  They're stupid.  Leave nothing unexplained.  Of course, we were often also told to 'keep things concise' and usually had to do this under the constraints of page or word limits.

Problematically, some people could have interpreted this as an opportunity to try to pull the wool over their marker's eyes, or baffle them with science.  Typically, the marker will know something about the subject, so that's probably not a great tactic.

I see where this advice comes from of course, and a couple of weeks ago I heard it phrased in a different way, that makes far more sense.

The examiner will be knowledgeable about the subject, but given to willful misunderstanding of what you're trying to say.

I think that gives a much better guide to how and when to explain things.  And also makes them more of an enemy to be conquered, than an inconvenient fool to be deceived.

And whilst I'm recounting advice I like, here's something that has stuck with me since 2010, from my manager at Google at the time.

When presenting to a large audience of people, don't think of them as a crowd.  Think of them as many individuals.  Make your presentation as you would to a single person.  It just so happens that there are lots of single people there all at the same time.

Since letting that sit in my subconscious, nerves before giving a presentation have shrunk to negligible levels.  It may be that over the past two years I've become more confident anyway, but I used to be thoroughly terrified of standing up in front of even a classroom of people the same age as me.  I wouldn't talk out in lectures for the most part of my undergraduate, and harbored a gut-wrenching fear of being picked on to answer a question.  As did most people, I imagine.

I make sure to actively observe how I feel when watching someone else present.  I look at other people in the audience too, and note the attitude and techniques of the presenter.  (Consequently I probably leave having no idea what the presentation was about).  The results are usually that a majority of people aren't listening properly.  A vast proportion certainly aren't angrily judging the presenter's every twitch.  Things people notice and get upset about include:

  • If they can't hear you.  
  • If you're just reading off slides, particularly if you try to act like you're not.
That sure is a short checklist of things to avoid.  There must be more that can go wrong.  Reasons people will stop listening, include:
  • If they can't hear you.
  • If you're just reading off slides, particularly if you try to act like you're not.
  • If they're not interested in what you're talking about. (Pro tip: make them interested).
  • If you're not making any sense at all.
  • If your slides are more interesting than what you're saying. (I really like presentations without slides.  So long as the speaker is engaging, of course.  Slides with lots of words are a definite negative, in my book though).
And some of my personal nitpicks include:
  • Drawing attention to a mistake by apologising for it.  From being on both sides of this situation, it usually feels right to do so at the time, but until you do, four fifths of the room won't have noticed, and the fifth that did will forget within the next few seconds.  Point it out, and everyone will remember.
    • If you say something wrong, just correct yourself and move on (but don't leave it uncorrected, this is usually noticeable).
    • If it's a technical problem, keep talking whilst it gets sorted.  This boils down to not relying on technology to keep your presentation interesting.  I am aware that there are some situations where this is impossible.
  • Weak intros and outros.  I've been guilty of both of these.  I plan to pay more attention to upcoming talks in order to fix this.  It's something I always forget to notice.  But for now:
    • Starting with filler words, like 'So, ...', 'Right then...' or 'Okay, ...'.
    • If you've been introduced, you don't need to repeat it, especially not in a way that draws attention to the fact you're repeating it. 
    • Make it clear when your talk is done.  Don't trail off with '...and that's that then.'  'Any questions?' is usually okay, but only if it follows a distinctive final sentence.  Jumping to that from what feels like half way through a paragraph is a bit rubbish in my head, but in all honesty will probably go unnoticed by the audience.  Except me.

Fortunately, my days of feeling agonisingly self-conscious whilst presenting are long gone.  On top of that, I find doing as little 'rehearsal' as possible boosts the natural fluidity of a presentation, and in turn my confidence.  If I haven't rehearsed, there's nothing to forget to say (and suddenly forgetting what comes next is the biggest killer of flow, something I discovered during French oral exams).  That only works if you're very familiar with the subject matter.  And if you're not, you probably shouldn't be presenting about it.


Disclaimer: I'm very early in my academic career, and haven't presented a whole lot.  The biggest audience I've talked in front of was about 120.  Despite the theory, I suffer from having neither a naturally loud voice, nor a naturally beaming expression.  So all round, I'm probably not very good.

Monday, November 26, 2012

Notes about Semantic Web tools for online communities


K. Faith Lawrence & Dr. Monica Schrafel (2007)  Amateur Fiction Online - The Web of Community Trust: A Case Study in Community Focused Design for the Semantic Web. Intelligence, Agents, Multimedia (IAM) Group, School of Electronics and Computer Science, University of Southampton.

NB. Need to read her full thesis, of the same name.  Will probably clear up some of the questions I scribbled whilst reading the paper.

Finding out if Semantic Web tools can be brought to hobbyist groups on the Web.
  • Uses the online fiction community; suggests they could benefit from:
    • improved searching
    • improved meta data
    • automatic recommendations
    • trust webs
    • personalisation.
  • A HCI project, so usability tests and comparisons with current systems are key.
Related work
  • Community centered design
    • to determine user needs - through continual interactions and user studies.
    • to consider how reader-facing apps present themselves and particular community
      • responsibility of being a portal - need clear affordences and points of failure.
  • Trust and Semantic Communities
    • The semantic web “provides a common framework that allows data to be shared and reused across application, enterprise, and community boundaries.” - Tim Berners-Lee, James Hendler, and Ora Lassila. The semantic web. Scientific American, May 2001.
    • Do we trust:
      • metadata
      • data
      • mechanism by which data is returned
      • person requesting data?
    • Many definitions of trust.
    • Jennifer Golbeck's trust onotology to go with FOAF (Jennifer Golbeck, Bijan Parsia, and James Hendler. Trust networks on the semantic web. In Proceedings of Cooperative Intelligent Agents 2003, 2003.): ratings of 1 - 9 for trust of associates.  Extended for FicNet (this paper) 
    • Here, trust: "the expectations that arise that an individual will not act in a way that is detrimental to another individual or community."
      • I might need to expand that for my stuff... maybe... maybe this will do.
Case study
  • Community predates the Internet. (duh)
  • Necessary to get opinions from people outside of the amateur writing community, because it is broad. Such as parents/guardians of members.
  • Questionnaire
    • General information
    • Reading habits
    • Community involvement
    • Access and distribution of materials
  • Questionnaire distributed by:
    • requests to archives to pass on to members
    • LiveJournal
    • emails to specific interested parties
    • mailing lists / bulletin boards of relevant special interest groups.
  • In two weeks, 1116 responses, from 30 countries.
  • Used to inform ontology design for FicNet, and OntoMedia.
Fan Online Persona (FOP)
  • Extension of FOAF, tailored for needs of online readers and writers.
  • foaf:person -> fop:persona
  • Separates environments for on and offline
  • fop:NomDe - context for name
  • Illusion of anonymity is fundamental to fanfic community (who are a large part of online amateur writers)
  • Most authors have one or more pseudonyms.
  • 80% said email address is the most personal information they should be asked for.  ('of the 80%, 15% said no personal information should be requested from anyone - does this make sense? Do they mean other than email address?  but that's the 80... If they're in the 80, they can't be part of that 15...)
  • Privacy is the main thing holding back FOAF (Joseph Smarr. Technical and privacy challenges for integrating foaf into existing applications. Presented at 1st Workshop on Friend of a Friend, Social Networking and the Semantic Web, September 2004.)
  • Personas aren't meaningless, because people become very attached to them, and only create new ones for specific reasons (says who? No citation..)
  • Expands foaf:document and foaf:groups
  • Creation, exchange and review of works is the point of these communities.
  • FOP dismisses FOAF info like work and school as irrelevant or potentially dangerous.
    • [Me]  I think the on/offline divide won't be so extreme for many amateur film makers (another story for consumers) because often their faces are in their movies... Also anecdotal evidence from my own experiences that I'm open to having proven to be a minority.  Actors vs characters is an interesting distinction too.  One amateur film maker can have many personas, even across one channel of output.
  • Options for FOP determined through long term study of metadata commonly attached to works. (Something I can do, too).
Trust
  • FilmTrust by Golbeck (just joined, it was closed last time I looked).
    • Could be prettier... but 2169 members!
    • Visualisations of the network - I need to get good at this.
  • Reader has to trust info from author falls within a certain level of accuracy.  In amateur writing, it is more acceptable to be over cautious than lenient.  Differing standards of acceptable content.
    • Less trust is lost if a story is underrated than overrated (resulting in disappointment)
    • A minority mislabeling work has a big effect on reputation of an archive/community (HelpingHands community members. A place to pitch in and help - a website creation resource and project. LiveJournal Community, 2005.)
  • Writer has to trust reader to make the right decision.
  • FicNet has a more specialised trust system than Golbeck's.
    • Largest contention in this field is adult material and younger readers (debated because this contrasts with IRL - no restricted areas in book stores, or suitability rating scheme for books).
    • Initially focussed on age.
    • Personas could vouch for each other.  Creating fake personae to validate another wasn't worth payoff?  Non malicious statements of distrust?
    • How to integrate trust and distrust webs?
Future
  • Ontologies developed, ready to be used by applications!
    • Ontologies will be continually refined.
    • Now designing applications.
      • Using info already gathered via quesitonnaire, re: UI, functionality.
  • Integrate with OntoMedia to describe works, and link works with people.

=> How did they get to talk to the parents of younger users?  Did they ask the members to put them in touch?  That doesn't seem like a realistic expectation to have, to me..

Week in review: Pancakes & Project management

19th November - 25th November

I read two papers about ontology development methodologies.

I read two articles by Bennett Haselton about decentralized social networking, which happened to pretty much sum up and beautifully articulate everything about that that has been floating in my subconscious for a couple of weeks.  I saw links to them in the latest Circumventor email, which I've been subscribed to since High School for bypassing the internal blacklist, and remain subscribed to because the jokes at the end are always laugh-out-loud funny.


I attended an all day course entitled 'practical project management for research students'.

  • It was attended by a diverse bunch of seemingly really lovely people.
  • The two ladies running it, from the IT Project Management department in the University, were lovely too.
  • The stuff covered was all obvious, common sense stuff (and pleasantly the organisers didn't try to claim otherwise) - but sometimes it's helpful to have it all written down and waved in your face.  And structured, in particular.  Made me actually focus on thinking about organising my project.  The main thing I hadn't much considered, even subconsciously, was formally identifying stakeholders for a project and their relative interest/power in the project.
  • There are a bunch of tools at projects.ed.ac.uk to aid in project management.
  • It prompted me to do these week-in-review posts, as I realised I haven't been recording properly everything I've been doing (an overview of my time goes on my calendar, but no detail).
  • The sandwiches weren't great, but fortunately when I got back to the Forum there were massive slabs of chocolate cake left over from some event.  I love the Forum.
I booked a place at the 1st International Open Data Dialogue in Berlin, and necessary flights.
  • Despite the short notice, it worked out logistically because I need to be in London on the 7th anyway, so I can simply go to London on the 4th, fly to Berlin from there for the 5th and 6th, and back to London for the 7th.
  • I'm particularly looking forward to "Open Statecraft: Openness as a Means (not an End)" by Philipp Müller, "The Open Data Movement vs. Business Models - is this a Contradiction?" by Dr. Peter A. Hecker, "Linked Open Data @ W3C-Vocabularies, Working Groups, Usage Scenarios" by Prof. Felix Sasaki, "The potential of Open Data for improving urban sustainability" by Dr. Marianne Linde and "Towards Trustworthiness: Establishing Transparency with Open Information Flows" by Dr. Edzard Höfig.
  • I'm also looking forward being in Berlin again, even if it is just for one evening, and I'll probably be too exhausted to appreciate it.
Ontologies with a View took place at a different place and time to usual.
I started preparing for presenting at Digital Methods as a Mainstream Methodology in London in a couple of weeks.
  • I scribbled lots of notes.
  • I skimmed a few papers by organisers/speakers but didn't read any in detail yet.  Mostly stuff about analysing data gathered from comments, tweets etc. 
  • There will be more about both of those things next week, I imagine.
I made a plan for the two weeks following the 20th.
  • It mostly consists of finishing my Digital Methods preparation.  I have a lot of non-PhD related things to do as well, plus lots of travelling.  Also graduation from my MSc, and subsequent parental visitation will get in the way.

Wednesday, November 21, 2012

Notes about ontology creation methodologies (2 papers)

Yesterday I unexpectedly read two whole papers about ontology development methodologies.  They were open in tabs I don't remember opening, but presumably did so during our weekly Ontologies With A View meeting last Friday.  There are still a bunch more tabs open with papers or articles about the same thing, so maybe I'll read those later..

The notes are here more or less as I scribbled them down whilst reading, and I haven't expanded with any analysis or discussion as of yet.

Notes in purple are things I intend/need to investigate further; colour-coding is just for me, really.


Jean Vincent Fonou-Dombeu & Magda Husiman (2011)  Combining Ontology Development Methodologies and Semantic Web Platforms for E-government Domain Ontology Development.  International Journal of Web & Semantic Technology (IJWesT) Vol.2, No.2, April 2011 

  • Start by describing ontology in a human-readable way, then turn to RDF (etc) to be machine readable.
  • Says there's not sufficient practical research around existing technologies or ontology development guidelines that would allow non-experts in e-government domain to make ontologies.
  • Uses framework from Uschold & King (see later in this post) to describe ontology - technique used here should be platform independent.
  • Then uses UML to semi-formally represent ontology
  • Uses Protege and Jena to convert to OWL and RDF
  • Paper's goal is to produce guidelines for e-government developers to create semantic content AND strengthen adoption of Semantic Web technologies in governments (particularly developing countries).
  • Outlines RDF, OWL, Protege and Jena (described as leading platforms; mentions other platforms: WebODE, OntoEdit, KAON1, Sesame).
  • Very critical of other literature; either ontologies have been produced but no practical information given; they've been developed with proprietary platforms; or they're only conceptual and don't say how they could actually be constructed with existing technologies.  other studies have not focused on a methodological approach, which means nothing is easily repeatable.
  • Detailed comparative studies of methodologies in:
    • M. Fernandez-Lopez, “Overview of Methodologies for Building Ontologies, ” In Proceedings of the IJCAI-99 workshop on Ontologies and Problem-Solving Methods (KRR5), Stockholm, Sweden, 2 August, 1999. 
    • H. Beck and H.S Pinto, “Overview of Approach, Methodologies, Standards, and Tools for Ontologies,” Agricultural Ontology Service (UNFAO), 2003. 
    • C. Calero, F. Ruiz and M. Piattini, “Ontologies for Software Engineering and Software Technology, ” Calero.Ruiz.Piattini (Eds.), Springer-Verlag Berlin Heidelberg, 2006.
  • Case study: Ontology for monitoring development projects in developing countries (OntoDPM)
    1. Create with Protege:
      • class heirarchies
      • slots
      • domain and range of slots
      • Based on the UML
      • Saved as OWL
    2. Then put content in RDF with Jena

Mike Uschold and Martin King (1995*) Towards a Methodology for Building Ontologies. Workshop on Basic Ontological Issues in Knowledge Sharing, IJCAI-95.
  • Steps:
    1. identify purpose
    2. build ontology
      • capture
      • coding
      • integrating existing ontologies
    3. evaluation
    4. documentation
  • Purpose
    • Many ontologies are intended for reuse
    • Should survey purposes to clarify options for future projects
  • Building
    • Capture
      • Identify key concepts and relationships in domain of interest
      • Produce unambiguous text definitions for these
      • Identify terms for these
      • Agree on all of the above
    • Coding
      • Explicit representation of conceptualisation in a formal language (choose a language)
    • When can capture and coding stages be merged?
    • Differences between building ontology and creating a general knowledge base (thinking about methodology will help with this)
    • Integrating (during either or both of above)
      • Work must be done in agreement between communities
      • Make explicit all assumptions underlying an ontology
  • Evaluation
    • Judge against requirements specification (and/or)
    • Judge against competency questions (and/or)
    • Judge against real life
    • This paper looks at knowledge base systems, and adapts for ontologies.
  • Documentation
    • Desirable to have established guidelines for documenting
    • Main barrier to effective knowledge sharing is inadequate documentation
    • ALL important assumptions should be documented
  • Case Study
    • Main emphasis is on capture phase
    • Initially:
      • define ontology (Gruber)
      • identify users and usage (initially abstract, then clarify with real life)
      • choose language (Ontolingua was chosen)
      • choose method for capture - BDSM (IBM) supported by others:
        • KADS
        • IDEF5
        • OO Analysis and Design techinques
        • Gruber's principles for ontology design
    • Categorisation is fundamental to the human condition (Lakoff)
      • Not heirarchical, but:
        • GENERAL
                 ^
             BASIC  ->  primary with respect to knowledge organisation
                 v
          SPECIFIC
        • eg.
                SUPER:   Animal    /   Furniture
                BASIC:    Dog        /   Chair
                SUB:        Retriever /   Rocker
        • Certain concepts used subconsciously, rather than understood intellectually.
          • These have a more important psychological status.
        • Therefore paper uses middle-out approach to capture terms
          • (bottom-up = too much detail unnecessarily,
          • top-down = risks imprecision)
        • BASIC concepts first because:
          • most important
          • used to define non-BASIC terms
          • increase clarity, especially for non-technical use
          • backed by BSDM experience of paper author
  • Scoping
    • Brainstorming
    • Consult corpora if there aren't enough domain experts to brainstorm
    • Grouping
      • structure terms into naturally arising sub-groups
      • collate synonyms
      • consider things that might refer to each other
  • Meta-ontology
    • Don't commit too early, can restrict thinking.  Let concepts and relationships themselves determine requirements.
    • Be consistent.
    • Use technologically neutral language ('thing' vs 'entity').
    • Start with areas where there's most overlap.
    • Work from basic terms to more abstract ones within an area.
  • Producing definitions
    • Agreeing on definitions (varying degrees of problems)
    • Handling ambiguous terms
      • clarify ideas without technical terms
      • use a dictionary!
      • label definitions, eg. x1, x2
      • determine most important concept
      • choose a term, avoiding original ambigious one
    • Avoid new terms
    • Terms get in the way (peoples' preconceived ideas) - concentrate on underlying meaning and concepts.

Friday, November 09, 2012

National Novel Procrastinating Month

It's day nine, and I'm on two thousand, eight hundred and seventy words.

A quick calculation might tell you that that means I'm quite behind schedule.  This may be my worst year yet.  There's still plenty of time to get back on track though!  Right..?!

I've only spent any time writing on about three or four of those nine days so far.  But I have been to a conference, organised some SocieTea events, read bits and pieces related to my PhD, cleaned my flat, watched a few episodes of Arrested Development, learnt some new crochet stitches and started crocheting a hat, and baked a lot.

I did meet the Edinburgh NanoBeans and had a great time at the write-in in Pulp Fiction last Wednesday.  We may have spent more time collaboratively developing the backstory of Pedro the Guide Bear (a troubled young grizzly attired in an Elvis costume and boater hat who constantly struggles against his estranged father, Yogi, the leader of an organised crime syndicate) than actually writing our novels though.

I have learnt one particularly important thing this year, that's never come up before.

Talking ideas through with other people is really useful!  

Last Sunday, Beth helped me explain the absence of a main character's mother and fix a potential looming plot hole with one fell swoop.  Telling Kit about the various civilisations and layout of the land in my world allowed him to pick holes and question things, raising, and partially solving, some things that didn't make sense or yet more potential looming plotholes.  And Caitlin (a new NanoBeans writing buddy) pointed out that just because a character had been anticipating reading a letter for the last thousand words, didn't necessarily mean the letter had to contain anything interesting... it could be a disappointment to the character... which helped, as I hadn't figured out what the letter said, and all of a sudden the character was opening it.

I sure wish blogging about Nano counted towards the word count.

Monday, November 05, 2012

Remediating the Social #elmcip

I spent the last few days in Edinburgh College of Art, helping out at the Remediating the Social conference.  I was in charge of making sure everyone's microphones were on, and slides were being projected, which turned out to be more work than anyone anticipated.  Only minor hiccups occurred though, usually when I unplugged something I shouldn't have by accident.  I couldn't have done it without my glamourous assistant José, who was the master of fiddling with Macbook screen resolutions to make them play nice with the projector.

More importantly, I saw some super interesting talks, and met and talked to some fantastic smart people about electronic literature, and other things.

I also presented about Palimpsest, in front of the biggest audience I have ever talked in front of.  Go me.

Videos of everything from the conference are here.

On the last day I implemented an idea that had been kicking around the back of my mind for a while, which was the Uninformative Twitter Wall, or Twitter Squares.  It's nothing particularly complex; it uses jQuery and probably has memory leaks.  I'd love for people to help themselves to the code and improve it. Converting a hash of a tweet text into a hex code, I generated coloured squares for the results of a search term.  If the feed you choose is updating a lot, then the squares move around quickly and it looks pretty funky.  If there are only occasional new tweets, then it looks less exciting, but is still equally useless for seeing what people are saying.  (Unless you hover over the squares).  That's okay though, because it's Art.

Wednesday, October 31, 2012

National Novel Writing Month

That's write right, it's the eve of Nanowrimo.

Last year, my MSc got violently in the way and I clocked out at about 15,000 words.  I'm hoping that this year, my PhD will make friends with my month of literary abandon, and both will come out better for it.

I stumbled into a new world last May, wandered around and met some characters over the summer, and have been mulling over them ever since.  I put pen to paper to draw a map today, and discovered that more of the world was there than I thought.

I have three viewpoint characters, and next I'm going to draw some squiggly lines on a piece of paper to figure out where their paths cross, and what might happen to them along the way.  Over the years I'm becoming more inclined towards plotting in advance, but a large part of me never really thinks it'll help.

I've been reading A Song of Ice and Fire, and am now gagging to create a world with half as much depth and drama as GRRM has done.  Mine will be fantasy, with a hint of sci-fi and a dash of Ancient Egypt (probably no medieval knights).

This year I'm going to work in yWriter in an attempt to keep on top of things as I expand settings and characters.

I'm hoping to attend more than just the launch party for the Edinburgh Nanobeans group this time round.  Though they do meet across the wrong side of town, so I might also start my own write-ins (consisting of just me) in Himalaya Cafe on South Clerk St. (it's ever so comfy, and the chai is the best).  If you're writing too, and in that neck of the woods, come and join me.  I'm tentatively saying I'll be there between 10 and 11am every day (except Sunday, they're closed), starting on the 5th.  (This is going to cost me a fortune in chai, isn't it..)

Stay tuned for progress updates.  Or lackthereof.

PS. I know I promised notes on papers related to my PhD... One day.  One day.


Monday, September 17, 2012

University 3.0

Even though I don't officially start my PhD until the 1st of October, today really felt like a proper first day of term.  I got up early went to a couple of fourth year/MSc classes that I've decided to sit in on (HCI and Text Technologies), went to training for tutoring/demonstrating, filled in some forms, got my new student card (that doesn't expire until 2016!) and most importantly, got the key to my office in the Informatics Forum.

My PhD ideas are vague at best right now, though it'll definitely be within the realms of the Semantic Web.  According to the proposal I wrote to apply for the position back in May, it'll be to do with provenance of and collaborative creation of digital media artefacts, like comics and films.  It'll be interesting to watch that morph and change.

Though I became aware of Semantic Web stuff during my undergraduate, I developed my knowledge during my MSc at Edinburgh.  Primarily by taking the Multi-Agent Semantic Web Systems course for credit, enjoying it a lot and doing pretty well.  (I should be TAing/marking for that this year).  I also learnt lots about linked data and other such things at conferences and hacks like Dev8D, and various open data meet-ups.  I'm super excited about the future of the Internet - particularly making sure it remains an open, public platform for uncensored expression and knowledge sharing (fingers crossed).  Since I'm a technologist, not a lawyer or policy-maker, I have to address this with theoretical and practical research around how people create and share things, and ways to improve connectivity (between people and data), which of course includes the big problems like privacy, security and identity.

Whilst I'm confident enough to say I know quite a lot about designing and developing for the Web these days, my Semantic Web knowledge really consists of a basic grounding, and a lot of enthusiasm.

There's a ton of research going on in various related areas, so I've decided to read one or two relevant academic papers a day... forever, I guess... and make notes on what I read.  Publishing my notes here works as a subconscious stimulant, to make sure I actually get it done.  A lot of them might be foundations, or basic stuff, but I intend to cram as much as possible - especially in the couple of weeks before I start propertly.  So look out for those! (If you're interested.  If not, ignore them).

Wednesday, June 13, 2012

Shifting focus (from readers to authors)


By sheer coincidence, I find myself neck-deep in creating a custom Interactive Fiction engine during a year in which all kinds of new engines for authoring and enjoying interactive narrative are popping up. Many of these by established programming and literary experts in the IF community.

It feels like every week Emily Short posts a review of something new. Most of the ones I've seen recently are more of a Choose Your Own Adventure format than parser-based. Nonetheless, there is a lot I can learn about world modelling and code-free authoring from these systems. Not all of them are open to the public for creating though. All of them seem to be open to Emily Short, however, and she is doing a great job of describing and reviewing her experiences.

It occurred to me a while ago that creating the interface for authoring pieces for Palimpsest is potentially more interesting and important than examining the reader experience when it's all done. Readers are much more fickle, and those who spend time writing and creating are by definition the committed and passionate. I'm not saying the readers aren't. But a solid interface for authoring is going to make it much more likely that good experiences for readers are to be created. If I rush through the design and build of the software itself so that I can focus on filling it with content then quizzing readers about the immersiveness of their experiences - as was my original plan - there's a good chance I won't spend nearly enough time testing with authors (as I'll be doing the authoring) and will miss the opportunity to allow people to create immersive or particularly engaging experiences, jeopardising the outcomes of the user study stuff in the end anyway.

I only have about two months left to work on this in the context of my MSc. Whilst I need something substantial to fill my final report with, I'm certain I'll be continuing to develop this in the future, and I'll only end up having to redo bits from scratch if I cut corners the first time.

I also have to think about what I want to get out of my MSc, and how I want to use the opportunities I have studying in an art department compared to what would be expected of me studying in a computing department.

I jumped on the idea of 'user study stuff' because I know how to do this in terms of software, and I know how to write a lot about it.  Not that I don't find the immersion thing interesting, but overall I think it would be more valuable to both myself and any eventual users of Palimpsest if I produced documentation of detailed musings about the perhaps unfinished or forever-ongoing development of a custom engine and authoring interface for Interactive Fiction, rather than gloss over that to fill my report with lots of nice statistics and quotes and analyses about reader experiences that are ultimately invalid anyway.

And I can do this for an 'art' project, where the journey is equally, if not more, important than the outcomes.  Where for a 'computing' project, the documentation of the journey has in my experience generally been intended to be a build up to an overarching theoretical or practical conclusion.

So next up, look out for a discussion of the various engines for different flavours of interactive narrative that have been popping up / I've been noticing recently, coming soon to a blog near you!  (This one.  This is the blog it'll be coming to.  Not any old blog near you).