Amy Guy

Raw Blog

Showing posts with label workshop. Show all posts
Showing posts with label workshop. Show all posts

Monday, April 22, 2013

[Notes] How to write a literature review workshop


Just notes!

Workshop by Dr Mimo Caenepeel on Monday 22nd April.

'Critical' does not mean you have to pass judgement, or say why it's good or bad.
Not taking things at face value.

Started with freewriting about what has particularly influenced / inspired our own research.  Five minutes, not allowed to stop or edit, don't worry about quality of writing, not for anyone else to read.  A good way to get ideas out of your head and start to organise your thoughts without censoring or constraining yourself.

How many pages will a review usually take up in a thesis?  My policy is to write what needs to be written and stop when you're done.  But apparently 20 to 30, sometimes more, is normal in sciences.

There's no consistent / right answer to 'how many publications to review'.  For some people it's in the tens, for some the hundreds.

Think about how to integrate literature review into the thesis.  You're unlikely to have a chapter that is just 'literature review' and no mention of the background reading elsewhere.

Good qualities for a lit review?
- Coherence (avoid fragmentation)
- Structure, clarity.
- Proof of novelty - purposeful.

A review can often be considered as an indicator of the quality of the rest of the research - demonstrating scholarship.

A good place to start:
1. Write your research question, formulated as a question.
2. Write up to five research areas that are relevant to your research question.
3. Note some related issues/areas that will not be considered in your review.

Think about balance of content.
1. Three studies influential in your field (I couldn't answer this, I clearly need to read more).
2. Two significan older contributions.
3. Five recent sources.
4. Two sources that have strongly influenced your thinking.

You don't need to consider all papers in the same level of detail.  Decide which papers are more important / useful than others.

For some papers (important ones) you should work through these questions in the same way every time you read something (this is 'SQ3R'):
1. Survey: What is the gist of the article? Skim the title, abstract, introduction, conclusion and section headings. What stands out?
2. Question: Which aspects of the research are particularly relevant for your review? Articulate some relevant questions the article might address.
3. Read: Read through the text more slowly and in more detail and highlight key points / key words.  Identify connections with other material you have read.
4. Recall: Divide the text into manageable chunks and summarise each chunk in a sentence.
5. Review: To what extent has the text answered the questions you formulated earlier?

Critical reading (these seem like really useful questions to work through whilst reading papers):
1. What is the author's central argument or main point, ie. what does the author want you, the reader, to accept?
2. What conclusions does the author reach?
3. What evidence does the author put foward in support of his or her conclusions?
4. Do you think the evidence is strong enough to support the arguments and conclusions, ie. is the evidence relevant and far-reaching enough?
5. Does the author make any unstated assumptions about shared beliefs with readers?
6. Can these assumptions be challenged?
7. Could the text's scientific, cultural or historical context have an effect on the author's assumptions, the content and the way it has been presented?

See Ridley, D. The Literature Review: A step-by-step guide for students.  Sage Study Skills Series. Sage Publications, 2011 (2008).

Thursday, April 18, 2013

[Notes] 'How to write a thesis' workshop

Just notes from a three-hour workshop about how to write an Informatics thesis, on the 16th of April.


State contributions (to knowledge) explicitly.  Intro, conclusions; each chapter should have some (probably not all) contributions discussed.  Be obvious; use headings.

Knowledge - background:

  • justify choices
  • explain methods
  • acknowledge alternatives
  • evaluate

Evidence, well-reasoned arguments, acknowledge limitations.

Clear openings for future work.  Be clear where they are.

Make it reproduceable.

Short / concise.  Examiners like short theses.

Introduce what's interesting and important.

When outline thesis, look at structure of main argument, not of document.

Background material must have point.  Only include as much detail as you need to make point.
Points, eg:

  • Explain method you use.
  • Novelty of your approach. Similarities with existing work.
  • Justify choices (evaluate other work).
  • Don't tear down others' work. 'Build on'.
  • Cite examiners, they've probably published something relevant.. (but not for the sake of it).


Then we had five minutes to write down what our PhDs are about and what we have already found out.  I wrote:

How do the futures of the Semantic Web and amateur digital content creation fit together?
Can Semantic Web tools and technologies be used to enhance collaborative creative partnerships and encourage fruitful outputs?

There are knowledge sharing systems and collaborative tools for scientific fields and in education, but nothing for creative artsy things.

Attitudes towards data sharing and privacy amongst content creators are in flux.  There are lots of projects and energy around open data and decentralised social networks that allow data to become portable and not tied to one platform.  One of TBL's visions for the Semantic Web is the dissolution of data silos and 'walled' applications that disadvantage the user, and as such the promotion of the 'ownership' of a user's data by the user themselves, rather than the software or organisation that uses the data.

There are lots of reasons people make content.  There are lots of reasons people don't make content (who could / would like to).

[Notes resume]
Use backreferences; don't repeat yourself.

Info / advice
...homepages.../sgwater/resources.html
..homepages.../imurray2/teaching/writing
Style: Toward Clarity & Grace (book)
The Craft of Research (book)

When to start writing thesis?

  • Do you already have papers?  Slot them into a thesis template asap.
  • Maybe a year beforehand.  Slower pace is better.

Don't assume appendices will be read.  More for extra info if needed by people trying to reproduce your work (not your examiners).

Too many direct quotes look like you don't understand and are avoiding explaining yourself.

Keep copies of web resources and cite access dates in case they change / disappear.
Figures might be copyright if you just copy them from papers, even if you cite them.  Remake them, and put 'adapted from' as citation.

Examiners?

  • Depends on your supervisor.  Discuss.  Student might be able to suggest someone to examine.
  • Maybe a balance between internal and external knowledge.
  • Won't be someone junior, even if they're considered an expert in the field.
  • Helpful if supervisor knows how that person will behave in viva.  Might be a good reason to avoid someone you think would be perfect from their background.
  • Conflict of interest regulations.  You can know them personally though.  External can't have been affiliated with UoE in the last three years, or substantially involved in your research (like co-authoring a paper).  No ex-supervisors, from any university.

No grading system (ie no different levels of passed PhD).  Might be external prizes if you want extra recognition.

Thursday, April 11, 2013

2nd UK Ontology Networks Workshop

The UK Ontology Networks Workshop took place over one day in the Informatics Forum.

There was a mix of people there; some talks were way over my head and very technical, and some talks were by people who confessed they had had to look up "ontology" that morning.  And things in between.

Lazy writeup, but following are notes as I scribbled them:



John Callahan

US navy research.
Focused information integration.
Human intervention to keep predictive part on track. Tweaking.

Alan Bundy

Interaction of representation and reasoning.
Changing world so agents must evolve. How to automate? What would trigger a need for change:
Inconsistency
Incompleteness
Inefficiency
how to diagnose which?
Interested in language and perception change.
Unsorted first order logic algorithm called Reformation. Based on standard unification algorithm.
Allows blocking and unblocking unification.


Phil Barker

Schema.org
Cetis (JISC funded)
learning resource metadata initiative.
Big names behind schema.org.
= ontology + syntax
Big and growing ontology.
Dumbed down for people.
LRMI adds to it. W3C go through it. It's creeping, how much do the big names actually care about stuff that's added?
don't know how Google uses it.
People should consider using it for more sophisticated search and disambiguation.

Gill Hamilton

Doing more with library metadata. Learnt from OKFN. Had to convince people in charge.
Dublin core, didn't like; not specific enough. Instead RDF > OWL. "We know best how to structure our data"

Hardest was convincing marketing people that there was no commercial value. Metadata is advert to actual resource.

Enrico Motta

Traditionally top down approach. So now so many people interacting with semantic structures, so should involve users.
Recognise there isn't a unique or best way of doing things.
Initial study included modeling task with binary relations.

Patterns that are more or less intuitive. 4D least, 3D+1 most.
N-ary most widely used by experts.

Relationship between reasoning power and intuitiveness of writing? More creativity needed for simpler ones. (Not really sure what he's saying)

Email him for copy of study.

Chris Mellish

Ontology authoring is hard. Better ways to do it.

Controlled language input (mature tech); responsive reasoning (also mature, information as you're editing); understanding the process (beginning to understand more).

Hypotheses:
users don't know what they're doing. What if questions.  Many answers, what is relevant? Depends on context.

Authoring as dialogue.
Todo list.

Useable in the same ways as protégé.

Peter Winstanley

UN classification schemes.
Various vocabularies.
Allow development of cross mapping between government administrations.

Mostly internal currently. Moves to bring externalizing data into the 21st century.

Peter Murray-Rust

Fight for your Ontologies.
Ontologies in physical sciences. Chemists don't want ontologies. They'll sue you.
Crystallography uses 'dictionary'. Written in CIF. 20 years to build CIF.

Compare physical sciences to government.

Every program author writes dictionaries that work for them. When different parties agree, promote to communal dictionary. Provide conventions to help disagreements.

Show a company can do it as opposed to a rabbiting academic ..

Jeff Pan

Tractable ontological stream reasoning.
Need to be more efficient, scaleable, as things change. Inputs from web.

Dealing with complexities: approximate owl2.
Dealing with frequent updates: to-add stream and to-do delete stream. Truth maintenance. Evaluation criteria.

Trowl.EU can use with protégé, also supports jena.

Edoardo Pignotti

Semantic web tech to support Interdisciplinary research.
ourSpaces VRE
Provenance crucial.
OPM prov ontology.

Deployed since 2009, 180 users. Comprehensive ontologies but people unwilling to provide metadata.
paper! Edwards et al. ourSpaces.

Tom Grahame (BBC) @tfgrahame

Content arrangement on BBC sport by tagging, automatic to free up editors to write.
LD API so systems don't need to know about each other.
Growing from simple rdfxml to more complex ontology.
Can ask much more general and much more detailed questions about sport.

Mapping incoming data is outsourced.
Lots of errors, sometimes system alerts, sometimes manual.

Working on opening the data. Maybe a dump, but licensing issues.

Ewan Klein

Mining old texts for commodities, adding place and time and putting in structured database.
Transcriptions of customs import records.

Skos for synonyms.
Dbp concepts.

Why? Want to query.
Visualisations.

Tools? Python script.

Janice Watson

Harnessing clinical terminologies and classifications for healthcare improvements.

Bob Barr

Geographical addressing.
Addressing and address geocoding is important and broad. Not always postal, but this not addressed (punlol) in ontologies.
Different contexts change meaning of address (for delivering, you only care about postbox; property sale whole building).
Loads of things to address. Loads of reasons why.
Work held up as national address file is owned by royal mail and might be sold!

Fiona McNeill

Run time extraction of data. Failure driven. Looking at extraction of specific information.
Emergency response. Lots of data, timely sharing of data required.
From domestic level to humanitarian disasters.
How can it be automated?
Multilayered incompatibility.
Format
Terminology
Structure
...

Richard Gunn

Towards an intelligent information industry.

Elena Simperl (Soton, sociam)

Crowdsourcing ontology engineering.

CSrc: Brabham 2008.

Distribute task into smaller atomic units.

Humans validating results that are automatically detected as not accurate.
What are the costs? What resources?

Games with a purpose. Like quizzes.
Micropayments or vouchers.
MTurk. CrowdFlower.
Paper about useage of microtask crowdsourcing.  ISWC 2012.

Claudia Paglieri

Ontologies in ehealth.

Enrico Motta - Rexplore
Klink algorithm mines relations between research topics.
Use this!  Nope, it's not public.   Uees MS Academic research.

Peter Murray-Rust

Content mining expands regular text mining.
Focus on academic stuff.
Chemical Tagger. Takes chemistry jargon and annotated it, knows actions, conditions, molecules etc.. NLP. Uses ontologies and contributes to ontologies.
In chemistry,  no need to put everything in rdf because there are already lots of formalisms.
Proper cool PDF to sensible format conversion. Amy the kangaroo. Looking for collaborators.

Yuan Ren

Ontology authoring in whatif project.

Reasoning with protégé and trowl .

Tractable reasoning. Trowl v fast.


Notes from conversations / breakout discussions:

BBC use owlm triplestore  .
Store all their datasets in svn. But they have reads and writes to the live triplestore all the time.

Lots of people saying minimise owl use because of unpredictable output.

Versioning ontologies (available in owl2) in case third parties change stuff you use. You're dependent on their software engineering practices. Only good if they're ahead of the game.

IRIs, Arabic characters in ontologies!
Semantic heavy, maybe make a decision to abstract away to ids and make heavier use of labels.

Difference between importing and using someone else's.

There's no (practically useful) software that lets you reason over stuff you haven't imported? (over HTTP?)

Build ontology from reality (data), don't start with no data.

Lode.

Problems with dbpedia URIs changing or disappearing.

Hard to visualize massive graphs. Relational, tabular much easier to understand.