'Museums meet the 21st century' – OpenTech 2010 talk

These are my notes for the talk I gave at OpenTech 2010 on the subject of 'Museums meet the 21st Century'. Some of it was based on the paper I wrote for Museums and the Web 2010 about the 'Cosmic Collections' mashup competition, but it also gave me a chance to reflect on bigger questions: so we've got some APIs and we're working on structured, open data – now what? Writing the talk helped me crystallise two thoughts that had been floating around my mind. One, that while "the coolest thing to do with your data will be thought of by someone else", that doesn't mean they'll know how to build it – developers are a vital link between museum APIs, linked data, etc and the general public; two, that we really need either aggregated datasets or data using shared standards to get the network effect that will enable the benefits of machine-readable museum data. The network effect would also make it easier to bridge gaps in collections, reuniting objects held in different institutions. I've copied my text below, slides are embedded at the bottom if you'd rather just look at the pictures. I had some brilliant questions from the audience and afterwards, I hope I was able to do them justice. OpenTech itself was a brilliant day full of friendly, inspiring people – if you can possibly go next year then do!

Museums meet the 21st century.
Open Tech, London, September 11, 2010

Hi, I'm Mia, I work for the Science Museum, but I'm mostly here in a personal capacity…

Alternative titles for this talk included: '18th century institution WLTM 21st century for mutual benefit, good times'; 'the Age of Enlightenment meets the Age of Participation'. The common theme behind them is that museums are old, slow-moving institutions with their roots in a different era.

Why am I here?

The proposal I submitted for this was 'Museums collaborating with the public – new opportunities for engagement?', which was something of a straw man, because I really want the answer to be 'yes, new opportunities for engagement'. But I didn't just mean any 'public', I meant specifically a public made up of people like you. I want to help museums open up data so more people can access it in more forms, but most people can't just have a bit of a tinker and create a mashup. “The coolest thing to do with your data will be thought of by someone else” – but that doesn’t mean they’ll know how to build it. Audiences out there need people like you to make websites and mobile apps and other ways for them to access museum content – developers are a vital link in the connection between museum data and the general public.

So there's that kind of help – helping the general public get into our data; and there's another kind of help – helping museums get their data out. For the first, I think I mostly just want you to know that there's data out there, and that we'd love you to do stuff with it.

The second is a request for help working on things that matter. Linkable, open data seems like a no-brainer, but museums need some help getting there.

Museums struggle with the why, with the how, and increasingly with the "we are reducing our opening hours, you have to be kidding me".

Chicken and the egg

Which comes first – museums get together and release interesting data in a usable form under a useful licence and developers use it to make cool things, or developers knock on the doors of museums saying 'we want to make cool things with your data' and museums get it sorted?

At the moment it's a bit of both, but the efforts of people in museums aren't always aligned with the requests from developers, and developers' requests don't always get sent to someone who'll know what to do with it.

So I'm here to talk about some stuff that's going on already and ask for a reality check – is this an idea worth pursuing? And if it is, then what next?
If there’s no demand for it, it won’t happen. Nick Poole, Chief Executive, Collections Trust, said on the Museums Computer Group email discussion list: "most museum people I speak to tend not to prioritise aggregation and open interoperability because there is not yet a clear use case for it, nor are there enough aggregators with enough critical mass to justify it.”

But first, an example…

An experiment – Cosmic Collections, the first museum mashup competition

The Cosmic Collections project was based on a simple idea – what if a museum gave people the ability to make their own collection website for the general public? Way back in December 2008 I discovered that the Science Museum was planning an exhibition on astronomy and culture, to be called ‘Cosmos & Culture’. They had limited time and resources to produce a site to support the exhibition and risked creating ‘just another exhibition microsite’. I went to the curator, Alison Boyle, with a proposal – what if we provided access to the machine-readable exhibition content that was already being gathered internally, and threw it open to the public to make websites with it? And what if we motivated them to enter by offering competition prizes? Competition participants could win a prize and kudos, and museum audiences might get a much more interesting, innovative site. Astronomy is one of the few areas where the amateur can still make valued scientific contributions, so the idea was a good match for museum mission, exhibition content, technical context, and hopefully developers – but was that enough?

The project gave me a chance to investigate some specific questions. At the time, there were lots of calls from some quarters for museums to produce APIs for each project, but there was also doubt about whether anyone would actually use a museum API, whether we could justify an investment in APIs and machine-readable data. And can you really crowdsource the creation of collections interfaces? The Cosmic Collections competition was a way of finding out.

Lessons? An API isn't a magic bullet, you still need to support the dev community, and encourage non-technical people to find ways to play with it. But the project was definitely worth doing, even if just for the fact that it was done and the world didn't end. Plus, the results were good, and it reinforced the value of working with geeks. [It also got positive coverage in the technical press. Who wouldn’t be happy to hear ‘the museum itself has become an example of technological innovation’ or that it was ‘bringing museums out into the open as places of innovation’?]

Back to the chicken and the egg – linking museums

So, back to the chicken and the egg… Progress is being made, but it gets bogged down in discussions about how exactly to get data online. Museums have enough trouble getting the suppliers they work with to produce code that meets accessibility standards, let alone beautifully structured, re-usable open data.

One of the reasons open, structured data is so attractive to museum technologists is that we know we can never build interfaces to meet the needs of every type of audience. Machine-readable data should allow people with particular needs to create something that supports their own requirements or combines their data with ours to make lovely new things.

Explore with us – tell museums what you need

So if you're someone who wants to build something, I want to hear from you about what standards you're already working with, which formats work best for you…

To an extent that's just moving the problem further down the line, because I've discovered that when you ask people what data standards they want to use, and they tell you it turns out they're all different… but at least progress is being made.

Dragons we have faced

I think museums are getting to the point where they can live with the 80% in the interest of actually getting stuff done.

Museums need to get over the idea that linkable data must be perfect – perfectly clean data, perfectly mapped to perfect vocabularies and perfectly delivered through perfect standards. Museums are used to mapping data from their collections management systems for a known end-use, they've struggled with open-ended requirements for unknown future uses.

The idea that aggregated data must be able to do everything that data provided at source can do has held us back. Aggregated data doesn't need to be able to do everything – sometimes discoverability is enough, as long as you can get back to the source if you need the rest of the data. Sometimes it's enough to be able to link to someone else's record that you've discovered.

Museum data and the network effect

One reason I'm here (despite the fact that public speaking is terrifying) is a vision of the network effect that could apply when we have open museum data.

We could re-unite objects across time and place and people, connecting visitors and objects, regardless of owing institution or what type of object or information it is. We could create highlight collections by mining data across museums, using the links people are making between our collections. We can help people tell their local stories as well as the stories about big subject and world histories. Shared data standards should reduce learning curve for people using our data which would hopefully increase re-use.

Mismatches between museums and tech – reasons to be patient

So that's all very exciting, but since I've also learnt that talking about something creates expectations, here are some reasons to be patient with museums, and tolerant when we fail to get it right the first time…

IT is not a priority for most museums, keeping our objects secure and in one piece is, as is getting some of them on display in ways that make sense to our audiences.

Museums are slow. We'll be talking about stuff for a long time before it happens, because we have limited resources and risk-averse institutions. Museum project management is designed for large infrastructure projects, moving hundreds of delicate objects around while major architectural builds go on. It's difficult to find space for agility and experimentation within that.

Nancy Proctor from the Smithsonian said this week: "[Museum] work is more constrained than a general developer" – it must be of the highest quality; for everybody – public good requires relevance and service for all, and because museums are in the 'forever business' it must be sustainable.

How you can make a difference

Museums are slowly adapting to the participation models of social media. You can help museums create (backend) architectures of participation. Here are some places where you can join in conversations with museum technologists:

Museums Computer Group – events, mailing list http://museumscomputergroup.org.uk/ #ukmcg @ukmcg

Linking Museums – meetups, practical examples, experimenting with machine-readable data http://museum-api.pbworks.com/

Space Time Camp – Nov 4/5, #spacetimecamp

‘Museums and the Web’ conference papers online provide a good overview of current work in the sector http://www.archimuse.com/conferences/mw.html

So that‘s all fun, but to conclude – this is all about getting museums to the point where the technology just works, data flows like water and our energy is focussed on the compelling stories museums can tell with the public. If you want to work on things that matter – museums matter, and they belong to all of us – we should all be able to tell stories with and through museums.

Thank you for listening

Keep in touch at @mia_out or https://openobjects.org.uk/

What would you change about your workplace? A survey for museum technologists

[Update: I've shared the data]

This week I launched a survey designed to help me understand and communicate the challenges faced by other museum technologists.

It's research for a chapter in a forthcoming book on museums on the web and social media in the first instance, but I'd left the terms and conditions fairly open as I wanted to be able to share and/or re-use the data in future – I wasn't sure if this would put people off, but I figured it was better to be upfront than to end up with great data I couldn't share.

Someone wrote to me to ask what the questions were – they didn't feel qualified to take it themselves but couldn't see all the questions without starting the survey.  I figure it'll also help with the bounce rate if I share them, so here you go:

1. As a museum technologist, what are the three most frustrating things about your job?

For this survey, I'm defining 'museum technologist' as someone who has expertise and/or significant experience in the museum sector and with the application or development of new technologies.

2. List any solutions for each of the problems you listed above

3. Any comments on this survey or on the issues raised?

4. What's your main job role? (if you don't mind it potentially being quoted)

5. Please enter your institution name and/or type (e.g. art gallery, history museum, local authority museum, science centre). (if you don't mind it potentially being quoted)

It's pretty simple – 'what are the three most frustrating things about your job' is the main question, the rest are aimed at providing just enough additional information to provide pointers to the effects of different factors. I didn't want to ask people for so much information that they'd be identifiable as I felt that might make people hold back.  I thought about saying 'things other than lack of resources/time/money' as they're pretty much a given and they're not unique to the museum sector, but I figured they're also too important too ignore.
For further context, when I posted it to the MCG and MCN lists I said:

I'm particularly interested in opportunities and problems that arise when (new) technologies meet (old) museums. … Your answers will help build a body of evidence that could help make a case for improvements in the way museums understand the issues and expertise around using technology to engage audiences, or at least help us understand what the solutions might be. And at the very least you get to vent a bit!

I'm running the survey until August 31 and my initial analysis will be completed by mid-September.  If you'd like to take the survey, or know someone who should, the address is http://www.surveygizmo.com/s3/348155/Challenges-facing-museum-technologists (or http://bit.ly/95oGtr if shorter is easier).

Finally, thanks to the person who suggested making the text boxes wider – I've done that now.

Ask a cultural heritage technologist?

I'm speaking at Open Tech 2010 (book your ticket now, only £5!) and it feels like the situation (and the mood) in the UK has changed since I first wrote my proposal and I'm not sure it suits anymore.  So I wanted to throw a few questions open to you to help me re-focus on the things that matter now:

  • what do you value about museums and technology, particularly the web, social media, open data? 
  • what do you want to know from someone working behind the scenes in museum technology?
  • what suggestions would you make if you were able to talk to museums?
  • what aren't museums asking our audiences (including our geek audiences) that we should be asking?
  • what's your favourite biscuit (or cookie)?
The title, by the way, is a play on 'ask a curator', an online event of some sort where you can ask whatever you've always wanted to ask a curator by using the hash tag #askacurator on twitter (or possibly also by commenting on a museum's blog, Facebook wall, twitter account, etc).

Linking museums: machine-readable data in cultural heritage – meetup in London July 7

Somehow I've ended up organising an (very informal) event about 'Linking museums: machine-readable data in cultural heritage' on Wednesday, July 7, at a pub near Liverpool St Station. I have no real idea what to expect, but I'd love some feisty sceptics to show up and challenge people to make all these geeky acronyms work in the real museum world.

As I posted to the MCG list: "A very informal meetup to discuss 'Linking museums: machine-readable data in cultural heritage' is happening next Wednesday. I'm hoping for a good mix of people with different levels of experience and different perspectives on the issue of publishing data that can be re-used outside the institution that created it. … please do pass this on to others who may be interested. If you would like to come but can't get down to that London, please feel free to send me your questions and comments (or beer money)."

The basic details are: July 7, 2010, Shooting Star pub, London. 7:30 – 10pm-ish. More information is available at http://museum-api.pbworks.com/July-2010-meetup and you can let me know you're coming or register your interest.

In more detail…

Why?
I'm trying to cut through the chicken and egg problem – as a museum technologist, I can work towards getting machine-readable data available, but I'm not sure which formats and what data would be most useful for developers who might use it. Without a critical mass of take-up for any one type, the benefits of any one data source are more limited for developers. But museums seem to want a sense of where the critical mass is going to be so they can build for that. How do we cut through this and come up with a sensible roadmap?

Who?
You! If you're interested in using museum data in mashups but find it difficult to get started or find the data available isn't easily usable; if you have data you want to publish; if you work in a museum and have a
data publication problem you'd like help in solving; if you are a cheerleader for your favourite acronym…

Put another way, this event is for you if you're interested in publishing and sharing data about their museums and collections through technologies such as linked data and microformats.

It'll be pretty informal! I'm not sure how much we can get done but it'd be nice to put faces to names, and maybe start some discussions around the various problems that could be solved and tools that could be
created with machine-readable data in cultural heritage.

'Game mechanics for social good: a case study on interaction models for crowdsourcing museum collections enhancement'

I've been very quiet lately – exams for my MSc and work on the digital infrastructure for two new galleries (and a contemporary science news website) opening next week at the Science Museum have kept me busy – but I wanted to take a moment to post about my dissertation project. (Which reminds me, I should write up the architecture I designed to extend our core Sitecore CMS with WordPress to support social media-style interactions with Science Museum-authored content.)

Anyway. This project is for my dissertation for City University's Human-Centred Systems MSc. I'm happy to share the whole outline, but it's a bit academic in format for a blog post so I've just posted an excerpt here. I'd love to hear your comments, particularly if you know of or have been involved in creating, crowdsourced museum projects or games for social good.

'Game mechanics for social good: a case study on interaction models for crowdsourcing museum collections enhancement' is the current title – it's a bit of a mouthful but hopefully the project will do what it says on the tin.

Project description
The primary focus of this project is the design and evaluation of interactions applied to the context of an online museum collection in order to encourage members of the public to undertake specific tasks that will help improve the website.

The project will include a design and build component to create game-like interfaces for testing and evaluation, but the main research output is the analysis of museum crowdsourced projects and 'games for social good' to develop potential models for game-like interactions suitable for museum collections, and the subsequent evaluation of the proposed interaction models.

Aims and Objectives
This project aims to answer this question: can game-like interactions be designed to motivate people to undertake tasks on museum websites that will improve the overall quality of the website for other visitors?

More specifically, which elements of game mechanics are effective when applied to interfaces to crowdsource museum collections enhancement?

Objectives

  • Design game-like interaction models applicable to cultural heritage content and audiences through research, analysis and creativity workshops
  • Build an application and interfaces to create and store user-created content linked to collections content
  • Evaluate the effectiveness of game-like interaction models for eliciting useful content

Theory
Recent projects such as Armchair Revolutionary[1] and earlier projects such as Carnegie Mellon University's 'Games with a purpose'[2] and InterroBang?![3] are indicative of the trend for 'games for social good'. Crowdsourced projects such as the Guardian newspapers examination of MPs expense claims[4], the V&A Museum's image cropping[5], Brooklyn Museum's tagging game[6], the National Library of Australia's collaborative OCR corrections[7]; Chen's (2006) study of the application of Csikszentmihalyi's theory of 'flow' to game design; and Dr Jane McGonigal's ideas about multiplayer games as 'happiness engines'[8] all suggest that 'playful interactions' and crowd participation could be applied to help create specific content improvements on museum sites. Game mechanisms may help make tasks that would not traditionally considered fun or relaxing into a compelling experience.

Within the terms of this project, the output of a game-like interaction must produce an effect outside the interaction itself – that is, the result of a user's interactions with the site should produce beneficial effects for other site visitors who are not involved in the original interactions. To achieve this, it must generate content to enhance the site for subsequent visitors. Methods to achieve this could include creating trails of related objects, entering tags to describe objects, writing alternative labels or researching objects – these will be defined during the research phase and creativity workshops.

Methods and tools

The project is divided into several stages, each with their own methodology and considerations.

Research
The preliminary research process involves a literature review, research into game mechanics and the theory of flow, and research into museum audiences online. It will also include a series of short semi-structured interviews with people involved in creating crowdsourced projects on museum sites or game-like interactions to encourage the completion of set tasks (e.g. games for social good) in order to learn from their reflections on the design process; and analysis of existing sites in both these areas against the theories of game design. This research will define the metrics of the evaluation phase.

Creativity workshop(s)
The results of this research phase sets the parameters for creativity workshops designed to come up with ideas and possible designs for the game-like interfaces to be built. Possible objectives for the creativity workshop include:

  • designing methods for building different levels of challenge into the user experience in an environment that does not easily support different levels of challenge when museum-related skills remain at a constant level
  • creating experiences that are intrinsically rewarding to enable 'flow' within the constraints of available content

Build and test
In turn, the creativity workshops will help determine the interfaces to be built and tested in the later part of the project. The build will be iterative, and is planned to involve as many build-test-review-build iterations as will fit in the allocated time, in order to test as many variant interaction models as possible and support optimisation of existing designs after evaluation. User recruitment in this phase may be a sample of convenience from the target age group.

The interfaces will be developed in HTML, CSS and JavaScript, and published on a WordPress platform. This allows a neat separation of functionality and interface design. Session data (date, interface version, tester ID) can be recorded alongside user data. WordPress's template and plug-in based architecture also supports clear versioning between different iterations of the design, allowing reconstruction of earlier versions of the interfaces for later comparison, and enabling possible split A/B trials.

Analysis and write-up
Analysis will include the results of user testing and user data recorded in the WordPress platform to evaluate the performance of various interface and interaction designs. If the platform attracts usage outside the user testing sessions it may also include log file or Google Analytics analysis of use of the interfaces.

[1] https://www.armrev.org
[2] http://www.gwap.com/
[3] http://www.playinterrobang.com/
[4] http://mps-expenses.guardian.co.uk/
[5] http://collections.vam.ac.uk/crowdsourcing/
[6] http://www.brooklynmuseum.org/opencollection/tag_game/start.php
[7] http://newspapers.nla.gov.au/ndp/del/home
[8] http://www.futureofmuseums.org/events/lecture/mcgonigal.cfm

Edward Tufte on 'Beautiful Evidence'

Tonight I went to see Edward Tufte lecture on 'visual thinking and analytical design' at the Royal Geographic Society tonight. The room was packed, perhaps because lots of people were in town for UX London already. The event was organised by Intelligence Squared, who are apparently 'dedicated to creating knowledge through contest', which is an interesting goal in its own right.

It's been a long week, so I suspect my notes are pretty sketchy, but I'm posting them in case they're useful. Let me know if you spot any corrections. On with the talk…

[His basic thesis:] The method of production interferes with production [of knowledge?].

We do most of our serious visual thinking inside, on unreal flatland [2D] screens, looking at representations of real things, instead of being outside looking at real things. In the real world, most of our thinking would be about way-finding, rather than deep analysis.

Evidence is evidence. Information doesn't care what it is. The intellectual tasks remain constant regardless of the mode of production/consumption – we need evidence to understand and reason about the materials to hand. We might care about mode of production but we shouldn't segregate the information by modes of production. We need content-oriented design.

In manuscripts, the hand directly integrates words and image. When marks all come from same source e.g. hands then material is integrated. When technology has different modes of production e.g. type and drawings, then content is segregated by modes of production.

[He showed a 9th century centaur – the image was made of Latin words, a unification of text and image.] Visual meaning of Latin centaur is still clear despite language. The universality of images, the stupefying locality of languages. Images are cosmopolitan, words are local and parochial. When the language is unknown to the viewer, image and language are separated into comprehensible image and incomprehensible language.

'Content indifference' is the result of teaching that only design matters. The essential test is how well each assists understanding the understanding of the content, not how stylish they are. You must know the meaning of the content to design it.

The point of information display is to assist analytical thinking. One common task is to make comparisons. Take the intellectual task (e.g. 'make smart comparisons') and turn them into design principles. Tasks become instructions to the design. Otherwise it's just based on fashion or latest technology.

[He then had the lights turned up so people could view copies of Minard's map of Napoleon's Russian campaign in 1812. He used this to illustrate his Six Grand Principles of Analytical Design.]

  1. First Grand Principle of Analytical Design: show comparisons, contrasts, differences e.g. difference between those who left and those who returned.  Principles guide the design but also the content. A lot of his work is secretly about analytical thinking.
  2. Second Grand Principle of Analytical Design – show causality, dynamics, mechanisms, explanation. For policy thinking, intervention thinking – to produce an effect, you need to know about and govern the cause – show causality.  Thinking about how these are derived are producer commandments. But as they're derived from fundamental intellectual task they're also consumer tasks – you should be asking, as an audience, what is your task?  Minard's map shows causality with temperature.
  3. Third Grand Principle of Analytical Design – show multivariate data. Three or more factors, variables. i.e. show more than two variables. Reality is inherently multivariable.Minard showed six dimensions – the size of army, direction, temperature, dates, location (lat,long). It's not about the method of display. The design is so good that it's invisible. Good web design has to be enormously self-effacing. The task is for users to understand information, not admire the interface. People should be too busy going about their business to notice the design.
  4. Fourth Grand Principle of Analytical Design – the principle of mode indifference – completely integrate all content. No segregation by mode of content. Videos and tables should be embedded in text, not set aside with captions elsewhere. Information don't care what it is, cognitive tasks don't care what the mode of production is. Minard had paragraphs of text, statistical bits, annotations all over.  Cognitive styles in approaching material – really good analysts are indifferent to the mode of evidence, spirit is 'whatever it takes to explain it'. Driven by explanation. Enormous difference between process driven and content driven explanation. Academics make industry from process driven design but it should be whatever it takes to explain it – not what's lying around, not what I'm good at.
  5. Fifth Grand Principle of Analytical Design – document everything and tell people about it. Document sources, scales, missing data. It's the credibility argument. Two things you need to get across in a presentation – what the story is and why the audience should believe you. Audience has two tasks – trying to figure out the story and whether they can believe the presenter. An important way to have credibility is to have care and craft with respect to the data. Minard's two paragraphs are about documentation – assumptions, scales of measurement. Minard was a great engineer who designed bridges and canals – this has the facticity of an engineer. Minard didn't want quibbles about the facts – he wanted people to appreciate the disaster of war. It was meant as an anti-war poster right from the start. It was meant to show the horrors of war. Precision and accuracy in evidence helps his credibility.  Documentation is part of the fundamental quality control mechanism for preservation of credibility. As a consumer, you should be sceptical if people don't say where the data came from e.g. URLs or full data sets. Very few people look at the original data but making sure it's available is important.  Cherry picking is a big threat to the credibility of presenters – am I seeing the results of evidence, or of evidence selection? Providing source is a way of showing you're not a cherry picker.  Overall, incompetence is more likely to be an explanation than conspiracy. Intimacy with the evidence helps convince of credibility.
  6. Sixth Grand Principle of Analytical Design – serious presentations largely stand or fall on the quality, relevance or integrity of the content. 'Just fancy that.' ' If that's not true of presentations where you work, maybe you want to work somewhere else.'  His great insight into design from his books is that content matters, but 'it's a shame we live in a world where that counts as an insight'.The best way you can improve your presentation is to improve the content.
  7. [Then I think a Seventh Grand Principle of Analytical Design snuck in:] We want to try as much as we can to show information as adjacent in space than flip backwards in time. If information is stacked in time e.g. one slide after another, you have to try and make comparisons between something that's gone away and something you're not seeing. Important comparisons should be try to be made in the eye span. This seventh principle has all kinds of consequences for design. Power users have multiple monitors – trying to get more content real estate adjacent in space.

Digression 'strange word, users', only two industries describe their customers as users, illegal drug industry and computer industry.

The principles makes you think about how you display information to viewers – put the data in front of the user, show the comparison. The human eye is really good at comparing things, so having content rich screens and use just about every pixel that you can to carry content.

Trust the users optical capacity. New York Times has 400 links on its homepage. When working on a model for site for site reporting transparency, he suggested news sites.

User testing is not 'having some temps come in and look your screen over' but rather how your website performs in the wild. Know New York Times and Google News work because of the sheer number of visitors. Websites that are very successful in the wild provide models. Look at people who report – first rate news websites.

[He then turned to the sparklines handout.]

Bonus real live 'Powerpoint sucks' comment.
Sparklines – intense, high resolution display.
Graphics should have resolution of typography [not thickness of pencil]. 'Graphics are no longer a special occasion.' They can be anywhere that a word, image or number can be.

Grey band [in the handout ] is based on the types of tasks that clinicians want to do – they don't care about normal values, only exceptions. [The whole visualisation is based on the task, the user and the context]

'We want to be approximately right rather than accurately wrong'. An approximate answer to the right question is better than an accurate answer to the wrong question. [I think I misheard 'accurately' for 'exactly', judging by people's tweets]

We're subject to the recency effect – sparklines show most recent change but put it in context – patterns in change, other similarly shaped changes.

Data analytical design is about showing reality in flatland.

In response to questions: a lot of websites get corrupted because they're pitching, bringing the ethics of the marketplace into the design – important alongside the intensity of a good news website is its spirit – reporting spirit, not pitching.

Qu re current trend for data visualisation e.g. artists animating visualisations
'I think they can do exactly what they please' One reason it's alright is they're not making other claims about the content, just using it as a found object.

[Finally, out of interest – you could compare these written notes to this sketched version of his talk by @lucyjspence.]

The challenge for museums in the 21st century?

Nick Serota, Director of the Tate, writes about modern museums in 'Why Tate Modern needs to expand'. I'm not sure he convinces me that the expansion needs to be physical, but it's a brilliant case for expanding Tate's online presence:

The world also sees museums differently. Wide international access, directly or through digital media and at all levels of understanding offers the opportunity for new kinds of collaboration with individuals and institutions.

The traditional function of the museum has been that of instruction, with the curator setting the terms of engagement between the visitor and the work of art. But in the past 20 years the development of the internet, the rise of the blog and social networking sites, as well as the more direct intervention in museum spaces by artists themselves, has begun to change the expectations of visitors, and their relationship with the curator as authoritative specialist. The challenge for museums in the 21st century is to find new ways of engaging with much more demanding, sophisticated and better informed viewers. Our museums have to respond to and become places where ideas, opinions and experiences are exchanged, and not simply learned.

The museum of the 21st century should be based on encounters with the unfamiliar and on exchange and debate rather than only on an idea of the perfect muse—private reflection and withdrawal from the "real" world. Of course, the museum continues to provide a place of contemplation and of protection from the direct pressures of the commercial and the market. It has to have some anchors or fixed points for orientation and stability, but it also has to be a dynamic space for ideas, conversations and debate about new and historic art within a global context.

The Tate's Head of Online, John Stack, has put the Tate Online Strategy 2010–12, including their 'Ten principles for Tate Online'.  Go read it – with any luck UK parliament will have managed to form a government by the time you're done.

So, do you agree with Serota? What are the challenges you face in your museum in the 21st century?

Slides and talk from 'Cosmic Collections' paper

This is a lazy post, a straight copy and paste of my presentation notes (my excuse is that I'm eight days behind on everything at work and uni after being grounded in the US by volcanic ash). Anyway, I hope you enjoy it or that it's useful in some way.

Cosmic Collections: creating a big bang?

View more presentations from Mia .

Slide 1 (solar rays – Cosmic Collections):

The Cosmic Collections project was based on a simple idea – what if we gave people the ability to make their own collection website? The Science Museum was planning an exhibition on astronomy and culture, to be called ‘Cosmos & Culture’. We had limited time and resources to produce a site to support the exhibition and we risked creating ‘just another exhibition microsite’. So what if we provided access to the machine-readable exhibition content that was already being gathered internally, and threw it open to the public to make websites with it?  And what if we motivated them to enter by offering competition prizes?  Competition participants could win a prize and kudos, and museum audiences might get a much more interesting, innovative site.
The idea was a good match for museum mission, exhibition content, technical context, hopefully audience – but was that enough?
Slide 2 (satellite dish):
Questions…
If we built an API, would anyone use it?
Can you really crowdsource the creation of collections interfaces?
The project gave me a chance to investigate some specific questions.  At the time, there were lots of calls from some quarters for museums to produce APIs for each project, but would anyone actually use a museum API?  The competition might help us understand whether or how we should invest in APIs and machine-readable data.
We can never build interfaces to meet the needs of every type of audience.  One of the promises of machine-readable data is that anyone can make something with your data, allowing people with particular needs to create something that supports their own requirements or combines their data with ours – but would anyone actually do it?
Slide 3 (map mashup):
Mashups combine data from one or more sources and/or data and visualisation tools such as maps or timelines.
I'm going to get the geek stuff out of the way and quickly define mashups and APIs…
Mashups are computer applications that take existing information from known sources and present it to the viewer in a new way. Here’s a mashup of content edits from Wikipedia with a map showing the location of the edit.
Slide 4 (APIs)
APIs (Application Programming Interfaces) are a way for one machine to talk to another: ‘Hi Bob, I’d like a list of objects from you, and hey, Alice, could you draw me a timeline to put the objects on?’
APIs tell a computer, 'if you go here, you will get that information, presented like this, and you can do that with it'.
A way of providing re-usable content to the public, other museums and other departments within our museum – we created a shared backend for web and gallery interactives.
I think of APIs as user interfaces for developers and wanted to design a good experience for developers with the same care you would for end users*.  I hoped that feedback from the competition could be used to improve the beta API
* we didn’t succeed in the first go but it’s something to aim for post-beta
Slide 5: (what if nobody came?)
AKA 'the fears and how to deal with them'
Acknowledge those fears
Plan for the worst case scenario
Take a deep breath and do it anyway
And on the next slides, the results.  If I was replicating the real experience, you’d have several nerve-biting months while you waited for the museum to lumber into gear, planned the launch event, publicised the project in the participant communities… Then waited for results to come in. But let’s skip that bit…
Slide 6: (Ryan Ludwig's http://www.serostar.com/cosmic/)
The results – our judges declared a winner and a runner-up, these are screenshots – this is the second prize winning entry.
People came to the party. Yay! I'd like to thank all the participants, whether they submitted a final entry or not. It wouldn't have worked without them.
Slide 7: (Natalie and Simon's http://cosmos.natimon.com/)
This is a screenshot from the winning site – it made the best use of the API and was designed to lure the visitor in and keep drawing them through the site.
(We didn’t get subject specialists scratching their own itch – maybe they don’t need to share their work, maybe we didn’t reach them. Would like to reach researchers, let them know we have resources to be used, also that they can help us/our audiences by sharing their work)
Slide 8: (astrolabe – what did we learn?)
People need (more) help to participate in a geektastic project like this
The dynamics of a competition are tricky
Mashups are shaped by the data provided – you get out what you put in
Can we help people bring their own content to a future mashup?
Slide 9: (evaluation)
I did a small survey to evaluate the project… Turns out the project was excellent outreach into the developer community. People were really excited about being invited to play with our data.  My favourite quote: "The very idea of the competition was awesome"
Slide 10: (paper sheet)
Also positive coverage in technical press. So in conclusion?
Slide 11: (Tim Berners-Lee):
“The thing people are amazed about with the web is that, when you put something online, you don’t know who is going to use it—but it does get used.”
There are a lot of opportunities and excitement around putting machine-readable data online…
Slide 12: Tim Berners-Lee 2:
But:  It doesn’t happen automatically; It’s not a magic bullet
But people won't find and use your APIs without some encouragement. You need to support your API users. People outside the museum bring new ideas but there's still a big role for people who really understand the data and audiences to help make it a quality experience…
Slide 13 (space):
What next?
Using the feedback to focus and improve collection-wide API
Adding other forms of machine-readable data
Connecting with data from your collections?
I've been thinking about how to improve APIs – offer subject authorities with links to collections, embed markup in the collections pages to help search engines understand our data…
I want more! The more of us with machine-readable data available for re-use, the better the cross-collections searches, the region or specialism-wide mashups… I'd love to be able to put together a mashup showing all the cultural heritage content about my suburb; all the Boucher self-portraits; all the inventions that helped make the Space Shuttle work…
Slide 14: (thank you)
If you're interested in possibilities of machine-readable data and access to your collections, join in the conversation on the museum API wiki or follow along on twitter or on blogs.  Join in at http://museum-api.pbworks.com/
More at https://openobjects.org.uk/ or @mia_out

Image credits include:
http://antwrp.gsfc.nasa.gov/apod/ap100415.html
http://antwrp.gsfc.nasa.gov/apod/ap100414.html
http://antwrp.gsfc.nasa.gov/apod/ap100409.html
http://antwrp.gsfc.nasa.gov/apod/ap100209.html
http://antwrp.gsfc.nasa.gov/apod/ap100315.html
http://www.sciencemuseum.org.uk/Centenary/Home/Icons/Pilot_ACE_Computer.aspx
http://www.prospectmagazine.co.uk/2010/01/mash-the-state/

MW2010 machine-readable data unconference session

This was originally posted on the 'Museums and the machine-processable web' wiki.

This is a rough report from an unconference session on RDFa, microformats and museum data held during Museums and the Web 2010.

I'm writing it up later than I intended (blame the volcano) so please excuse any mistakes in writing up, misattributions, etc – you can sign in to edit them yourself, leave a comment or drop me a line (contact details on the register your interest page).

I'm also writing it up just before I head to the airport, so this first version won't be complete so do jump in and add your own notes if you were there (or wanted to be).

We started by introducing ourselves and briefly describing our interest in the session.

Those present were: Richard Urban, Nate Solas, Paul Hagon, Peter Goodall, Bart Grob (?), Ilya, Piotr Adamczyk, Richard Morgan, Paul Rowe, Darren Scott, Erich Schroeder, Patrick Schmitz, Gunter Waibel…

Interests included: included inference rules based on metadata, embedding metadata in webpages, breaking through the 'analysis paralysis' and choosing a standard to implement (even if it wasn't perfect),

What problems are people having? Picking a standard!

What issues arose during the unconference?

There was an interesting tension between the 'just do it, near enough is good enough' and the 'let's wait until we've got the standard right' impulses – as museum technologists I guess many of us are a mixture of both. But there was also a feeling that we should find a way to move beyond the questions to the point where we start implementing something, with an eye to having a demonstrator project available by this time next year (so April 2011).

We made a useful distinction between a lightweight shared 'standard' that aimed to increase the discoverability of content, and more heavyweight standards that might be used internally or implemented with particular uses in mind. This distinction allows us to keep working through the issues to come up with a suitable (usable, robust, sustainable, implementable, accurate) long-term solution while trying out existing or ad hoc standards in the shorter term.

The voices of reason

One of the reasons I was so happy with this unconference session is that all kinds of people contributed commonsense warnings from their various domains and experiences.  Piotr and Richard said they were still looking for the things that could be done in RDFa that couldn't be done with existing infrastructure.

The use cases

Providing use cases helps everyone understand what we each want to do with the data as well as what we have in our collections.

Peter Goodall wants to make it easy for museums to do mashup collections.

Piotr is still looking for what can be done in RDFa that can't be done with existing infrastructure…

Ilya – neighbourhood project – Open Source Software Foundary – implemented RDFa as a demonstration – FOAF is format to describe social networks and DOPE – description of a project. What kind of aggregation service could we endorse to harvest from our collections?

One of mine: Caroline Herschel (1750 1848) is an astronomer, and there's content about her in lots of museums across the world. I've encountered her in Brooklyn Museum, the National Maritime Museum, the National Portrait Gallery… I'd love to link to images and content from all those other museums from our page about her – but how would I find that content, and how could I reliably link to it?

Erich from Illinois state museum – was working on oral historyproject  on agriculture, indexed to really detailed level – wants to provide user with a proper citation for an interview clip. Found zotero but only got as far as that.

Gunter: OAI-PMH and CDWA-Lite on last project; writing tips for museums working on stuff like this.

FOAF? Richard, V&A – just done collections online with an API that wasn't really standards-based.  Is with Piotr – we should just be able to do this stuff with NLP and text mining – also interested in FOAF.  FOAF sounds like a winner as we know there are people out there lookig for people's names. 

Peter Goodall – large db of people to disambiguate names.  Paul – playing with FOAF – someone made a FOAF generator from their API.  Paul Rowe – NZ museums project – looking at terminologies and overlaps.

Or maybe not FOAF… Patrick from CollectionSpace and UC Berkeley – in past life has done lots of semantic work but has reservations about RDFa. Worries about vocabs e.g. Dublin Core that turns out to be irreconcilable but once embedded make it hard to do more serious things. Interested in reasoning and inferencing across collections.  Ontologies are a point of view, doesn't believe can have a universal point of view.  Use NLP (natural language processing) to index collections from a given community. Interesting to explore more specifically the use cases e.g. compelling cases around events. FOAF doesn't let you model different types of relationships and roles that one person may fulfil. e.g. of how it's hard to shift a community to something more refined once a model is in place.  Potential to generate multiple points of view with different vocabs, use cases will help him understand.

What next? AKA, getting on with it

Testing standards – I'm really up for implementing something on our existing pages – I was thinking that a comparison of two different standards, both marked up as RDFa on existing Science Museum/NMSI web pages (Dublin Core on Ingenious and LIDO on Making the Modern World) , would help provide some useful data on the utility of the approach and the beginning of a comparison between standards.  I've written about it a bit at http://museum-api.pbworks.com/Science-Museum-linked-data – it's a very unfinished document but if you've got suggestions how making it better I'd love to hear them.

[My notes get sketchy from here on it because I'm returning to them after a few months, and some use cases may have ended up in this section, but that's probably ok]

It was suggested that versioning could be a way of dealing with the fact that we don't have a perfect standard right now – it could allow us to iterate through various prototypes and demonstrators until we get something good, while not breaking projects that are built in the meantime.

Microformats – Paul Hagon has used them on event (and other stuff?), Nate pointed out that they're used by Google and Yahoo.

Richard – maybe work on a new 'do one thing' challenge.

Dublin Core is 'messy'.  Patrick: 'is a little better than tagging'.
Peter – interested in using really dumb taxa cos people catalogue inconsistently anyway.
Patrick – taxa even in life sciences don't agree.
Something that's good enough vs something perfect.
Map to shared system with mapping to the authorities used to back things up.
PS: instead of describing a free concept, e.g. a pig, but 'a pig' and when we say pig, we mean it as in this name authority.
GW: identifier-based systems.
How much do we aim for perfection?
PS: don't tie yourself to a syntax that doesn't allow for that.
NS: What can we solve today?
PS: don't want to say figure everything out before you start but consider later options.
NS: let's do something lightweight – add RDFa to marked up pages.
Peter G: interested in something really simple… really interesting thing is the objects – being able to refer to the identity of an object from a pictorial represntation.

LIDO as vocab that works for social history museums and not just art galleries; Dublin Core as quick win.


NS: if we provide enough good enough markup… PA: satisficing approach.
WordNet as term, authority list.
Grappling with issues around how lightweight/heavyweight to go that allows useful exchange of records/assertions.
PS: can I pivot across museums based on some RDFa tags?

[So as you can see, there were no solid conclusions and we didn't leave with an agreement "let's all try implementing x".  I still like the idea of an MW2010 challenge, ideally something you can participate in as a publisher or consumer of data… Suggestions?]

Some thoughts on linked data at the Science Museum – thoughts in progress

I originally posted this on the Science Museum developers blog.

I’ve posted on twitter and my personal blog but forgot to post over here (tsk) – I’ve written some very-much-in-progress thoughts on how the Science Museum could work with linked data/APIs to improve our machine-readable data offerings at the museum data wiki.

I’m particularly interested in finding the balance between a solution we can achieve in the medium-term and something that works with standards as much as possible.

It’s nearly time for the Museums and the Web 2010 conference, where questions like this might be addressed in one of the unconference sessions so I’d love to hear your thoughts.

Additional content from the 'Museums and the machine-processable web' wiki: Science Museum linked data

This is very much a work in progress, and in fact I suspect it's not even the latest version, but hopefully at least it's more useful up here than on my hard drive, even in a very draft-ish state.

February, 2010.

This is a thoughts-in-development piece on how the Science Museum/NMSI could provide re-usable, interoperable, structured machine-readable data for use as linked data or APIs.

I've made it a document rather than a blogging it on http://openobjects.blogspot.com/ or putting it on http://museum-api.pbwiki.com/ directly because it's a bit too long (and probably a bit too incoherent) right now. I'd love to hear your thoughts though – twitter (http://twitter.com/mia_out) or as comments/edits here.

URIs and concepts we could model

Concepts we could model:

  • objects,
  • types of objects,
  • people/organisations,
  • events,
  • places,
  • narratives (stories, themes, topics – typically more subjective, contextualised, interpretive),
  • science subjects (science-y concepts like physics, chemistry, engineering, maths, psychology, astronomy)
  • news stories

Each of these would form part of a URI  e.g. http://sciencemuseum.org.uk/objects/[identifier]

I'm including here things that we generally have enough information about for it to make sense for us to link them. I'll talk about ways to link to the rest of the world below.

Objects – we have lots of these. Yay! Each record is about a specific accessioned object. As you can see from the diagram above, objects can be related to everything else (and to each other, in various ways). An object might be as big and iconic as Robert Stephenson's Rocket or as small as a spark plug.

Types of objects – a more generic view. It allows us to solve two problems – our collections don't cover everything we want to talk about, and we have lots and lots of certain types of objects.  So a page on spark plugs is a user-friendly layer of content about spark plugs for general readers and provides links to all 8000 spark plugs in the collection (I totally made that number up).

It lets us discuss topics that our collections don't cover comprehensively, and to create a user-friendly layer between the detail of our collection (8000 spark plugs) and general information about spark plugs.

[If you're not familiar with museum collections  – coverage varies according to what was collectable or collected – our collections may represent fashions in history of collecting more than an ideal uber-collection. Unlike, say, an art gallery, not every single item in our collection is a precious and unique diamond – for the general user, it might be enough to know what we have some information about dental forceps and a picture of one – but for the specialist researcher, browsing our collection of 300 of them might be the highlight of their week. (Maybe).]

Places – in our collections databases, we can look at the place an object was made, used, designed, destroyed, collected, restored, redesigned, invented, etc, etc. People and events also have various possible relationships to places.

People/organisations – ideally, we'd like to Wikipedia for every person and place, but not everyone we refer to in our collections has Wikipedia notability.  

Images – we also have lots of related images, which are a major asset but work better in relation to other things (like objects) than as concepts on their own.

Other hooks in our content include dates and materials – these might be particularly useful for facetted browsing or mashups made with our data, but don't particularly make sense as concepts on their own. We also produce contemporary science news through our (re-opening in June) Antenna gallery, and marking this up with hNews seems a no-brainer. Working out how to link to the original news stories, whether in Nature, the BBC, whatever, would be good – something we can build into the publishing platform (WordPress MU) to make it nice and easy for our content authors would be even better.

Linking concepts and microsites, creating a canonical object home

I'm proposing a model that should allow us to make the most of all the data we've got online already as well as designing around concepts.

[see notes below for some background]

As well as 'objects' as a basic concept, museums come with a handy set of stable concepts built into our collections management systems.  Sometimes these are called 'subject authorities'.  They cover things like people and organisations, places, events and the relationships between them.  We often build various interpretative narrative layers on top of them – themes, topics, stories, whatever.

If we build permanent URIs around those concepts, we can link to them from the existing microsites. We can also wrap metadata around the elements already on the pages of those microsites so that the data is meaningfully machine-accessible in situ.

As an example, we'd have http://sciencemuseum.org.uk/objects/1956-152 as the 'home page' for the Pilot ACE computer in our collection. This page would contain the basic 'tombstone' information – when, where, what, etc, and link to every known instance of the object in other sites, as below.  These other sites might be exhibitions, subject-specialist sites, cross-institution collections. Often they'll contain information written specifically for that site, particularly tailored for its scope and audiences.

This object is represented in various microsites. The image below shows up we might mark up those sites with links to our Science Museum concepts:

The object home page could also link to the Pilot Ace page on Ingenious and on our Centenary site, and they could link back to the object home. They could also link to our Alan Turing page, National Physical Laboratory page, etc.

It'd be great if we could link to other content about that object – this BBC article on Pilot ACE is a pointer to more content.

Vocabularies

This is one of the places I get stuck… Do we go general or specific? There's lots of stuff out there for visual resources but that doesn't describe our collections well.  There's some discussion of this on various pages here, including Authority Lists, Implementation formats, and RDFa (the names get out of control fairly quickly!).

Notes on URIs

Some of our accession numbers are going to make things difficult because they contain '/'.

The objects we currently have online in this format are divided by collection, which is possibly a less permanent concept, so my preference would be for http://www.sciencemuseum.org.uk/objects/1878-3 rather than http://www.sciencemuseum.org.uk/objects/computing_and_data_processing/1878-3 (1873-3 is the accession or inventory number – these are about as permanent an identifier as you can get [insert museum-y discussion of the exceptions]).

Background-y bits

On Wednesday [you can tell how long ago I started this because that was February 24] I went to the second London Linked Data meetup, held during dev8D.

For a while I've been wondering what we (Science Museum/NMSI) could do with linked data, but it's also taken a while for the issues to bubble up. 

The first two issues are data standards and vocabulary.  As the saying goes, 'the good thing about standards is that there are so many to choose from'.  http://museum-api.pbworks.com/Implementation-formats and http://museum-api.pbworks.com/RDFa bear witness to the difficulties of… finding out what developers prefer to work with (if they care at all), finding out what other museums can output to try and get some critical mass going…

The third is machine-readable interface design.  Tom Scott [Apis and APIs] advocates building APIs so that you're linking people to the concepts that matter to them, and making your website your API.  I think this is the right way to go, but it's made trickier by the fact that we're not a greenfield site – we've got exhibition microsites that are over ten years old.  We're gradually migrating all that data into a central repository, but it'd be good if we could make the data already online in those sites re-usable too.

Other earlier notes… When designing the Cosmic Collections API last year, I'd considered building it into the 'human-facing' website architecture, so that a device could request XML or JSON versions of the pages alongside the (X)HTML pages.  In the end I went for a standalone API as an interim solution.  The Cosmic Collections competition was designed in part to answer some of my questions about the formats preferred by developers.

Comments on the Science Museum linked data wiki page

Comments (18)

Mia said

at 3:02 pm on Mar 21, 2010

Wow – I only tweeted this a few minutes ago and I've had lots of useful feedback.

Jim suggested 'collections' as a concept (http://twitter.com/pekingspring/statuses/10821178733) and he's absolutely right. It'd be great to be able to link our King George III collection (http://www.sciencemuseum.org.uk/onlinestuff/stories/the_king_george_iii_collection.aspx) with that at the British Library (http://www.bl.uk/reshelp/findhelprestype/prbooks/georgeiiicoll/george3kingslibrary.html)

This made me realise I've also completely missed out 'exhibitions' as a concept – we do cover this for current exhibitions to an extent, but there's a lot of information hidden in the choices made for previous exhibitions that could be useful. It also contributes to really making the object home the definitive resource.

Mia said

at 3:06 pm on Mar 21, 2010

Tony (http://twitter.com/psychemedia/statuses/10821277577 http://twitter.com/psychemedia/statuses/10821106102) also suggested 'there's also the design of BBC URI sets; eg if you take a programme episode to be an object, does that lead anywhere?', which is something I'd been thinking about – I really need to finish writing up my notes from the London linked data meetup; and using http://writetoreply.org/ukgovurisets/ 'as a framework for museum/collections URI sets?' – which I hadn't even known about, but will read up on.

Mia said

at 5:39 pm on Mar 21, 2010

More comments:

DavidHaskiya (http://twitter.com/DavidHaskiya/status/10824735162) suggested 're vocabularies:General ones e.g. Geonames, LCSH, VIAF should work for you. A science object theaurus you'll have to do yourselves!' (http://en.wikipedia.org/wiki/Library_of_Congress_Subject_Headings http://www.oclc.org/research/activities/viaf/default.htm)

Wilbert Kraan (http://twitter.com/wilm/statuses/10823953962) suggested 'I'm not an expert in cultural heritage, but CIDOC seems a good, rdf based ontology to adopt or plunder http://cidoc.ics.forth.gr/'

Mia said

at 7:20 pm on Mar 21, 2010

And another comment – can you tell I should be doing something else today? It's all about constructive procrastination.

Richard Morgan from across the road at the V&A commented (http://twitter.com/rmorg/status/10831225400), 'linked data vocabularies tricky for me too. For V&A I'm tending towards just geo, foaf and dbpedia – more about links than data' which I think is a useful perspective. There is a level at which the precise application of term lists matters, but if it means we spend the next ten years trying to get it perfect rather than doing something now, I'd rather we did something now. The two aren't mutually exclusive technically, but pragmatically I only have limited time/brain space in which to get something done.

andy.powell@… said

at 2:14 pm on Mar 22, 2010

Mia,
hi… I think you'll need to model both real-world objects and web documents as part of this. So, for example… for any particular artefact, say the lunar lander, you have the thing itself (a real-world object which is assigned one URI) and the description of that thing (a Web document which is assigned a different URI).

To get from the 'object' URI to the 'description' URI requires an HTTP 303 redirect response (unless you choose to use hash URIs).

The 'description' URI can offer multiple representations, e.g. HTML with embedded RDFa and RDF/XML.

So, if http://sciencemuseum.org.uk/objects/1956-152 is the URI of a real-world object then it does NOT directly serve a representation of that object. Rather, it issues a 303 redirect to a URI that serves representations of that object, e.g. http://sciencemuseum.org.uk/documents/1956-152.

Apologies if you knew this already and I missed it above. I think this applies to most of the entities in the diagram above.

andy.powell@… said

at 2:19 pm on Mar 22, 2010

Sorry… I should have said, "Rather, it issues a 303 redirect to a URI that serves representations of a description of that object, e.g. http://sciencemuseum.org.uk/documents/1956-152.".

Bill Roberts said

at 2:22 pm on Apr 5, 2010

I like your list of "URIs and concepts we could model" and the idea of how the web page about an object in the collection can be linked to relevant people, places, images etc.

There's a lot of scope for this approach to help people to explore the collection from different perspectives and via different dimensions.

Vocabularies: this is an area where it makes sense to re-use existing work where possible, but if there is nothing out there that fits your purpose, don't be afraid to invent a new specialist vocabulary of your own. It's easy (and normal practice) to 'mix and match' terms from multiple vocabularies/ontologies as required.

Mia said

at 6:23 pm on Apr 9, 2010

Thanks for your really useful comments, Bill. I've been horribly busy preparing for a conference next week but will respond properly when my feet are back on the ground!

John S. Erickson, Ph.D. said

at 7:16 pm on Apr 13, 2010

This is an excellent start!

Try to keep in mind that an important reason for publishing the museums artifacts, whether real or digital, is to enable data about them to be "meshed" with other data (from the museum and from elsewhere) and republished, possibly in unanticipated ways, and the "mashed" applications that are created from those datasets. So the answer to whether you are doing it "correctly" will depend on the feedback you get!

The most important thing for you to do is ensure that you make it easy for your community of users to provide you with feedback, wiki a wiki or whatever. Make sure this is obvious and easy, AND that you adapt as they provide that feedback!

You might consider using OpenVocab http://open.vocab.org/ as a means for your community to add new terms.

Good luck!

John

Raj said

at 11:51 pm on Apr 15, 2010

There's already a great authoritative reference for places:
GeoNames Ontology
http://www.geonames.org/ontology/
"over 6.2 million geonames toponyms now have a unique URL with a corresponding RDF web service"

eatyourgreens said

at 11:40 am on Apr 16, 2010

Descriptions depend on context, so might need their own URLs, separate from objects. A record typically has a description that's written for the collections management system
eg. http://www.nmm.ac.uk/collections/explore/object.cfm?ID=BHC0719 but a short label when the object is on display eg. http://www.nmm.ac.uk/visit/exhibitions/past/turmoil-and-tranquillity/gallery/?item=51

Eric Kansa said

at 5:54 pm on Apr 17, 2010

Great discussion of the linked data issues.

I think we can add a point that a RESTful web services (esp. based on simple common standards like Atom) can be useful for bridging between more "Plain Web" design approaches and linked data approaches. Here's a<a href='http://www.alexandriaarchive.org/blog/?p=497'> paper</a> I gave at the Computer Applications in Archaeology conference about this issue.

Eric Kansa said

at 5:55 pm on Apr 17, 2010

OK. Try this again, since HTML doesn't work in the comments.

Great discussion of the linked data issues.

I think we can add a point that a RESTful web services (esp. based on simple common standards like Atom) can be useful for bridging between more "Plain Web" design approaches and linked data approaches. Here's a(http://www.alexandriaarchive.org/blog/?p=497) I gave at the Computer Applications in Archaeology conference about this issue.

Richard Light said

at 12:08 am on Jul 8, 2010

Notes on 7 July 2010 meetup (part 1)

These thoughts are my own "take homes" from the discussion, rather than any sense of the meeting's overall conclusions.

What data do museums have?

Database content, mostly fielded and designed mainly for collections management support. Textual materials, much of it in a
non-accessible "grey literature" format. Images.

The database content is typically (reasonably) self-consistent within a given environment. Thus we have known properties (from the field name) with usable string values. The challenge from a Linked Data perspective is the cost-effective generation of URLs from the string values currently held, e.g. for people and places, given that different museums will have different vocabularies to control their content.

Who wants to use this data?

The public, who are typically interested in classes of objects (rather than individual objects), or in objects with certain properties (e.g. coming from a place of interest to them). Educators, or more specifically people who create resources for educators to use. Students, if relevant objects could be easily accessed as "follow up" to formal learning materials.

Richard Light said

at 12:09 am on Jul 8, 2010

Notes on 7 July 2010 meetup (part 2)
How do we improve the data?

There is nothing to stop every museum publishing URLs, and whatever associated Linked Data they have to hand, for each object in their own collection, and thereby giving them a "hook" onto which others can hang added-value information and assertions of their own. They should treat this task as an urgent priority.

Where possible, convert string values in data to URLs, ideally widely-used (not just local) ones. Could use e.g. geonames.org for place names, or dbpedia for object class names. Interest in Portsmouth's historical gazetteer for "old" place names.

There is a clear need for a sector-specific ontology which represents the properties found, i.e. the types of information recorded in
museum databases. This will act as the "predicate" in Linked Data triples/assertions. It could be based on an existing agreement
about these semantics, e.g. CIDOC CRM or LIDO.

Axis-based data such as geographical co-ordinates or dates/date ranges could be treated as purely numerical data, or "pixellated" by assigning a URL which imposes a certain level of precision (e.g. year for dates). Or both approaches could be adopted.

What's the museum take on Linked Data?

Simple assertions are not enough; we care about the attribution of those assertions (i.e. who is making the assertion). We also want a framework which allows the expression of uncertainty and doubt.

We are not particularly bothered about the specific format (RDF/XML, RDFa, JSON, Topic Maps) in which Linked Data is published, but we would like to be able to "do the job once" and have done with it.

Joshan Mahmud said

at 12:18 am on Jul 8, 2010

Thanks for the minutes Richard – seems like it was a really interesting discussion – shame I couldn't be there – particularly as we've been working with the author of CIDOC to start mapping our data! Look forward to the next meeting. Josh

Shaun Osborne said

at 4:28 pm on Jan 28, 2011

hi Mia

I been wondering about identifiers, pref. UUID types
this sort of fits in where you have [insert museum-y discussion of the exceptions] in your doc.
given we have loads of object numbers full of illegal characters (for both file systems and URIs) I thought the concept of MuseumID may be very helpful as we moved toward linked data..
http://museumid.net/about

Mia said

at 11:21 pm on Jan 31, 2011

Hi Shaun, that's a really interesting proposal, thanks for sharing the link. Do you know wherther ICOM would support it and guarantee permanence?

Cheers, Mia