User-generated mashups in natural language?

In case you missed it elsewhere, check out Mozilla Lab's video and blog post on Introducing Ubiquity – 'An experiment into connecting the Web with language'.

It's a framework that brings together lots of the bits of functionality that are available with browser extensions and bookmarklets and lets the user run them with natural language commands. One of the goals is to "enable on-demand, user-generated mashups with existing open Web APIs. (In other words, allowing everyone–not just Web developers–to remix the Web so it fits their needs, no matter what page they are on, or what they are doing.)".

It's a long way from being ubiquitous, but it does show that it's increasingly worth publishing your data in re-usable formats. They show an example of address being picked up from microformats in apartment listings and mapped for the user – that kind of mashup was possible before and they're a huge step forward in themselves, but how many users have the skills and time to do it? Being able to use natural language to pull together and use data could bring mash-ups to the general public in a massive way.

Mashed museum and UK MW 2008 write-up

A report I wrote on 'The 2008 Mashed Museum Day and UK Museums on the Web Conference' is now live on the Ariadne site. I've already reported on most of the sessions and the mashed museum day here, but the opportunity to reflect on the day and write for a different audience was useful. The review really made me appreciate that time and space away from all the noise of every day life in which to learn, try and think is incredibly important, whether you call it a workshop or an away day or something else entirely:

One lesson from the Mashed Museum day was that in a sector where innovation is often hampered by a lack of financial resources, time is a valuable commodity. A day away from the normal concerns of the office in 'an environment free from political or monetary constraints' is valuable and achievable without the framework of an organised event. An experimental day could also be run with ICT and curatorial or audience-facing staff experimenting with collections data together.

The Ariadne issue is packed full of articles I've marked 'to read', so you might also find them interesting.

BBC experimenting with inline links in articles

I noticed the following link when reading a BBC article today:

BBC: We are trialling a new way to allow you to explore background material without leaving the page.
If you turn on inline links, they appear as subtly blue text against the usual grey. Some have icons indicating which site the link relates to (YouTube, Wikipedia), others don't. Links with an icon open the content directly over the article; links without icons open the link in the same window, taking you from the BBC story. Screenshot below:


The 'Read more' links to a page, Story links trial, that says:

For a limited period the BBC News Website is experimenting with clickable links within the body of news stories.

If you click on one of these links, a window will appear containing background material relevant to that word that is highlighted. The links have been carefully chosen by our journalists.

We are doing this trial because we want to see if you enjoy exploring background material presented in this way. It's part of our continuing efforts to provide the best possible experience.

In addition to background material from the BBC News website, we are also displaying content from other sites, including Wikipedia, You Tube and Flickr.

I'd be really interested to know what the results of the trial are, and I hope the BBC share them. I've been thinking about inline links and faceted browsing for collections sites recently, and while the response would presumably vary if the links were only to related content on the same site, it would be useful to know how the two types of links are received.

The story I noticed the link on is also interesting because it shows how content created in a 'social software' way can be (probably wilfully, in this case) misinterpreted:

"Downing Street has been accused of wasting taxpayers' money after making a jokey video in response to a petition for Jeremy Clarkson to be made PM.

A Conservative Party spokesman said: "While the British public is having to tighten its belts, the government is spending taxpayers' money on a completely frivolous project.""

Software with free licenses still has copyright

I'm highlighting this story because it might help to answer institutional issues with the use of open source and Creative Commons licenses. The emphasis below is mine, and it's an American case so local relevance will vary, but the understanding of the importance of recognition or attribution is a milestone.

BBC, Legal milestone for open source:

Advocates of open source software have hailed a court ruling protecting its use even though it is given away free.

The court has now said conditions of an agreement called the Artistic Licence were enforceable under copyright law.

"For non-lawgeeks, this won't seem important but this is huge," said Stanford Law Professor Larry Lessig.

"In non-technical terms, the Court has held that free licences set conditions on the use of copyrighted work. When you violate the condition, the licence disappears, meaning you're simply a copyright infringer.

"Open source licensing has become a widely used method of creative collaboration that serves to advance the arts and sciences in a manner and at a pace few could have imagined just a few decades ago," Judge White said.

The ruling has implications for the Creative Commons licence which offers ways for work to go into the public domain and still be protected.

"This opinion demonstrates a strong understanding of a basic economic principle of the internet; that even though money doesn't change hands, attribution is a valuable economic right in the information economy."

The Age also has an article that might help you make sense of it, Even free software has copyrights: judge

Freebase meetup, London, August 20

As she explains on the Freebase blog, Kirrily from Freebase will be in London for a little while this month and she'd having an informal meet up with Freebase users and those who might be interested to learn more about it:

We'll be meeting at the Yorkshire Grey Pub in Holborn from 6:30pm, having a few drinks, and talking about open data, building communities around free information, mashups, and more. If you're interested, please stop by. There'll be free wifi available, so bring your laptops if you've got them.

You can RSVP on upcoming.org. I'm going because I think Freebase could be really useful for a personal project but also because it's another way of helping people make the most of their digital heritage.

If you don't know much about Freebase, or haven't seen it lately, this video on Parallax, their new browsing interface should give you a pretty good idea of how useful it can be for cultural heritage and natural history data. It's 8 minutes long, and it's really worth taking the time to watch particularly for the maps and timelines, but if you're pressed for time then skip the first two minutes.

You can also get more background at The Future of the Web or Freebase: Dispelling The Skepticism. There are lots of possibilities for museums, archaeology and other cultural content so come along for a chat and a pint.

[Update: if you're not in London but have some questions about Freebase and digital heritage that you think might be useful for discussion or need some context to explain, drop me a line via the form on miaridge.com and I'll take them along.]

It's a good week for search engine gossip

Dare Obasanjo quotes Nick Carr as a lead in to a post on Google's Assault on Wikipedia:

Clearly Nick Carr wasn't the only one that realized that Google was slowly turning into a Wikipedia redirector. Google wants to be the #1 source for information or at least be serving ads on the #1 sites on the Internet in specific area. Wikipedia was slowly eroding the company's effectivenes at achieving both goals. So it is unsurprising that Google has launched Knol and is trying to entice authors away from Wikipedia by offering them a chance to get paid.

What is surprising is that Google is tipping it's search results to favor Knol. Or at least that is the conclusion of several search engine optimization (SEO) experts and also jibes with my experiences.

After looking at some test cases he concludes:

Google is clearly favoring Knol content over content from older, more highly linked sites on the Web. I won't bother with the question of whether Google is doing this on purpose or whether this is some innocent mistake. The important question is "What are they going to do about it now that we've found out?"

It's early days for Knol so maybe the placement of Google search results will settle down over time.

Via other links I found confirmation that '[f]or years, Google's link: command (and see here) has deliberately failed to show all the links to a website.' Old news but I missed it at the time, but since I'd always wondered why the link: thing never seemed to work properly I thought it was worth mentioning.

One step closer to intelligent searching?

The BBC have a story on a new search engine, Search site aims to rival Google:

Called Cuil [pronounced 'cool'], from the Gaelic for knowledge and hazel, its founders claim it does a better and more comprehensive job of indexing information online.

The technology it uses to index the web can understand the context surrounding each page and the concepts driving search requests, say the founders.

But analysts believe the new search engine, like many others, will struggle to match and defeat Google.

Instead of just looking at the number and quality of links to and from a webpage as Google's technology does, Cuil attempts to understand more about the information on a page and the terms people use to search. Results are displayed in a magazine format rather than a list.

From the Cuil FAQ:

So Cuil searches the Web for pages with your keywords and then we analyze the rest of the text on those pages. This tells us that the same word has several different meanings in different contexts. Are you looking for jaguar the cat, the car or the operating system?

We sort out all those different contexts so that you don't have to waste time rephrasing your query when you get the wrong result.

Different ideas are separated into tabs; we add images and roll-over definitions for each page and then make suggestions as to how you might refine your search. We use columns so you can see more results on one page.

They also provide 'drill-downs' on the results page.

Cuil will direct you to this additional information. By looking at these suggestions, you may discover search data, concepts, or related areas of interest that you hadn’t expected. This is particularly useful when you are researching a subject you don't know much about and aren't sure how to compose the "right" query to find the information you need.

I haven't used it enough to work out exactly how it differentiates concepts (tabs) and 'additional information' (drill-downs/categories).

It does a good job on something like the Cutty Sark. Under 'Explore by Category' it offered:

  • Buildings And Structures In Greenwich
  • Sailboat Names
  • Museums In London
  • Neighbourhoods Of Greenwich
  • School Ships

It picked up search results for Cutty Sark whisky and news of the Cutty Sark fire but they weren't reflected in the categories, and the search term didn't trigger the tabs. The tabs kick in when you search for something like 'orange'.

It didn't do as well with 'samian ware' – the categories picked up all sorts of places and peoples, (and randomly 'American Films'), but while the search results all say that it's 'a kind of bright red Roman pottery' that's not reflected in the categories. Fair enough, there may not be enough information easily available online so that 'Types of Roman pottery' registers as a category.

Incidentally, most of the results listed for 'samian ware' are just recycled entries from Wikipedia. It's a shame the results aren't filtered to remove entries that have just duplicated Wikipedia text. The FAQ says they don't index duplicate content I guess the overall site or page is just different enough to be retained.

It might take a while for museum content to appear in the most useful ways, but it looks like it might be a useful search engine for niche content. From the FAQ again:

We've found that a lot of Web pages have been designed with a small audience in mind—perhaps they are blogs or academic papers with specific interests or pages with family photos. We think that even though these pages aren't necessarily for a wide audience, they contain content that one day you might need.

Our job is to index all these pages and examine their content for relevancy to your search. If they contain information you need, then they should be available to you.

It's all sounding a bit semantic web-ish (and quite a bit 'reacting to Google-ish') and I'll use it for a while to see how it compared to Google. The webmaster information doesn't give any indication of how you could mark up content so the relationships between terms in different contexts is clear, but I guess nice semantic markup would help.

Refreshingly, it doesn't retain search info – privacy is one of their big differentiators from Google.

'Annoying adverts affect website traffic'

Via the BCS:

Nearly three quarters – 73 per cent – of internet users clicked away from a favourite website because of an annoying advert, according to research.

The survey, carried out by Opinion Matters for HowTo.tv, also revealed that 59 per cent no longer visited a particular website because of its advertising.

I use AdBlock for a serene and calm web experience, so when I use someone else's computer I'm always amazed at the sheer level of noise on the web and the crappiness of pages plastered with ads.

I was using Add-Art in conjunction with AdBlock, and will again when it support Firefox 3 because it's a lovely idea. If you haven't heard of Add-Art before, check out this Webmonkey article until you can install it on Firefox 3.

Giant squid dissection via live video

I've been watching the recording of the live stream of the first ever public dissection by Museum scientists of a giant squid.

Congratulations to everyone involved at Museum Victoria, it's a great use of technology and a great approach to openness. The explanations were beautifully clear, and did a great job of contextualising the research, the process and the animal itself.

I love the paparazzi-style photo flashes as they rolled the trolley out onto the main floor.

Portable mapping applications make managers happy

This webmonkey article, Multi-map with Mapstraction, about an 'open source abstracted JavaScript mapping library' called Mapstraction is perfectly on target for organisations that worry about relying on one mapping provider.

How many of these have you heard as possible concerns about using a particular mapping service?

  • Current provider might change the terms of service
  • Your map could become too popular and use up too many map views
  • Current provider quality might get worse, or they might put ads on your map
  • New provider might have prettier maps
  • You might get bored of current provider, or come up with a reason that makes sense to you

They're all reasonable concerns. But look what the lovely geeks have made:

The promise of Mapstraction is to only have to change two lines of code. Imagine if you had a large map with many markers and other features. It could take a lot of work to manually convert the map code from one provider to another.

And functionality is being expanded. I liked this:

One of my favorite Mapstraction features is automatic centering and zooming. When called on a map with multiple markers, Mapstraction calculates the center point of all markers and the smallest zoom level that will contain all the markers.

Open source rocks! Not only can you grab the code and have someone maintain it for you if you ever need to, but it sounds like a labour of geek love:

Mapstraction is maintained by a group of geocode lovers who want to give developers options when creating maps.