Sunday, 22 May 2011

Debunking Bad Journalism and the Last Mile Problem

There are some great blogs out there that regularly take on a piece of bad reporting and pull it to pieces, often they focus on their own particular sphere of expertise, such as Ben Goldacre's BadScience.net.
However most of these blogs, forums and other 'Amateur Journalism Police' efforts run into what I call the Last Mile Problem, and that Last Mile is the distance from the offending article to their blog.

Can we counter bad journalism through debunking articles under the nose of the consumer?

What I envision is a browser plugin which allows users to tap into a crowd sourced discussion of the facts in an article. This plugin not only provides context when you read an article, but also allows you to tag things you think are wrong, and then connects you to a community where you can interact with others passionate about that topic to collaborate and work through finding and reporting back the truth of the matter.

This plugin would need to enable the following:

  • Atomic quoting of parts of the article, so that you can precisely deconstruct the problems with it.
  • Ability to cite in counterarguments and proof of what is wrong.
  • Ability to discuss the topic beyond the article itself, such as delving into the sources cited (probably through a discussion community)
  • A robust discussion to avoid Wikiality level of truth, with all work proven and cited
  • Provide an in-article level of context to the reader as to it's reliability, the amount of discussion it's generated or even marking individual facts under discussion with color coding to show how 'true' they are.

The discussion system would need to provide systematic focus on the difference between Fact and Opinion, but allow for the inclusion of both. By bringing together all of the modern developments in both forum and commenting software, such has been discussed in the Knight-Mozilla MoJo challenge, I think something truly effective could result.

The plugin is the tool, but the collaborative discussion is the goal; I hope that a community could build up a Semantic Web of metadata over time. This metadata would gain great depth of reliability through Eigenfactor-like webs of trust leading all the way up from proven science in respectable journals to the quoted snippets in journalistic articles. Clouds of articles, proof trees and links would form around topics, and each new article that is published can be assimilated into the collected knowledge.
Then any kind of user can connect in to find out what is real and what is fake - like an upgraded Wikipedia that you consume with your news.


This gets the controversy in front of the reader at the time of consumption, and it subverts the bad journalism coming out of some publications which degrade trust in the whole media industry.

First to Post beats Best Post

There's a specific problem with a lot of comment systems: You have to post early to get eyeballs.

This problem is with systems that list posts chronologically or only offer an uncollapsed single threaded view with the most popular point at the top (so most of them!). Then as you go down it can be difficult to spot other posts which have spawned their own interesting discussion - valid arguments are often pushed too far down the stack to be seen by most people, or even onto a separate page where seen by even less people.

Use of discussion monitoring metrics to track which comments are generating discussion and then pushing them up to form 'Roots' of a new discussion thread on the topic, all collapsible to allow easy navigation, would go a long way to combating this issue.

Having more than one Root to the discussion presented at the top of the page represents a small UI challenge of making it easy to see the shape of the discussion and allow a reader to work down the threads that interest them. Potentially this could be solved by combining a diagram of the discussion Tree with a heat map of the activity, then allowing the user to click on a hot section to skip to reading that Root and thread.

Techniques to spot potential discussion spawning late-coming posts could be based around the history of the poster, checking for indicators of 'effort' in the post such as links and cited sources. New posts could be arbitrarily bumped up the reading view for users known to moderate wisely in order to check the value of the content.

More than one way to rank a comment

Why do we need to have only a single ladder for ranking comments; 'like' or 'dislike' (youtube), +1 or -1 (slashmod) - the list is really long but the result is all the same: A single ranked ladder which is exposed to the user.
What's more, this only accounts for active systems - what about exposing passive systems like the number of times a comment has been viewed, collapsed, expanded or generally interacted with outside of our active ranking system? Some sites do this (Youtube and others) but provide little ability to directly interact with the statistic.

Where does it say we can't combine the best of them and enable some truly emergent behavior?

We need to track multiple vectors of community interest and influence on a comment or thread in order to model behavior and seek to improve signal to noise ratios.

The 'Vectors' can be the combinations of trackable metrics, both passive and active, across a discussion thread and individual comments.
As a simple example: Let's say we combine the various read/expand/collapse metrics into a 'Heat' vector, take Youtube's like/disklike meter and slashdot's SlashMod system of +1/-1.

One of the primary actions that can be taken from this is looking for Tension patterns across the vectors. Tension patterns could be situations where a comment gets lots of negative and positive nudges and is also modded as insightful. Perhaps the system can allow for recognised bias within the community and look for situations where it's being exerted systematically.
Tension generating Threads can then be pushed up the visibility stack to ensure robust discussion.

What if you could apply filters on each of these Vectors on the page as you read it through AJAX controls? Control your own signal to noise preference without loading a different page. Add the ability to remember presets and you can build yourself a set of views to flick through and gain a wide understanding of the discussion.


How would you combine these Vectors to empower the community towards better discussion?