Friday, 8 September 2017

Some highlights of Genome 10k and Genome Science 2017

Everyone comes away with different highlights from a conference, as we each see it from the perspective of our own research, but here were some of the highlights for me of Genome 10k and Genome Science 2017. The conference had multiple tracks so there were many talks I missed. The conference was also live tweeted by many under hashtag #g10kgs2017.
  1. New technology: Nanopore sequencing was mentioned by many speakers, but often from people who were just about to use it for the next step in their research. Long reads were mentioned frequently, and PacBio was still a contender for long read sequencing in talks and posters. Optalysys was advertising a hardware approach to sequence alignment, using light detected after passing through two images representing the sequences ("comparison as fast as the speed of light", except for the time it takes to refresh the images). 
  2. Assembly of long reads and assembly analysis: Those who were using long reads were often using this to produce a whole genome, and were therefore attempting assembly, though many fragments remain even with long reads. Canu was mentioned regularly during the talks, as was FALCON, and miniasm during informal chat. John Davey's talk describing detective work to understand the genome of red algae extremophile Galdieria sulphuraria stood out. After assembly he counted chromosomes by looking for telomeres and the end-of-chromosome read alignment, and still found puzzling questions: Could this 14Mbp organism have 72 chromosomes? Do some of them share regions? 
  3. Haplotyping: Sam Nicholls gave an excellent talk about the Metahaplome and resolving haplotypes in a metagenome. PacBio users now have FALCON-Unzip to phase diploid genomes assembled with FALCON.
  4. Comparative genomics, genome alignments and lineage tracing: After the genomes are assembled, the eukaryote researchers are busy comparing their species with other species (often using Cactus). This seemed to be a common topic. Comparing genomes across species is compute-intensive/expensive. Within-species comparison/alignment was discussed by Bernardo Clavijo, who had many wheat genomes to merge and used skip-mers for the alignment seeds. Graph genomes were mentioned briefly but are not yet used. Alternative RNA splicing was a topic of interest for several speakers.
  5. GC content and GC biases: seem to be responsible for everything, including undersampling for short read sequencing, missing genes in the fat sand rat, photosynthetic efficiency and the efficacy of natural selection. Steve Kelly's talk was fascinating. Plants need different amounts of nitrogen for photosynthesis and this corresponded to the GC content and codon usage of their genomes. So photosynthetic efficiency can be predicted by GC content. He went on to describe how increasing atmospheric CO2 will lead to increased mutation and speciation rate in plants.
  6. Animals with superpowers: Researchers studying animals have all the best stories. There were bats that live forever and don't get cancer (well at least 43 years), mice that can have their fur or limbs removed and have no problem regenerating them, Tasmanian devils transmit cancer by biting and passenger pigeons were once the most abundant bird in the US, migrating in flocks so dense that the sky was darkened, but are now extinct.
  7. Sketches by Alex Cagan: He drew each of the talks in the main room at lightning speed, uploading the sketched to Twitter immediately after each talk. The main room concentrated on the eukaryote genomes, so if you're interested in his summaries, see them all at https://twitter.com/search?q=%23g10kgs2017%20atjcagan
Other observations:
  • The sponsors and stall holders were mostly selling lab automation of one kind or another. Clearly, automation is what genomics researchers are likely to buy.
  • "We are a chemostat for microbes" - Lindsay Hall
  • "Strain resolution is not clustering but deconvolution" - Chris Quince
  • PerkinElmer have put a lot of work into understanding the biases of 16S kits for different hypervariable regions. 
  • Several presentations that confessed "Most of this work was done by my student" (then let your student give the presentation!). Also several people suggesting a useful paper by X where X was the last author rather than the first author. Give credit to the first author!

Friday, 4 August 2017

Joy plots

Today is the birthday of John Venn (of Venn diagrams fame). He was born on this day, 4th August 1834. So I was thinking about diagrams today and finally got around to looking up code for the trendiest-plots-of-2017, 'Joy plots'. As of July 2017 we now have ggjoy plotting code in R (https://github.com/clauswilke/ggjoy) and they're also beautiful in Python's Seaborn (https://github.com/mwaskom/seaborn/pull/1238). 

They're named after the Joy Division album cover. But finding out what 'Joy Division' actually means (https://en.wikipedia.org/wiki/House_of_Dolls) is always going to make me less joyful about the joy in using them.

Thursday, 25 May 2017

Releasing research software

How should researchers be releasing their software?

The BBSRC together with the Software Sustainability Institute and Elixir-UK recently held a workshop to find out. This was the "Developing Software Licensing Guidance for BBSRC Workshop" (April 2017). The BBSRC wanted to ask the community what guidance it should be providing for grant applicants, grant reviewers and the developers and users of research software.

Here's some of the points that were raised during the day: 

Ownership

Who owns the software? University academics often have contracts that say that everything we do belongs to the uni, even if it's in the evenings or weekends. Students, on the other hand tend to own their own IP. On collaborative projects, especially those collaborating with businesses, the ownership of software becomes more complex again. If the software is truly open source, then hopefully multiple people will contribute ('random strangers on the internet'). Who owns it then?

Licenses

Three main options: very permissive, copyleft or commercial. If you need to choose a software license, some universities are well supported by technology transfer staff who understand all the implications and some are not. There are websites to help people choose, but the issues are complex and so far they haven't helped me choose. The consensus for permissive licenses seemed to be that while MIT was good (simple, easy to read and understand), the Apache 2.0 license actually deals with accepting contributions from other developers in future ("Unless You explicitly state otherwise, any Contribution intentionally submitted for inclusion in the Work by You to the Licensor shall be under the terms and conditions of this License, without any additional terms or conditions."). If you don't use this license then you may have to use a contributer licensing agreement and that can scare off the casual patch submitter. The "for academic use only" kind of license was widely seen to be awful, restricting future collaboration, restricting project expansion and being unclear about who is and isn't an academic.

Software is produced at different scales

There are quick scripts, solid code, and whole platforms. The BBSRC will be funding work that produces all three types of software. It wants all grant-holders to think about how they will release their software and support the creation of software. We know that documenting, testing and packaging is time consuming. How do we ensure that projects allow time for this? How can software be cited properly (see Software citation principles)? Is there any such thing as a 'throwaway script'? Will the project management team include someone who knows about software?

Existing advice

There is a collection of existing advice from various places about how to choose licenses and release software.
The release of software as an output of research is clearly an issue that is being raised by multiple organisations now and it's great to see software being taken seriously. 2017 will see the second Research Software Engineers Conference (RSE2017) and a June 2016 Dahgstul Workshop produced an Engineering Academic Software Manifesto, containing pledges such as
  • I will make explicit how to cite my software.
  • I will cite the software I used to produce my research results.
  • When reviewing, I will encourage others to cite the software they have used.
and more. The BBSRC now has the task of pulling together all the discussion from the workshop and other places and creating a guidance document to help grant applicants, reviewers and panellists, and also the software developers. This will then become part of the grant proposal process along with other docs such as the data management plan, pathways to impact, justification of resources, etc.

Thursday, 27 April 2017

Ada Lovelace and her contribution to computer science

Ada Lovelace (1815-1852) is known as the first computer programmer. What did she really contribute to computing? So many blog posts, articles and books get bogged down describing her childhood, and her famous father and overbearing mother. But what did she actually do?

A while ago I got stuck in and actually read the paper she actually wrote.  "Sketch of the Analytical Engine invented by Charles Babbage... with notes by the translator. Translated by Ada Lovelace"

In fact, she translated into English a paper by an Italian engineer, Luigi Menabrea, and clearly felt that his description did not go far enough, because she added her own notes at the end, which amount to more detail than the original translation.

The original paper by Menabrea was supposed to sell the idea of the Analytical Engine, a machine designed by Charles Babbage but not yet constructed. Money needed to be raised for this endeavour. Menabrea explains what the machine can do. To some extent he does try also to sell the idea, but his writing can be quite dry: 'To give an idea of this rapidity, we need only mention that Mr. Babbage believes he can, by his engine, form the product of two numbers, each containing twenty figures, in three minutes.'

Lovelace wanted to add more, and her writing is glowing with excitement about the machine. She added 7 detailed discussion notes, labelled A to G, and also 20 numbered footnotes. In the numbered footnotes she comments on issues with the translation. She also comments sometimes on how she feels that Menabrea hasn't quite understood completely the novelty and potential (for example 'M. Menabrea's opportunities were by no means such as could be adequate to afford him information on a point like this'). In the notes labelled A to G she explains many of the concepts that are the fundamentals of computing today.

Note A

This note spells out that the Analytical Engine is a general purpose computer. It doesn't just calculate fixed results, it analyses. It can be programmed. It's flexible. It links 'the operations of matter and the abstract mental processes of the most abstract branch of mathematical science'. This is not just a Difference Engine, but instead offers far more. She is very clear that 'the Analytical Engine does not occupy common ground with mere "calculating machines". It holds a position wholly its own'. It's also not just for calculations on numbers, but could operate on other values too. It might for instance compose elaborate pieces of music. 'The Analytical Engine is an embodying of the science of operations'. Lovelace has a beautiful turn of phrase that evokes her delight in imagining where this could go: it 'weaves algebraic patterns just as the Jacquard loom weaves flowers and leaves'.

Note B

This note is about memory, the 'storehouse' of the Engine. It is about variables, and retaining values in memory. She describes how the machine can retain separately and permanently any intermediate results, and then supply these intermediate results as inputs to further calculations. Any particular function that is to be calculated is described by its variables and its operators, and this gives the machine its generality.

Note C

This note is about reuse: the Engine uses input cards, as used in the Jacquard looms. Lovelace had toured the factories of the north with her mother and seen the Jacquard looms in action. They would have been the cutting edge of technology at the time. Intricate patterns were woven by machines, with thousands of fine threads, all controlled by punched cards. The cards needed to be created by card punching machines, with detailed tables of numbers that would be translated into the patterns of holes. The cards allowed a design to be repeated, and they were bound together in a sequence. She explains that these input cards, can be reused 'any number of times successively in the solution of one problem'. And furthermore, that this applies to just one card, or a whole set of cards. The train of cards can be rewound until they are in the correct position to be used again.
Jacquard loom, Nottingham Industrial Museum

Punched cards used in Jacquard loom, Nottingham Industrial Museum
 

Note D

This note is about the order and efficiency of operations. She explains that there are input variables, intermediate variables, and final results. Every operation must have two variables 'brought into action' as inputs and one variable to use as the result. We can then trace back and inspect a calculation to see how a result was obtained. We can see how often a value was used. We can think about how to batch up operations to make them more efficient. She begins to imagine not just how the machine will operate, but how programmers will have to think carefully about their programs, how they will debug them and how they will optimise the algorithms they implement.

Note E

In this note Lovelace explains how loops or 'cycles' can be used to solve series, for example to sum trigonometrical series. With a worked example (used in astronomical calculations), she abstracts out the operations needed to produce terms in the series and works on a notation for expressing cycles that include cycles. Note 18 expands further to explain that one loop can follow another, and there may be infinitely many loops.

Note F

This note begins with the statement 'There is in existence a beautiful woven portrait of Jacquard, in the fabrication of which 24,000 cards were required.' Babbage was so taken with this woven portrait that he bought one of the few copies produced. The amount of work that machines could do, unfailingly, without tiring, was changing the face of industry. The industrial revolution was sweeping through the country. Lovelace explains how the Engine will be able to solve long and intensive with a minimum of cards (they can be rewound, and used in cycles). The machine can calculate a long series of results without making mistakes, solving problems 'which human brains find it difficult or impossible to work out unerringly'. It might even be set to work to solve as yet unsolved and arbitrary problems. 'We might even invent laws for series or formulae in an arbitrary manner, and set the engine to work upon them and thus deduce numerical results which we might not otherwise have thought of obtaining; but this would hardly perhaps in any instance be productive of any great practical utility, or calculated to rank higher than as a philosophical amusement.'
 

Note G 

She concludes with the scope and limitations of the Analytical Engine. She is keen to stress the machine's limitations and about using the machine for discovery: it doesn't create anything. It has 'no pretensions to originate anything'. Lovelace then enumerates what it can do, but warns against overhyping it. Alan Turing references her concerns in his 1950 paper Computing Machinery and Intelligence (which is also well worth a read). In this paper he discusses 9 objections that people may raise against the possibility of AI. One of these 9 objections is Lady Lovelace's Objection. He writes:
Our most detailed information of Babbage's Analytical Engine comes from a memoir by Lady Lovelace (1842). In it she states, "The Analytical Engine has no pretensions to originate anything. It can do whatever we know how to order it to perform" (her italics).
In Note G we also get the detailed trace of a substantial program to calculate the Bernoulli numbers, showing the order of operations combining variables to make results. The trace shows looping and storage, showing intermediate results, and the correspondence with the mathematical formulae being calculated at each step. She inspects the program for efficiency, working out the number of punched cards required, the number of variables needed and the numbers of execution steps.

Lovelace wonders, in the third paragraph of this Note, whether our minds are able to follow the analysis of the execution of the machine. Will we really be able to understand computers and their abilities? 'No reply, entirely satisfactory to all minds, can be given to this query, excepting the actual existence of the engine, and actual experience of its practical results.'

Ada Lovelace is buried in Hucknall Church, Nottingham. She died aged 36, and never got to see the Analytical Engine constructed.

Monday, 10 April 2017

Visiting JG Mainz University

I've just returned from a two week visit to Johannes Gutenberg University Mainz, where I was hosted by Andreas Karwath in the Data Mining group of the Dept of Informatik. I was there to pick up new research ideas, to get away from admin/teaching duties for a while and to share what we've been working on lately. 

The University at Mainz is a large campus based university on the outskirts of the town, on the beautiful river Rhine. It's named after the inventor of the printing press in the west, Johannes Gutenberg. The Gutenberg museum is excellent, tracing books from wax tablets through handwritten parchments, to moveable type print, then the industrial revolution, typewriters and high volume printing. This is a museum about the value of information and its transmission, and the technology invented to do this. There was even a special exhibition about the Futura font, as an added bonus. This is the geometric circles-and-lines font used in posters for "2001 A Space Odyssey", and on the Apollo 11 moon plaque ("We came in peace for all mankind"), and in so much future-looking advertising and propaganda in the Art Deco inspired 1930s era, both in Germany and beyond.

It was interesting to be embedded in a data mining group, rather than bioinformatics for a change. They have a broad range of application areas, but also happily switch technology (neural nets, relational learning, topic models, matrix decomposition, graphs, rs-trees and more) as the application area needs. Also very interesting to see in which ways a different country's research culture is different. It's not REF-dominated like the UK, so more they're more free to focus on quick-turnaround peer-reviewed compsci conference publications, and perhaps more hierarchically structured, as only the few professors have permanent positions. And yet it's still the same. University departments are international places and share much in common, whichever country you're in: same grant applications, student supervisions, seminar talks, dept silos, etc.

It was great fun to be there, and they were excellent hosts. At the same time, it was strange and sad to be a British person on exchange in Germany during the week that the UK sent article 50 to the EU. International collaborations are so important to research that leaving the EU is bound to be hugely detrimental to us UK academics. We need more exchange, not less.

Tuesday, 4 April 2017

The metahaplome

A sample of water, soil or gut contents will contain a whole community of microbes, cooperating and competing. We now have the sequencing technology to begin to explore these communities, to find out the variation that they possess. This can be useful in the search for new anti-microbials, or in the search for better enzymes for biofuels. However, the sequencing technology is not quite there yet. The very short length of the reads, together with the errors introduced, combine to make the problem of reassembling the underlying genomes much more complex.

We introduce the concept of the 'metahaplome': the exact sequence of DNA bases (or "haplotype") that constitutes the genes and genomes of every individual present. We also present a data structure and algorithm that will recover the haplotypes in the metahaplome, and rank them according to likelihood.

Our preprint about the metahaplome is now available at bioRxiv: Probabilistic Recovery Of Cryptic Haplotypes From Metagenomic Data.

Monday, 20 March 2017

Is academic impact useful as a proxy measure for real world impact?

In academia we have various measures of "success". Some of the more common measures are "how many papers did we publish in reputable journals" and "how many other academics went on to use my work or refer to my work". We count citations, check out our h-index, and perhaps even check the number of other academics talking about our work on social media. We'd all like our work to be useful to others. Of course, any measure of success can be gamed. Academics may cite themselves to boost citation counts, use clickbait paper titles to attract attention, and select journals for prestige rather than availability to the community.

However, to be useful to other academics is not the same as to be useful to the rest of the world. Are we also having impact outside of academia? Some blue sky basic research is unlikely to do this (but perhaps likely to be cited by academics). Some applied research can have immediate impact on medical outcomes, law and policy, societal attitudes, civil rights, environmental strategies and business practices.

The UK tries to measure this kind of research impact by asking universities to submit REF Impact case studies. These are summary documents, written by academics, describing what impact they've had. They're not easy to write, or to evidence, and yet they are used as part of the REF exercise measuring research quality, whereby Higher Education funding bodies distribute research money to the universities.

How can we make it easier for academics to find and present their impact? Before we answer that, is the whole exercise worth doing anyway? We need to know if we could just cheaply count citations and use this as a proxy for real world impact. After all, if a paper is popular with academics, surely it's also going to be useful to the rest of the world too?

We've done the analysis in Measuring scientific impact beyond academia: An assessment of existing impact metrics and proposed improvement. TL;DR: the answer is 'no'. Academic impact is not correlated with more comprehensive impact in the real world. We're not surprised, but we needed to prove it. But don't take our word for it: we've made all the data available. Try it yourself.

On with the next steps now... James will now be data mining the real world impact of scientific research, from the collection of unstructured documents out there in news archives, parliamentary proceedings, etc. We expect NLP, data mining and machine learning challenges ahead as we trace the movement of academic ideas and results out from academia and into the rest of the world.

Wednesday, 3 August 2016

Holding an internal research workshop

We have just held the 4th Aberystwyth Bioinformatics Workshop. It's a one-day workshop, held with no budget, and intended to be a mostly internal informal research networking event.

We call for 5 min lightning talks, 20 minute longer talks, demos of software, and posters. We end up with a good mixture of both. We especially encourage new PhD students to present, and for all attendees to be friendly and supportive rather than combative in their questions. Registration is done by a very simple Google form (name, email, what kind of talk, title of talk/poster, any other comments). Registration closes one week before the workshop. Tea and coffee is acquired somehow, a room is booked, talks are arranged into a programme, and then away we go.
Aber Bioinformatics Workshop attendees July 28th 2016. Photo by Sandy Spence.

 Each time we've done this we have ended up with a full day of talks. People use it to let others know what they're working on, to practise a talk they're preparing for an external conference, to ask for advice on their work, to describe the state of the compute cluster facilities and to just introduce new people. Bioinformatics at Aberystwyth is mostly done within the biology departments of IBERS, but this meeting allows Computer Science and Maths people to join in, and make interdisciplinary links. Finally we go down to the pub, and continue the discussions there.

It's a very low cost minimal preparation way to bring together a group of otherwise independent researchers. Many bioinformaticians feel that they are either the only one in their group, or else, that they're not really a bioinformatician at all and somehow masquerading as one. I've learned a great deal from each workshop that we've had, and its just great to find that we do have a surprisingly strong local support network in such a specialist field.

Yr Eisteddfod

Dw i'n dysgu Cymraeg. I'm a Welsh learner (still making many mistakes). Wythnos diwetha es i i'r Eisteddfod i helpu yn yr Pafiliwn Gwyddoniaeth a Thechnoleg (dydd Gwener i dydd Sul). Last week I went to the Eisteddfod to help in the Science and Technology Pavilion (Friday to Sunday). Mae hi'n fy Eisteddfod gyntaf. It was my first Eisteddfod. Bendigedig! Bydda yn mynd eto blwyddyn nesaf. Fantastic! I'll go again next year.




Thursday, 16 June 2016

Aros yn Ewrop : the metagenome of our countries

My thoughts on the parallels between our work in metagenomics and the referendum on June 23rd 2016.

A community of species
sharing cross-genomic pieces,
co-existing in the rumen
work to chew the grass around.

A community of nations
with historical relations
can assemble a consensus
to agree on common ground.

The variation that we're seeing
gives an accent to each being
and to the metagenome union,
and the medley swings in sound.

Vote to stay, aros yn ewrop!
Exchange people, plans and workshops.
In a globe of sequence differences
we can learn from each one found.

Wednesday, 20 April 2016

Gregynog Statistical Conference 2016

The Gregynog Statistical Conference is a long running conference, now in its 52nd year. This conference has been running since 1965. The conference has such a long history that its origins predate box-and-whisker plots, bootstrapping and the R language. But statistics has clearly been relevant and important for the last 52 years and will no doubt remain so for the next 52 years.

Gregynog Hall, where the conference is held every year, is in the heart of mid Wales. It is a beautiful old mansion bequeathed to the University of Wales by the Davies sisters, and now used for conferences, music festivals and educational activities, such as our computer science undergraduate weekends away.


This year the conference main themes seemed to be modelling of epidemics, using variants of S-I-R models, MCMC and Markov models in general. Another topic for discussion was p-values, following the Friday evening after-dinner talk on this subject by David Colquhoun. The statistical power of experiments and meta-studies to combine data from smaller studies was also a recurring theme. Some of the talks I enjoyed were by Ruth King, who described how to include time spent in each state (dwell time) in a Markov model, and Simon Spencer who explained S-I-R epidemic models and went on to use MCMC and importance sampling to estimate his model parameters. Also Chris Jewell, who described the challenges of modelling vector-borne disease outbreaks in cattle in New Zealand, while at the same time providing real-time advice to government on how to manage the course of the disease.


The poster session was a little haphazard. Somehow the posterboards hadn't arrived so posters were bluetacked to the cupboards, blackboards and walls. But the range of topics was good, from Sam Nicholls' work on modelling the metahaplome in metagenomics, to students from Warwick working on the approximation of integration and partial derivatives using Gaussian functions, and a meta-analysis of studies on delayed rewards and delayed penalties (receiving £10 today instead of £20 next week, vs minus £10 today instead of minus £20 next week).

Hopefully another new statistics lecturer will be joining our maths department shortly, as we're recruiting at the moment. Statistics underlies almost every area of research now, particularly in the sciences. We do need to make sure that we keep talking to the expert statisticians regularly.

Tuesday, 5 April 2016

Lovelace Colloquium 2016

This year's Lovelace Colloquium was held at Sheffield Hallam University, last week (March 31st). Sheffield Hallam proved to be a great venue. It's convenient for most people in the UK to get to, with a smart building right by the train station, providing a large poster-exhibiting hall, a modern lecture theatre and a cafeteria, all next to each other.


Every year the Lovelace is an inspiring event. I've now been (and blogged) in 2012, 2013, 2014 and 2015 and it gets bigger and better each year. I'm particularly impressed by the first year undergraduates who are up there presenting posters alongside everyone else, talking to employers and thinking about their future careers.


I didn't get to attend many of the talks this year because I spent more time on the desk and doing organisational jobs (and fixing my posters-numbering error, oops!). But there were some really strong posters covering a wide range of computer science topics, including several with live Arduino demos. The end of the day panel session featured questions and advice on where the field is going in the future, the pros and cos of a career in industry vs academia, the challenges of running your own business and questions about recruitment.

This year I was also impressed to chat to many interesting people during the evening social, for example Claire and Emily from Relish Learning. After graduating from uni they worked for others until deciding one day that they could do it themselves. They set up their own business in Sheffield, and now provide digital e-learning, for a wide range of topics. They described how they've been recently training people in the Army on how to change the wheel on a tank (imagine animations of the components required, and the order in which to remove parts, etc). They are now keen to help others to succeed, to encourage them to believe that they can and to talk to others about how they did it.

If you are not recommending this event to your women undergrads in computer science, then they are missing out. Poster presenters get expenses refunded and may come away with a prize, thanks to all the sponsors. Employers are keen to meet them, so they will also come away with contacts to help them apply for a job or placement. The photos of the event give a great impression of what it's really like if anyone needs any further reassurance.

Other summaries of the day:

Monday, 28 March 2016

Goldilocks: census your genomes

Goldilocks is a new tool written by Sam Nicholls for counting interesting properties of genomes. It's very easy to install ("pip install goldilocks") and has a detailed user manual.

So, let's have a look at GC count across each of the chromosomes of Sorghum. Sorghum is a plant that is a reasonably close relative to Miscanthus, which is extensively studied here in Aberystwyth. I downloaded the chromosome assembly of sorghum from RefSeq. Here's the plot, showing amount of GC on the y-axis and position along the chromosome on the x-axis. The 10 Sorghum chromosomes are all shown stacked up in one plot panel.
The Python code for this using Goldilocks to do this plot is as simple as:

sequence_data = { 
    "sorghum" : {"file": "./sorghum.fna.fai"},
}
g = Goldilocks(GCRatioStrategy(), sequence_data, length="500K",
               stride = "1000K", is_faidx = True)
g.plot("sorghum", title="GC content of sorghum chromosomes")

The dip in GC for the centromere of each chromosome is obvious, except for chromosomes 2 and 6.

A similar but inverted pattern can be seen if we look at the number of Ns along the genome:

So, what's different about the centromeres in chromosomes 2 and 6? Why are they not so visible? Another way to spot them would be to look for a motif known to be in the centromeres. Centromeres have many repeats, and a repeat region known to be found in sorghum centromeres is CEN38. Let's choose a short motif from the sequence for CEN38, say "CCTAATG", and census that.

There's clearly plenty of this motif found in chromosomes 2 and 6, and found where we might expect a centromere to be (also lots of this motif in the centromeres of chr 3 and chr 5 too). But it's not found in all chromosomes. Could it be that CEN38 varies its sequence in the other chromosomes, and so doesn't have precisely that motif? Or that too many Ns in the other chromosomes stop CEN38 being characterised?

This is just a simple demonstration of how Golidlocks can be used to explore questions. And questions lead to more questions, and then many a happy hour can be spend browsing your genomes. Goldilocks can also be used to export details about which regions are the most interesting (hence the name: it finds regions that are "just right", for whatever your "just right" criterion might be).

Enjoy browsing your genomes! Goldilocks paper, Goldilocks docs, Goldilocks source code.

Saturday, 30 January 2016

Playful coding: computing activities for schools

In many schools, computing is a topic that needs more encouragement. The Playful Coding project wants to make practical activities that can be run in schools to explore ideas in computer science. I've just been along to one of their meetings and seen it in action. It was extremely inspiring to be in a room full of people who didn't see running computing engagement activities as a chore, but as fun. They had all put a lot of thought into making fun activities and all wanted to run their activities with the groups of children.
 

It's an EU project involving teachers and university researchers from Spain, Romania, Italy, France and us in Aberystwyth, Wales. Each project partner had developed several activities and the purpose of the meeting was to tested out many of these activities on children and their teachers, and to start to develop a guide for teachers to explain how to use them. Until that guide is produced, you can still browse the activities and have a go with them. Try out for example:
There are lots more activities to choose from. Some just take an hour, some take a day, and some span a term. Some use robots, some use no particular equipment. They can be embedded into other lessons (languages, maths, science, art), or just standalone. And they are adaptable for different age ranges.

To follow the project see the Playful Coding website, follow #playfulcoding on Twitter or find Playful Coding on Facebook.

Wednesday, 23 December 2015

Seamless gene deletion

2015 is the year that genome editing really became big news. A new technique, "CRISPR/CAS", was named as Science magazine's breakthrough of the year as voted by the public from a shortlist chosen by staff.

However, people have been manipulating DNA through many useful methods long before CRISPR/CAS made headlines. Gene deletion is an important tool when trying to understand the function of genes. Take out a gene and see what effect it causes. Genes can be disrupted (by removing a portion of the DNA or inserting some extra DNA) or can be interfered with, for example via their RNA production, or they can be entirely deleted. It's common practice when removing a gene to insert a marker, so that we can easily select for the cells where this procedure has been successful. For example, to insert an antibiotic resistance gene as a marker, so that we can now grow the cells on a plate with an antibiotic. Then only those that have lost our gene of interest and gained antibiotic resistance will now grow. The trouble with this is that many gene deletions have no visible effect by themselves. If we also want to delete a second gene and a third, then we need more markers, or we need to be able to remove and reuse the marker we inserted. We also don't want the process to leave any scars behind that could destabilise the genome. We've just published a paper to help solve this problem.

This process of 'swap a gene of interest for a marker gene' can be achieved in many organisms by homologous recombination. This is a process used by many cells to repair broken strands of DNA. If we provide a piece of DNA that has a good region of similarity to the region just downstream of our gene of interest, and also a good region of similarity to the region just upstream of the gene of interest, but instead of the gene of interest, has the marker gene between these regions, then the normal cellular processes of homologous recombination will exchange the two. Some organisms perform homologous recombination very readily (S. cerevisiae for example). Others may need a little more encouragement, such as creating a double stranded break.

Our new paper A tool for Multiple Targeted Genome Deletions that Is Precise, Scar-Free and Suitable for Automation with Wayne Aubrey as first author uses a 3-stage PCR process to synthesise a stretch of DNA (a 'cassette') that will do everything. It will have good regions of similarity to the regions upstream and downstream of the gene of interest. It will contain a marker gene. And (here's the good bit), it will contain a specially designed region ('R') before the marker gene that is identical to the region that occurs just after the gene of interest. In this way, after homologous recombination has done its thing and inserted the DNA cassette instead of the gene of interest, there will be two identical R regions, one before the marker gene, and one after the marker gene. Sometimes the DNA will loop round on itself, the two R regions will match up and homologous recombination will snip out the loop, including the marker gene.



We can encourage this to happen and select for the cells that have had this happen if our marker is also 'counter-selectable'. That is, we'd like a marker for which we can add something to the growth medium so that now only cells without the marker will now grow. That is, we'd like to use a marker or marker combination for which we can first select for its presence and then counter-select for its absence. When we have this we can select for cells that have had the marker replace the gene, and then counter-select for cells that have now lost the marker too. So we have a clean gene deletion.

Of course we're always standing on the shoulders of giants when we do science. Our method is an improvement on a method by Akada 2006, so that no extra bases are lost or gained and the method requires no gel purification steps. Just throw in your primers and products and away you go. It's not fussy about quantity. No purification steps means that it could be automated on lab robots. And it could be used to delete any genetic component, not just genes. Give it a try!

Thursday, 17 December 2015

Data science and a scoping workshop for the Turing Institute

In November I went to a workshop to discuss the remit of the Alan Turing Institute, the UK national institute for data science, with regard to the theme of "Data Science Challenges in High Throughput Biology and Precision Medicine". This workshop was held in Edinburgh, in the Informatics Forum, and hosted by Guido Sanguinetti.

The Alan Turing Institute is a new national institute, funded partly by the government, and partly by five universities (Edinburgh, UCL, Oxford, Cambridge, Warwick). The amount of funding is relatively small compared with that of other institutes (e.g. the Crick) and seems to be enough to fund a new building next door to the Crick in London, together with a cohort of research fellows and PhD students to be based in the new building. What should be the scope of the research areas that it addresses and how should it work as an institute? There are currently various scoping workshops taking place to discuss these questions.

Data science is clearly important to society, whether it's used in the analysis of genomes, intelligence gathering for the security services, data analytics for a supermarket chain, or financial predictions for the city. Statistics, machine learning, mathematical modelling, databases, compression, data ethics, data sharing and standards and novel algorithms are all part of data science. The ATI is already partnered with Lloyds, GCHQ and Intel. Anecdotal reports from the workshop attendees suggest that data science PhD students are snapped up by industry, ranging from Tesco to JP Morgan, and that some companies would like to recruit hundreds of data scientists if only they were available.

The feeling at the workshop seemed to be a concern that the ATI will aim to highlight the UK's research in the theory of machine learning and computational statistics, but risks missing out on the applications. The researchers who work on new and cutting edge machine learning and computational statistics don't tend to be the same people as the bioinformaticians. The people who go to NIPS don't go to ISMB/ECCB. And KDD/ICML/ECML/PKDD is another set of people again. These groups used to be closer, and used to overlap more, but now they rarely attend each others' conferences. Our workshop discussed the division between the theoreticians who create the new methods but prefer their data to be abstracted from the problem at hand, and the applied bioinformaticians, who have to deal with complex and noisy data, and often apply tried and tested data science instead of the latest theoretical ideas. To publish work in bioinformatics generally requires us to release code and data, and to have shown results on a real biological problem. To publish in theoretical machine learning or computational statistics, there is no particular requirement for an implementation of the idea, or to demonstrate its effectiveness on a real problem. There is also a contrast between the average size of research groups in the two areas. Larger groups are needed to produce the data (people in the lab to run the experiments, bioinformaticians to manage and analyse the data, and these groups are often part of larger consortia) whereas the theoreticians are often cottage-industry style research with just a PI and a PhD student. How should these styles of working come together?

Health informatics people worry about access to data: how to share it, get it, and ensure trust and privacy. Pharmaceuticals worry about dealing with data complexity, such as how to analyse phenotype from cell images in high throughput screening, having interpretable models rather than non-linear neural networks, and how to keep up with all the new sources of information, such as function annotations via ENCODE. GSK now has a Chief Data Officer. Everyone is concerned about how to accumulate data from new bio-technologies (microarrays then RNA-seq, fluorescence then imaging, new techniques for measuring biomarkers of a population under longitudinal study). Trying to keep up with the changes can lead to bad experiment design, and bad choices for data management.

There was much discussion about needing to make more biomedical data open-access (with consent), including genomic, phenotypic and medical data. There seemed to be some puzzlement about why people are happy to entrust banks with their financial data, and supermarkets with purchase data, but not researchers with biomedical data. (I don't share their puzzlement: your genetic data is not your choice, it's what you're born with, and it belongs to your family as much as it belongs to you, so the implications of sharing it are much wider).

All these issues surrounding the advancement of Data Science are far more complex and varied than the creation of novel and better algorithms. How much will the ATI be able to tackle in the next five years? It's certainly a challenge.

Tuesday, 29 September 2015

An executable language for change in biological sequences

A discussion on Twitter about whether there was a language for representing sequence edits prompted me to post my draft proposal for such a language. http://figshare.com/articles/Draft_proposal/1559009

Comments, criticism, collaboration and competition welcome. Hopefully I'll submit it shortly.

Saturday, 5 September 2015

Notes from workshop on Computational Statistics and Machine Learning

I've just attended "Autonomous Citizens: Algorithms for Tomorrow's Society", a workshop as part of the Network on Computational Statistics and Machine Learning (NCSML). That's an ambitious title for a workshop! Autonomous Citizens are not going to hit the streets any time soon. The futuristic goals of Artificial Intelligence are still some way off. Robots are still clumsy, expensive and inflexible. But AI has changed dramatically since I was a student. Back in the days when computational power was more limited, AI was mostly about hand-coding knowledge into expert systems, grammars, state machines and rule bases. Now almost any form of intelligent behaviour from Google translation to Facebook face recognition makes heavy use of computational statistics to infer knowledge.

Posters: there were some really good posters and poster presenters who did a great job of explaining their work to me. In particular I'd like to read more about:
  • A Probabilistic Context-Sensitive Model of Subsequences (Jaroslav Fowkes, Charles Sutton): a method for finding frequent interesting subsequences. Other methods based on association mining give lots of frequent but uninteresting subsequences. Instead, define a generative model, then go on to use data and EM to infer the parameters of the model. 
  • Canonical Correlation Forests (Tom Rainforth, Frank Wood): a replacement for random forests that projects (a bootstrap sample of) the data into a different coordinate space using Canonical Correlation Analysis before making the decision nodes.
  • Algorithmic Design for Big Data (Murray Pollock et al): Retrospective Monte Carlo. Monte Carlo algorithms with reordered steps. There are stochastic steps and deterministic steps. The order can have a huge effect on efficiency. His analogy went as follows: imagine you've set a quiz with a right answer and a wrong answer. People submit responses and you need to choose a winner. You could first sort them all into two piles (correct, wrong) and then pick a winner from the correct pile (deterministic first, then stochastic). Or you could just randomly sample from all results until you get a winner (stochastic first). The second will be quicker.
  • MAP for Dirichlet Process Mixtures (Alexis Boukouvalas et al): a method for creating a Dirichlet Process Mixture model. This is useful as a k-means replacement where you don't know in advance what k should be, and where your clusters are not necessarily spherical.
Talks: these were mostly full talks (approx one hour), but then we had a short introduction to the Alan Turing Institute with Q&A at the end.

The first talk presented the idea of an Automated Statistician (Zoubin Ghahramani). Throw your time series data at the automated statistician and it'll give you back a report in natural language (English) explaining the trends and extending a prediction for the future. The idea is really nice. He has defined a language for representing a family of statistical models, a search procedure to find the best combination of models to fit your data, an evaluation method so that it knows when to stop searching, and a procedure to interpret/translate the models and explain the results. His language of models is based on Gaussian processes with a variety of interesting kernels, together with addition and multiplication as operators on models, and also allowing change points, so we can shift from one model combination to another at a given timepoint.

The next two talks were about robots, which are perhaps the ultimate autonomous citizens. Marc Deisenroth spoke about using reinforcement learning and Bayesian optimisation as two methods for speeding up learning in robots (presented with fun videos showing learning of pendulum swinging, valve control and walking motion). He works on minimising the expected cost of the policy function in reinforcement learning. His themes of using Gaussian processes, using knowledge of uncertainty to help determine which new points to sample were also reflected in the next talk by Jeremy Wyatt about robots that reason with uncertain and incomplete information. He uses epistemic predicates (know, assumption), and has probabilities associated with his robot's rule base so that it can represent uncertainty. If incoming data from sensors may be faulty, then that probability should be part of the decision making process.

Next was Steve Roberts, who described working with crowd sourced data (from sites such as zooniverse), real citizens rather then automated ones. He deals with unreliable worker responses and large datasets. People vary in their reliability, and he needs to increase accuracy of results and also use their time effectively. The data to be labelled has a prior probability distribution. Each person also has a confusion matrix, describing how they label objects. These confusion matrices can be inspected, and in fact form clusters representing characteristics of the people (optimist, pessimist, sensible, etc). There are many potential uses for understanding how people label the data. Along the way, he mentioned that Gibbs sampling is a good method but is too slow for his large data, so he uses Variational Bayes, because the approximations work for this scenario.

Finally, we heard from Howard Covington, who introduced the new Alan Turing Institute which aims to be the UK's national institute for Data Science. This is brand new, and currently only has 4 employees. There will eventually be a new building for this institute, in London, opposite the Crick Institute. It's good to see that computer science, maths and stats now have an discipline-specific institute and will have more visibility from this. However, it's an institute belonging to 5 universities: Oxford, Cambridge, UCL, Edinburgh and Warwick, each of which has contributed £5million. How the rest of us get to join in with the national institute is not yet clear (Howard Covington was vague: "later"). For now, we can join the scoping workshops that discuss the areas of research that are relevant to the institute. The website, which has only been up for 4 weeks so far, has a list of these, but no joining information. Presumably, email the coordinator of a workshop if you're interested. The Institute aims to have 200 staff in London (from Profs to PhDs, including administrators). They're looking for research fellows now (Autumn 2015), and PhDs soon. Faculty from the 5 unis will be seconded there for periods of time, paid for by the institute. There will be a formal launch party in November.

Next year, the NCSML workshop will be in Edinburgh.

Thursday, 30 July 2015

ISMB/ECCB 2015

ISMB/ECCB 2015 (and HitSEQ 2015) was held in Dublin, just across the Irish Sea from us here in Aberystwyth. So off we went, to find out the latest research in bioinformatics. There were many parallel tracks, but the recurring themes of the talks I attended were:
  • lots of work on human genomics, particularly disease, particularly cancer
  • single cell analysis, finding variation (SNVs) from clonal populations, haplotype resolution
  • sequencing technologies: RNA-seq, sequencing of methylation, Hi-C sequencing, ultra deep sequencing and lots of promise for long reads
  • reference sequences: most people were working with a reference rather than de-novo
  • training bioinformaticians, maintaining software, keeping a core of bioinfomaticians
  • the Burrows Wheeler transform - does it solve every large-data problem?
  • graphs, and ways of cleaning up graphs, adding weights to graphs, finding minimal/maximal components of graphs
There wasn't very much about the following topics, though they did make occasional appearances:
  • text mining
  • metagenomics
  • multi-omics
Too hard? Dropping out of favour? Or perhaps people working in these areas just don't attend this particular conference?

Aberystwyth PhD students with their posters: Stefani Dritsa, Sam Nicholls, Tom Hitch, Francesco Rubino

The keynote talks tended to be of the kind that long-established group PIs do well. They're the "Here's a summary of all the work my group has been doing for the past 5-10 years to answer this particular biological question" talk. While I admire their determination and group size, I feel that they're speaking only to a subgroup of the audience with this kind of talk, and that a keynote should somehow also aim to more generally inspire the audience to go out and do great work, have new ideas, think in new directions, and not just to have learned a little more about that specific subject area. The far more off-the-wall non-keynote talk by David Searls about a bioinformatic analysis of James Joyce's book Ulysses fascinated the audience, and provided exactly that. He received a huge round of applause.


The jobs notice boards were full (below are just 2 of the notice boards). More bioinformaticians are clearly needed!


Many of the conference talks are now online, but you need to be a member of ISCB to see them http://www.iscb.org/ismb-mm/media-ismbeccb2015. The papers are also collected in a special ISMB/ECCB issue of Bioinformatics.

Monday, 8 June 2015

Aber Bioinformatics Workshop

Last week we had the 2nd Aber Bioinformatics Workshop. It's an internal workshop for work-in-progress talks, posters and networking and the aim is for us all to keep up with what's going on in Aberystwyth in bioinformatics across departments and institutes. We had a wide range of talks on genomics and sequence analysis, metabolomics, optimising proteins, population and community modelling, data infrastructure and other topics. Here's the programme for the day.

Photo of all the attendees, taken by Sandy Spence
It was great to see that we now have so many people interested and working in bioinformatics, despite the difficulties in trying to understand all sides of the story (the biology, the computing, the statistics, etc). We talked about the range of modules and courses that were available to help people get up to speed with this, and how we should do more to let new PhD students know what is available. Also, now that we've had the workshop, hopefully we're more aware of the expertise and facilities available here in Aber, so we now know who to approach with questions and ideas.

At the end of the day we moved down to the pub, and continued to discuss more random topics: beetles, plant senescence, hens, temperature sensing wires for computer clusters, and concordance in Shakespeare texts. I'm sure this all helps in the long run.