Showing posts with label genome. Show all posts
Showing posts with label genome. Show all posts

Sunday, 10 October 2010

Biggest Genome ever? It's plant vs amoeba

A few days ago there was an article in Science magazine that they had discovered a plant (Paris japonica) with the largest genome ever found with 149 billion base pairs...

Except, I remembered there's an amoeba (Polychaos dubium - what a bloody weird name by the way!) with an even bigger genome of 670 billion base pairs.

So it got me wondering - how could Science magazine get it so wrong?! Well, this comment on the science mag site gave the reason why they would reject dubium:


"While this and some other Amoeba have been reported to have such very large genomes, some caveats regarding their reliability are perhaps in order.
The measurement for Amoeba dubia and other protozoa which have been reported to have very large genomes were made in the 1960s using a rough biochemical approach which is now considered to be an unreliable method for accurate genome size determinations. The method uses whole cells rather than isolated nuclei and thus will include not only DNA from the mitochondria but also any DNA in engulfed food organisms. Also some of the species are multinucleate.
The accuracy of the genome size estimates are also called into question given that a related species, Amoeba proteus, which was reported to have a genome size of 300 pg was more recently shown to be an order of magnitude smaller (34 - 43 pg DNA per cell). Like the situation for dinoglagellates (see below), the genomes of these Amoeba are clearly large but to know just how big requires their genome sizes to be estimated using modern best practice techniques available today. Only then will it be possible to know just how their genomes compare in size to those of Paris japonica." {Quoted from Ilia}
The reasons given in the quote can be ignored except for the clincher which is in bold. If we assume the same experimental conditions and level of accuracy, dubium will be found to have about 69pg DNA.

Game over, the plant wins.

Thursday, 3 September 2009

At least 100 new mutations in each of us

Each of us has at least 100 new mutations in our DNA, according to research published in the journal Current Biology.

Scientists have been trying to get an accurate estimate of the mutation rate for over 70 years.

However, only now has it been possible to get a reliable estimate, thanks to "next generation" technology for genetic sequencing.

The findings may lead to new treatments and insights into our evolution.

In 1935, one of the founders of modern genetics, JBS Haldane, studied a group of men with the blood disease haemophilia. He speculated that there would be about 150 new mutations in each of us.

Others have since looked at DNA in chimpanzees to try to produce general estimates for humans.

However, next generation sequencing technology has enabled the scientists to produce a far more direct and reliable estimate.

They looked at thousands of genes in the Y chromosomes of two Chinese men. They knew the men were distantly related, having shared a common ancestor who was born in 1805.

By looking at the number of differences between the two men, and the size of the human genome, they were able to come up with an estimate of between 100 and 200 new mutations per person.

Impressively, it seems that Haldane was right all along.

Unimaginable

One of the scientists, Dr Yali Xue from the Wellcome Trust Sanger Institute in Cambridgeshire, said: "The amount of data we generated would have been unimaginable just a few years ago.

"And finding this tiny number of mutations was more difficult than finding an ant's egg in an emperor's rice store."

New mutations can occasionally lead to severe diseases like cancer. It is hoped that the findings may lead to new ways to reduce mutations and provide insights into human evolution.

Joseph Nadeau, from the Case Western Reserve University in the US, who was not involved in this study said: "New mutations are the source of inherited variation, some of which can lead to disease and dysfunction, and some of which determine the nature and pace of evolutionary change.

"These are exciting times," he added.

"We are finally obtaining good reliable estimates of genetic features that are urgently needed to understand who we are genetically."

http://news.bbc.co.uk/1/hi/sci/tech/8227442.stm

Thursday, 6 August 2009

Structure of HIV genome 'decoded'

Scientists say they have decoded the entire genetic content of the HIV-1 virus, a key source of Aids infection.

They hope this will pave the way to a greater understanding of how the virus operates, and potentially accelerate the development of drug treatments.

HIV carries its genetic information in more complicated structures than some other viruses.

The US research, published in Nature, may allow scientists the chance to look at the information buried inside.

HIV, like the viruses which cause influenza, hepatitis C and polio, carries its genetic information as single-stranded RNA rather than double-stranded DNA.

The information enclosed in DNA is encoded in a relatively simple way, but in RNA this is more complex.

We are also beginning to understand tricks the genome uses to help the virus escape detection by the human host
Ron Swanstrom
study author

RNA is able to fold into intricate patterns and structures. Therefore decoding a full genome opens up genetic information that was not previously accessible, and may hold answers to why the virus acts as it does.

The team from the University of North Carolina at Chapel Hill said they planned to use the information to see if they could make tiny changes to the virus.

"If it doesn't grow as well when you disrupt the virus with mutations, then you know you've mutated or affected something that was important to the virus," says Ron Swanstrom, professor of microbiology and immunology.

"We are also beginning to understand tricks the genome uses to help the virus escape detection by the human host."

Deep inside

Dr David Robertson from the University of Manchester welcomed this "definitive analysis".

"What this may reveal is some of the proteins operating at a level below the structures, which may have all sorts of functions within the virus.

"More generally, if we can unpick the structures then we can compare the systems of different viruses and gain new understanding of how they work."

Keith Alcorn of the HIV information service NAM added: "Encouraging the virus to mutate is not a new idea, but it is one of a number of options on the table.

"How important this information will be for the development of new drugs remains to be seen, but it is a useful addition to what we know."

http://news.bbc.co.uk/1/hi/health/8186263.stm

Tuesday, 16 December 2008

Primer 2: Genomes

A genome is the collective genetic material, the DNA or RNA, of an organism. It is all the Nuclotide bases found in DNA/RNA - A(T/U)CG - in the order found in the organism. All organisms on earth use DNA/RNA and if they can be used between species - it's all compatible - you could stick mouse DNA into a fly's genome or even a human's genome and it would do something at least. Almost all cells in the human body have the full set of chromosomes in them - the whole genome.

Prokaryotic-celled species such a bacteria have a single long strand of DNA floating freely in the cytoplasm of the cell. Viruses have an RNA (occasionally DNA) genome floating in their capsule (head).

Eukaryotic-celled species such as mammals have numerous chromosomes with DNA wound onto them. Not all the DNA is useful - in humans only 5% of 3 Billion bases codes for proteins! The coding parts are called Exons. The rest of the DNA is often called "junk DNA" but some people think it still plays an role in the cell but they aren't exactly sure what. Non-coding parts are called Introns.

The first genome to ever get sequenced was that of the virus, a Bacteriophage, a virus that infects bacteria. It had an RNA genome of 3569 bases. Viruses are actually not living organisms but instead are a structure made of proteins in which some genetic material is stored in the head, and when the virus manages to latch on to a cell, it creates a pore in the cell membrane and inserts its RNA genome, which the host cell unwittingly takes and replicates to create more viruses. Viruses tend to cause cell death and the cell can explode or bud-off packages to release the virus copies. No-one is really sure where viruses came from or how they came to be.

The first genome of a living organism to be sequenced was that of the bacteria Haemophilus influenzae in 1995, done using the Shotgun method, as a proof of concept. It had 1.83 Million bases (Mb) on one chromosome!

The shotgun method involves cutting up the whole genome using a restriction enzyme (an enzyme that cuts DNA when it finds a specific base sequence) and then analysing the pieces and then sticking them back together like a puzzle. The reading of the bases and the puzzle solving is done by computers. For example, say we cut the genome and we get lots of strands that are similar to each other but overlap each other - we line them up like so:

ATCCCGGATGCTCTG
-----------ATGCTCTGAAAAAATTCCCCCC

The hyphens represent spaces, to fit the overlap, in order to align the sequences. The first fragment shares some of the code from the second fragment so we assume they come from the same place and so in reality the original DNA strand contains the code ATCCCGGATGCTCTGAAAAAATTCCCCCC. The computer does this lots of times with lots of fragments and eventually everything lines up and it publishes the final sequence.

c.f. Eric D. Green (2001); STRATEGIES FOR THE SYSTEMATIC SEQUENCING OF COMPLEX GENOMES; Nature Reviews: Genetics, Aug 2001 Vol2. p573 (here)

What's amazing is that the process is repeated over and over to fill any inconsistancies and gaps in the data and to validate the original sequencing. Some parts of the human genome were repeatedly sequenced up to 12 times. When a draft sequence is published it is sometimes possible to find an X nucleotide base in the middle of a sequence - this signifies that particular base was either unclear or the data was not collected successfully. It is a long and laborious endevour but well worth it as it has led to the genomics revolution which we are currently at the start of - you will soon see so many discoveries thanks to the availability of the genomes of humans and model organisms such as mice, zebra fish, nematode worms, drosophlia flies and the Arabidopsis plant.

The first draft of the human genome was published in late 2000 and then the final draft was published in 2003. It was dicovered the human genome had between 25000 and 30000 genes which was much less than anyone had predicted since even simpler organisms like the rice plant has about 30000 genes (but a smaller genome). Finally, take a look at the table of species that have been sequenced and look at how huge the genome of the single-celled Amoeba is!!! (here)

A good resource, some slides, although slightly innacurate, can be seen here.