How do proteins develop from long strands of bonded amino acids into the complex shapes that they must take on in order to perform their biochemical tasks? This question has puzzled scientists for decades and not until recently have the necessary resources become available to help solve this problem. The process by which proteins assemble themselves into their final shape has been termed “folding”, which comes from the way that they bend and fold in on themselves to reach their final state (Brown, 1999). This extremely complex process happens in fractions of a second yet it takes an average computer over one day to simulate one-billionth of a second of this activity. With modern day advancements in computer technology this tedious process can be sped up more than one hundred thousand times that of a personal computer (Pande, 2004).
Proteins are long chains of amino acids joined together by peptide bonds. Billions of these chains are produced in the ribosome of cells. Once they leave the ribosome the proteins fold into the shape that is needed for them to perform their final task in the body. Exactly how the proteins fold is called “the protein-folding problem” and has been researched since it was proposed in 1935 by chemist Linus Pauling (Brown, 1999). There are twenty different amino acids that make up proteins and the sequence that they appear in the chain is a crucial factor in how the protein folds (Brown, 1999).
The commonly accepted fundamental basis of how proteins fold lies in the concept of “energy states”. Energy state refers to the protein’s natural tendency to reach the lowest possible tension between all of its atoms (Brown, 1999). A marionette dangling from its strings can be used as an example for this concept. If the marionette is hung up such that the control is level, all of the doll’s appendages will settle so that there is the smallest amount of strain throughout the doll. The process of reaching the minimum energy state begins with the interaction of atoms that are located close to each other on the chain. These atoms will attract or repel each other early on in the folding process and form the basic shape of the final product (Brown, 1999). The remaining atoms of the chain fall into place in the same manner until the lowest possible energy state has been reached. If two proteins have an identical makeup with the exception of two amino acids being switched, the end products will be completely different because the initial fold will completely change the dynamics of attraction and repulsion among the remaining atoms (Brown, 1999).
In order to understand how diseases such as Alzheimer’s and Mad Cow arise scientists need to understand why certain proteins occasionally misfold. Alzheimer’s disease is caused when misfolded proteins form a plaque in the brain; the symptoms are a result of the intended protein being in short supply since it has been folded into a useless form (Thomasson, 2004). The use of computer programs can help find out why these proteins are not folding into the correct end product. Such computer programs try to simulate the way proteins fold by using a trial and error method that folds amino acids close to each other many times and then keeps the fold that results in the lowest energy state. Once the first fold is made the computer reassesses the protein and continues by making multiple trials of the second fold and taking the best result. This process is continued until there are no more folds possible (Brown, 1999).
The process of simulating protein folding is extremely computationally intensive and requires supercomputers that are capable of trillions of calculations per second to make such a task feasible. Distributed computing and massively parallel supercomputers are the two current approaches. Under distributed computing hundreds of thousands of separate computers run specialized software that simulates protein folding. Each computer only computes a small amount of the total dataset and then returns its findings to a centrally located server. Once the entire protein has been folded the results are combined so that scientists can analyze the steps involved in that particular fold. The second method is to use localized supercomputers that are capable of performing calculations millions of times faster than a desktop computer (Ong, 2004). The fastest supercomputer in the world today is the IBM Blue Gene/L which is capable of a sustained performance of 70.72 trillion calculations per second (Anonymous, 2004). Each method has advantages and disadvantages. The use of localized supercomputers allows the process to be closely monitored since it is all contained in a small area; however the cost to purchase and run such computers is beyond what even a well financed study can afford. Distributed computing allows researchers to cheaply fold protein since most computer time is donated, but because it is such a widespread operation it can be very hard to efficiently manage (Ong, 2004).
The time and money requirements of accurate protein folding simulations are the main factor working against scientists. According to Moore’s law which states that computers generally double their computational capabilities every two years, it would seem as if it would take half as much time to fold a given protein (Anonymous 2, 2004). However with increased processing speed the algorithms used to fold the proteins can be refined to include variables that were ignored previously resulting in similar folding times (Ong, 2004).
Ever since the concept of protein folding was introduced in the 1930s biochemists have known that fully understanding how and why proteins fold as they do is critical to understanding many diseases. With the advent of the computer age scientists are now able to simulate these complex chemical processes in detail and use this information in the development of new drugs. As time progresses the information gained from computer simulations of proteins will possibly revolutionize the way we diagnose and treat many diseases.
Anonymous. IBM Blue Gene/L Tops List. IBM. 2 Dec. 2004
<http://www.research.ibm.com/bluegene/>.
Anonymous 2. Moore's law. 24 Nov. 2004. Wikipedia.
<http://en.wikipedia.org/wiki/Moore's_law>.
Brown, David. Deciphering The Message of Life's Assembly. 19 Apr. 1999.
<http://wsrv.clas.virginia.edu/~rjh9u/protfold.html>.
Ong, Emil. Approaches to Protein Folding. 2 Dec. 2004
<http://www.cs.berkeley.edu/~emilong/classes/cs267/assignment0.shtml>.
Pande, Vijay. Science of Folding@Home. Stanford University. 2 Dec. 2004
<http://folding.stanford.edu/science.html>.
Thomasson, Bill. Unraveling the Mystery of Protein Folding. FASEB Office of Public
Affairs. 2 Dec. 2004 <http://www.faseb.org/opar/protfold/protein.html>.