Background: There are no published data on peanut sensitization in Egypt and the problem of peanut allergy seems underestimated. We sought to screen for peanut sensitization in a group of atopic Egyptian children in relation to their phenotypic manifestations.
Methods: We consecutively enrolled 100 allergic children; 2-10 years old (mean 6.5 yr). The study measurements included clinical evaluation for site of allergy, possible precipitating factors, consumption of peanuts (starting age and last consumption), duration of breast feeding, current treatment, and family history of allergy as well as skin prick testing with a commercial peanut extract, and serum peanut specific and total IgE estimation. Children who were found sensitized to peanuts were subjected to an open oral peanut challenge test taking all necessary precautions.
Results: Seven subjects (7%) were sensitized and three out of six of them had positive oral challenge denoting allergy to peanuts. The sensitization rates did not vary significantly with gender, age, family history of allergy, breast feeding duration, clinical form of allergy, serum total IgE, or absolute eosinophil count. All peanut sensitive subjects had skin with or without respiratory allergy.
Conclusions: Peanut allergy does not seem to be rare in atopic children in Egypt. Skin prick and specific IgE testing are effective screening tools to determine candidates for peanut oral challenging. Wider scale multicenter population-based studies are needed to assess the prevalence of peanut allergy and its clinical correlates in our country.
The prevalence of peanut allergy is not sufficiently studied in many developing countries including Egypt. Peanut allergy was estimated to affect 0.8% of children and 0.6% of adults in the US, showing a twofold increase over a 5-year period [1,2]. A recent 11 year follow up survey showed that the prevalence of peanut allergy in children in the US in 2008 was 1.4% compared to 0.8% in 2002 and 0.4% in 1997 [3]. In the United Kingdom, the total estimate for clinical peanut allergy was1.5% of 3-4 year-old children [4]. Relevant studies estimated the prevalence of peanut allergy to be 1.34% among primary school children in a Canadian province [5] and 1.15% among 3 year olds in the Australian capital territory with a trend for rise in prevalence between 1997 and 2005 [6]. Peanut was the third most common sensitizing allergen in an Asian community especially in young atopic children with multiple food hypersensitivities and a family history of atopic dermatitis [7].
Studies to address the reasons for increased prevalence and persistence of food allergies, focusing primarily on peanut, have included the hygiene hypothesis; changes in the components of the diet, including antioxidants, fats, and nutrients, such as vitamin D; the use of antacids, resulting in exposure to more intact protein; food processing, such as for peanut roasting and emulsification to produce peanut butter compared with fried or boiled peanut; and extensive delay of oral exposure, thus increasing topical (possibly sensitizing) rather than oral (possibly tolerizing) exposure to food allergens [8,9].
The evaluation of a child with suspected allergy to peanut should include a careful history taking, skinprick testing (SPT), measurement of serum-specific IgE, and, confirmation by an oral food challenge [10,11]. The prevalence of confirmed food allergy based on confirmatory tests is lower than perceived allergy which is based on self report [12,13]. Diagnostic cut-off values for SPT and specific IgE results have improved the diagnosis of food allergy and thereby reduced the need to perform oral food challenges [14]. Overall, a negative peanut SPT has a negative predictive value of more than 95% [15]. The positive predictive value, however, is significantly lower, reaching only 60% in patients with a convincing history of an allergic reaction [16]. Published values on positive SPT and specific IgE values vary from one series to another depending on several factors [12,[17][18][19][20]. Cut-off values do not always have general acceptability and sometimes need to be individualized in the context of clinical impression [21].
There is an impression that peanut allergy is uncommon in Egypt and there are no published data on its incidence or prevalence. We sought to investigate the frequency of peanut sensitization and allergy in a group of atopic Egyptian infants and children in a pilot attempt to uncover its importance as an allergen in our country.
This cross sectional study comprised 100 children diagnosed to have allergic diseases. They were enrolled consecutively after getting informed oral consent was obtained from the parents or care-givers. The study protocol gained approval from the local ethics committee.
Inclusion criteria: -Age at enrollment between one and 18 years.
-A physician made diagnosis of allergic diseases including asthma, allergic rhinitis, urticaria, and eczema.
-Patients who cannot stop antihistamine therapy.
-Extensive skin lesions, scars, positive dermatographism, and very dark skin.
-Treatment with systemic corticosteroids for more than 7 days.
-Other chronic or debilitating illness.
All patients included in the study were subjected to the following:
Detailed history was taken for the possible precipitating factors, peanut consumption (starting age and last consumption), duration of breast feeding, and family history of allergy. Patients were subjected to a general clinical examination, as well as chest, skin, and ENT examination to verify the diagnosis.
Serum total IgE was measured by enzyme linked immunosorbent assay (Genzyme Diagnostics, Medix Biotech Inc, San Carlos, CA, USA). A serum IgE level was considered elevated if it exceeded the highest reference value for age [22]. The value used in correlation analysis was the percentage from the highest normal value for age (patient's actual value/highest normal value for age multiplied by 100) Peanut Specific IgE was measured in children with positive peanut SPT results using the CLA allergenspecific IgE Assay according to the manufacturer's instructions (Hitachi Chemical Diagnostics, Mountain View, California 94043, USA). The concentration of ≥ 15 kUA/L was considered positive.
Complete blood counting was done using an automated cell counter (Coulter MicroDiff 18, Fullerton, CA, USA) and manual differential.
Skin prick test (SPT) was performed for each patient using a commercial peanut allergen extract, positive histamine control, and negative control (Omega Laboratories, Montréal, Canada). First generation short-acting antihistamines were avoided for at least 72 hours and second generation antihistamines were avoided for at least 5 days before testing. The test sites were marked and labeled at least three cm apart to avoid the overlapping of positive skin reactions. The marked site was dropped by the allergen and gently pricked by sterile skin test lancet. Positive and negative control solutions were similarly applied. The patient waited for at least 20 minutes before interpretation of the results. Largest and orthogonal diameters of any resultant wheal and flare were measured. A wheal diameter of 8 mm or greater was considered positive.
Children with proven peanut sensitization were subjected to open oral peanut challenges under close medical observation taking all the precautions needed to treat anaphylaxis. A second informed consent was obtained from the parents or care-givers prior to the challenge. Children were given gradually increasing amounts of roasted peanuts at 30 min intervals and symptoms and physical signs were closely monitored. The total dose of peanut before considering that the challenge was negative was 15 grams of roasted whole peanuts. An open feeding of a larger portion (ageappropriate serving) followed negative challenges and then children were kept under observation for 2 more hours. The cases with negative challenges were supplied with a contact phone number to report any reactions that might develop within the next 24 hours while cases with positive challenge were treated and kept in hospital for 12-24 hours under observation.
Data were analyzed by a standard computer program (SPSS version 13 for Windows, Chicago IL, USA). The mean, standard deviation (SD), median, and interquartile (IQ) range presented the descriptive data. Groups were compared using the students t-test for parametric and the Kruskal-Wallis and Mann-Whitney Z tests for nonparametric data. Fisher's Exact and Chi square (X 2 ) tests were used for comparison of categorical data. Pearson and Spearman coefficient tests were used to correlate the numeric data. For all tests, p values less than 0.05 were considered statistically significant.
The studied sample comprised 55 boys and 45 girls. Their ages ranged between two and 10 years [median (IQR) = 6.28 (4.0); mean (SD) = 6.51 (2.35) years]. The duration of exclusive breast feeding ranged from 3 to 8 months [median (IQR) = 6.00 (2.00); mean (SD) = 5.42 (1.32) months] and the age of stoppage of breast feeding ranged from 9 to 24 months [median (IQR) = 16.00 (7.00); mean (SD) = 16.30 (4.26) months]. None of the subjects gave a history suggestive of peanut allergy and all of them started consuming roasted peanuts before the age of two. The duration since last peanut consumption ranged between one and 15 days [median (IQR) = 4.00 (3.00); mean (SD) = 4.63 (8.94) days]. The diagnoses included bronchial asthma in 63 children, urticaria in 57, allergic rhinitis in 22, atopic dermatitis in five, and history of anaphylaxis due to unknown cause in one patient. Fifty six children had one, 40 had two, and four had three of the aforementioned allergic diseases.
Skin prick testing with peanut extract gave positive results (wheal diameter ≥ 8 mm) in seven children (7%). The specific IgE results of these children confirmed sensitization [range = 17 -24; median (IQR) = 21.0 (4.00); mean (SD) = 20.9 (2.41); 95% CI = 16.9 -24.5 kUA/L]. Six out of the 7 peanut sensitized patients consented for an open oral challenge with roasted whole peanuts. Three out of the six children showed immediate allergic manifestations after consumption (two developed urticaria and respiratory manifestations and one developed urticaria only) and were thus proven to have allergy to peanut. The remaining three developed no symptoms or signs upon oral challenge with peanuts (Table 1). Neither the positive nor negative OFC cases developed late phase reactions to peanut. Two of the peanut allergic children were brothers (Table 1; patients 1 and 2). They had bronchial asthma and attacks of urticaria and the elder one also suffered from allergic rhinitis. The peanut specific IgE levels could not be correlated to any of the clinical or laboratory data of the sensitized children. Children with positive oral challenge were statistically comparable to those with negative results as far as their clinical and laboratory data are concerned.
Seventy percent of the studied sample gave a positive family history of allergy with no statistically significant relation to peanut sensitization. However, the 7 children sensitized to peanut had positive family history of allergic diseases. None of the subjects gave a history suggestive of peanut allergy in the family. Peanut SPT results did not vary according to gender. Positive results were obtained in 5 out of 55 boys and 2 out of 45 girls. Peanut sensitization rates were not influenced by the duration of exclusive breast feeding, age at complete weaning from the breast, last peanut consumption, serum total IgE level, or peripheral blood eosinophil count (Table 2).
The SPT results were not influenced by the target organ affected whether respiratory or cutaneous. Also, peanut sensitization did not vary according to the number of target organs affected in the studied sample (X 2 = 2.714; p = 0.257). Ten children had confirmed allergy to other foods (egg allergy in two, fish in three, cow milk in two, sesame in one, banana in one, and prunes in one); 9 of them were not peanut sensitized while one was sensitized to peanut and allergic to bananas. The relation between peanut sensitization rates and the presence of other food allergies did not reach statistical significance (X 2 = 0.154; p = 0.695).
The wheal diameter of the peanut sensitized children was not correlated to age at complete weaning from the breast, days since last peanut consumption, days since last antihistamine consumption, serum total IgE, serum specific IgE, peripheral blood eosinophil count, or histamine wheal diameter.
The frequency of peanut sensitization in our series was 7%. The sensitized children were subjected to open oral challenges except for one child whose parents did not consent to the challenge due to a past history of anaphylaxis of unknown etiology. Only three patients were proven to be allergic to peanut by oral challenge giving a 3% rate of peanut allergy in the studied sample. The studied sample comprised children with physiciandiagnosed allergy and therefore does not represent the general population. The Peanut specific IgE was only estimated in the 7 children with positive SPT for confirmation of sensitization. It would be worthwhile to estimate its expression in children with negative SPT. However, this was not among the objectives of the current study.
Population based data from some other parts of the world show different rates of sensitization ranging between 3.3% in England [4] and 4.6% among children from the Netherlands [23]. Cross sectional random telephone surveys revealed self reported peanut allergy in 0.6-1.4% in the US [2,3], 0.93-1% in Canada, and 1.5% in the UK [12]. The prevalence of peanut allergy is relatively low in Asian children (0.43% -0.64%) [24].
We considered a SPT wheal size of 8 mm and specific IgE level of 15 kUA/L as our cut off values for predicting peanut allergy [17,25]. Nevertheless, only half of the sensitized children had clinical reactions upon oral challenge. We performed an oral challenge despite the fact that our cut off values may diagnose peanut allergy because none of the sensitized subjects gave a history suggestive of peanut allergy. It seems that the cut off values should be tailored to the levels of consumption in different geographical locations. A recent study from the UK reported a 22.4% prevalence of clinical peanut allergy among sensitized subjects. The authors used the same wheal diameter and specific IgE cut off values that we used [25]. The published peanut specific IgE cut off values are variable ranging between 5 up to 57 kUA/L [26,27]. A wheal diameter of 16 mm was considered to have a positive predictive value of 100% for allergy to peanuts [27].
Although a prior probability from the history is an important starting point in the diagnosis of peanut allergy, none of our subjects gave a history suggestive of peanut allergy and all of them started consuming roasted peanuts before the age of 2 years. Roasted peanut is a popular snack in Egypt and its allergy is not a public concern due to underestimation of its significance even from the health care workers' point of view. The history was reported to be notoriously poor (approximately 30% verified) in identifying causal foods for chronic disorders such as atopic dermatitis [16]. The early peanut consumption by children does not seem to be a risk factor for peanut allergy and was postulated to be even protective [28].
There was no significant relation between the duration of breast feeding and peanut sensitization in the current study. A relevant study noted that neither the maternal peanut consumption during pregnancy and lactation nor the duration of breast feeding was associated with the development of peanut allergy [29]. On the other hand, a recent case-control study assumed that exposure to peanut allergens in utero or through breast milk may increase the risk of developing peanut allergy [30].
A family history of atopy, especially of food allergy, is a good screening test to identify an individual at risk of food allergy and the rate of allergy in a sibling of an allergic person is known to be higher than the rate in the general population. Although we could not demonstrate a statistically significant relation between the family history of allergy and peanut sensitization, all peanut sensitive children came from atopic families. Also, two of the peanut allergic children were brothers.
The current study did not demonstrate a significant gender difference in peanut sensitization but the boys outnumbered girls in the whole sample. Higher prevalence of peanut allergy among male children was previously reported [31].
The peanut sensitization in the present study did not vary according to the site of allergy. However, the 7 peanut sensitized children had physician diagnosed skin allergy combined with respiratory allergy in five. A relevant study on Asian children revealed that most (89.5%) of first reactions featured skin changes and that respiratory and GI symptoms did not occur as the sole manifestation [32]. One child with positive oral challenge in our series had allergy in one target-organ (skin) and a past history of one attack of unexplained anaphylaxis and two children had symptoms in more than one system (skin and respiratory). According to a voluntary registry in the Isle of Wight, about half of all children with peanut allergy had allergic manifestations in one target-organ system, 30% in two, 10% -15% in three, and 1% had symptoms in four systems [33].
Self-reported food allergy is an independent risk factor for potentially fatal childhood asthma [34]. Peanut allergy was reported in 28.8% of children with asthma in a recent investigation [35]. One of our peanut sensitized children was allergic to bananas. Her peanut oral challenge yielded a negative result. The girl was not latex sensitive and the parents did not consent to any further investigations. Evaluation of pollen cross reactivity would have been worthwhile [36].
From this pilot study, it seems that peanut allergy in Egypt is underestimated and that the sensitization rates are even higher. Skin prick and specific IgE testing aided by history are good screening tools to determine candidates for peanut oral food challenge. Peanut allergy can be associated with any clinical form of allergy and the causal relationship needs extensive evaluation. The conclusions are limited by the sample size and study design which targeted physician-diagnosed allergy rather than the general population. Further wider-scale populationbased studies as well as a national registry are needed to be able to outline the real magnitude of peanut allergy and its clinical correlates in our country.
AEC: absolute eosinophil count; BF: breast feeding; IQR: interquartile range; Peanut +: peanut sensitive; Peanut -: not sensitized to peanut. * Peanut sensitization: SPT wheal diameter ≥ 8 and specific IgE ≥ 15 kUA/L
The authors declare that they have no competing interests.
Authors' contributions EH put the study design and coordination, performed the statistical analysis, and drafted the manuscript. GG participated in collection of the study sample and data analysis. AS performed the total and specific IgE assay and differential blood cell counting. AE collected the study sample and performed the skin prick test. All authors read and approved the final manuscript.
Background: Elizabethkingia spp. are opportunistic pathogens often found associated with intravascular devicerelated bacteraemias and ventilator-associated pneumonia. Their ability to exist as biofilm structures has been alluded to but not extensively investigated.
Methods: The ability of Elizabethkingia meningoseptica isolate CH2B from freshwater tilapia (Oreochromis mossambicus) and E. meningoseptica strain NCTC 10016 T to adhere to abiotic surfaces was investigated using microtiter plate adherence assays following exposure to varying physico-chemical challenges. The role of cellsurface properties was investigated using hydrophobicity (bacterial adherence to hydrocarbons), autoaggregation and coaggregation assays. The role of extracellular components in adherence was determined using reversal or inhibition of coaggregation assays in conjunction with Listeria spp. isolates, while the role of cell-free supernatants, from diverse bacteria, in inducing enhanced adherence was investigated using microtitre plate assays. Biofilm architecture of isolate CH2B alone as well as in co-culture with Listeria monocytogenes was investigated using flow cells and microscopy.
Results: E. meningoseptica isolates CH2B and NCTC 10016 T demonstrated stronger biofilm formation in nutrientrich medium compared to nutrient-poor medium at both 21 and 37°C, respectively. Both isolates displayed a hydrophilic cell surface following the bacterial adherence to xylene assay. Varying autoaggregation and coaggregation indices were observed for the E. meningoseptica isolates. Coaggregation by isolate CH2B appeared to be strongest with foodborne pathogens like Enterococcus, Staphylococcus and Listeria spp. Partial inhibition of coaggregation was observed when isolate CH2B was treated with heat or protease exposure, suggesting the presence of heat-sensitive adhesins, although sugar treatment resulted in increased coaggregation and may be associated with a lactose-associated lectin or capsule-mediated attachment.
Conclusions: E. meningoseptica isolate CH2B and strain NCTC 10016 T displayed a strong biofilm-forming phenotype which may play a role in its potential pathogenicity in both clinical and aquaculture environments. The ability of E. meningoseptica isolates to adhere to abiotic surfaces and form biofilm structures may result from the hydrophilic cell surface and multiple adhesins located around the cell.
Members of the genus Elizabethkingia are aerobic, nonmotile, Gram-negative rods that display a light yellow pigmentation or may be non-pigmented [1]. The absence of gliding motility and the presence of flexirubin pigments differentiate these genera from other genera in their family Flavobacteriaceae. Only two species have been identified to date, i.e., Elizabethkingia meningoseptica and E. miricola [1].
E. meningoseptica is the most significant species for human clinical infections, although E. miricola has been associated with sepsis [2]. Elizabethkingia-related infections occur in severely immuno-compromised and postoperative patients as well as neonates [1]. E. meningoseptica has been implicated in endocarditis, cellulitis, abdominal infection, septic arthritis and eye infections in severely immuno-compromised patients [1] suffering from malignancy, end-stage hepatic and renal disease, extensive burns and acquired immune deficiency syndrome as well as community-acquired necrotizing fasciitis, pneumonia, and bacteraemia [3]. These infections constitute a major clinical concern, since together with Chryseobacterium spp., Elizabethkingia spp. isolates are constitutively resistant to multiple antibiotics [1,4].
Elizabethkingia spp, isolates constitute a further threat, being able to thrive in aqueous environments and are associated with intravascular device-related bacteraemias, wound sepsis, and ventilator-associated pneumonia by virtue of their ability to contaminate and persist in fluid-containing apparatus [2,3,5]. E. meningoseptica has been found in the hospital environment in such sites as water supplies, saline solution used for flashing procedures, disinfectants, and medical devices, including feeding tubes and arterial catheters [6]. Outbreaks have been documented following administration of contaminated medicine, use of devices contaminated via water or more sporadic infections in immuno-compromised patients or post-trauma and -surgery patients.
The bacterium has been isolated from such medical devices as the respirator, vaporizer and artificial ventilation tubing. E. meningoseptica strains isolated from slimy biofilm communities inside spouts of sink taps of a hospital have been implicated in a neonatal meningitis outbreak [7].
Elizabethkingia spp. have also been isolated from diverse ecological niches, including eutrophic lakes, soil, freshwater sources, spent nuclear fuel pools, and water condensation on the Russian space laboratory Mir [1]. E. meningoseptica have been recovered from diverse eukaryotes, including amoebae, frogs, turtles, birds, cats, dogs, and fish. The first E. meningoseptica infection in fish was diagnosed in farmed koi carp with skin lesions and hemorrhagic septicaemia. Fish-associated members of the genus Elizabethkingia may represent pathogenic or spoilage organisms or belong to the normal bacterial flora that colonize the mucus at the surface of the skin and gills and the intestine of healthy fish [1].
In the aquaculture environment, two challenges may be posed by E. meningoseptica, i.e., ability of these multidrug-resistant species to evade eradication following antimicrobial treatment and persistence in tanks due to biofilm community formation, leading to disease and associated economic losses; and their potential role as opportunistic human pathogens. The ability of these organisms to act as potential zoonotic pathogens, via transmission from fish and fish farm environments to immuno-compromised workers and consumers should not be underestimated [4] and necessitates investigation into their ability to adhere to surfaces and form biofilms.
Although Elizabethkingia spp. isolates have been isolated from clinical biofilm communities [7,8], the factors involved in initiating biofilm formation by these nonmotile bacteria has not been elucidated. The present study investigated the ability of Elizabethkingia meningoseptica isolate CH2B from farmed freshwater tilapia and strain NCTC 10016 T , to adhere to polystyrene under various physico-chemical conditions using the microtiter plate assay. Hydrophobicity as well as coaggregation and autoaggregation abilities were also investigated. The role of extracellular cell components in adherence was determined using reversal and inhibition of coaggregation assays, while the effect of cell-free supernatants from diverse bacteria in inducing enhanced adherence was investigated using microtiter plate adherence assays. Biofilm architecture of E. meningoseptica isolate CH2B was examined using a flow cell system, as was the ability of E. meningoseptica isolate CH2B to form a mixed biofilm structure with Listeria monocytogenes.
Creamish-yellow pigmented isolate CH2B was cultured from diseased freshwater tilapia (Oreochromis mossambicus) from a South African aquaculture facility. Isolate CH2B was presumptively identified as E. meningoseptica using the following tests: Gram stain, colony characteristics, and flexirubin pigment production [1]; and this identification was confirmed by 16S rRNA gene PCR and sequencing [9] (GenBank: EU598807).
E. meningoseptica isolate CH2B and type strain NCTC 10016 T were maintained on enriched Anacker and Ordal's agar (EAOA) [10] at ambient temperature (21°C ± 2°C). For long-term storage, cultures were placed in 20% glycerol and enriched Anacker and Ordal's broth (EAOB) and stored at -80°C.
E. meningoseptica isolates CH2B and NCTC 10016 T were cultured overnight in EAOB at room temperature (21°C ± 2°C) and centrifuged for 2 min at 12000 rpm. Cell pellets were washed and re-suspended in phosphate-buffered saline (PBS, pH 7.2) to a turbidity equivalent to a 0.5 M McFarland standard [11]. In order to determine bacterial microtitre plate adherence, wells of sterile 96-well U-bottomed polystyrene microtiter plates (Deltalabs S.L, Barcelona, Spain) were each filled with 90 μl EAOB/tryptic soy broth (TSB; Merck Chemicals, Gauteng, RSA) and inoculated with 10 μl of standardized cell suspensions, in triplicate [12]. Negative control wells containing only broth or PBS were included in each assay while a Vibrio mimicus isolate (VIB1; isolated from cultured trout) was used as a positive control. Plates were placed on a C1 platform shaker (New Brunswick Scientific, Edison, NJ, USA) and/or the benchtop to simulate dynamic and static conditions, respectively, and incubated aerobically at room temperature (21°C ± 2°C) and/or 37°C for 24 h, in either nutrient-poor EAOB/nutrient-rich TSB media. An optical density (OD) reading of each well was obtained at 595 nm using an automated microtiter-plate reader (Microplate Reader model 680, BioRad Laboratories Inc., Hercules, California). Tests were done in triplicate on three separate occasions and the results averaged [12].
Biofilm formation was classified as non-adherent, weakly-, moderately-or strongly-adherent. The cut-off OD (OD c ) for the microtiter plate test was defined as three standard deviations above the mean OD of the negative control. Isolates were classified as follows: ODOD C = non-adherent, OD C < OD(2 × OD C ) = weakly adherent; (2 × OD C ) < OD ≤ (4 × OD C ) = moderately adherent and (4 × OD C ) < OD = strongly adherent [12]. Statistical significance of differences (p < 0.05) due to altered variables (temperature, medium, agitation) in the microtiter adherence assays were determined using one way repeated measures analysis of variance (ANOVA; SigmaStat, V3.5, Systat Software, Inc., USA).
Surface hydrophobicity was assessed using the bacterial adherence to hydrocarbons (BATH) assay, with xylene (BDH, VWR International, Leicestershire, UK) as the hydrocarbon of choice [11]. E. meningoseptica isolates CH2B and NCTC 10016 T grown in EAOB at room temperature (21°C ± 2°C) were harvested during the exponential growth phase (18 h old cultures), washed three times and resuspended in sterile 0.1 M phosphate buffer (pH 7) to an OD of 0.8 at a wavelength of 550 nm (A 0 of 10 8 CFU/ml). Samples (3 ml) of the bacterial suspension were placed in glass tubes with 400 μl of xylene, equilibrated in a water bath at 25°C for 10 min and vortexed [11,13]. After a 15 min phase separation, the lower aqueous phase was removed and its OD 550 determined (A 1 ). Strains were considered strongly hydrophobic when values were >50%, moderately hydrophobic when values were in the range of 20-50%, and hydrophilic when values were <20% [14]. Each value represents the mean of experiments done in triplicate and on two separate occasions. PBS was used as a negative control and V. mimicus isolate VIB1 was used as a control [11].
For the modified salting aggregation test (SAT) assay, overnight EAOB cultures grown at room temperature (21°C ± 2°C) were harvested, washed twice and resuspended in PBS (pH 7.2). E. meningoseptica isolate CH2B was 'salted out' (aggregated) by combining 25 μl volumes, containing 2 × 10 9 bacteria, with 25 μl volumes of a series of methylene blue-containing ammonium sulphate [(NH 4 ) 2 SO 4 ] concentrations (0.2, 0.5, 1, 1.5, 2, 2.5, 3, and 4 M) on microscope slides [11]. The lowest final concentration of (NH 4 ) 2 SO 4 causing aggregation was recorded as the SAT value and classified as follows: < 0.1 M = highly hydrophobic, 0.1 M -1.0 M = hydrophobic and >1.0 M = hydrophilic [15]. Experiments were done in triplicate on two separate occasions and respective (NH 4 ) 2 SO 4 concentrations were used as negative controls.
For the autoaggregation assay, E. meningoseptica isolates CH2B and NCTC 10016 T were grown in 20 ml EAOB at room temperature (21°C ± 2°C), harvested after 36 h, washed and re-suspended in sterile distilled H 2 O to an OD of 0.3 at a wavelength of 660 nm. The percentage of autoaggregation was measured by transferring a 1 ml sample of bacterial suspension to a sterile plastic 2 ml cuvette and measuring the OD after 60 min using a DU 640 spectrophotometer (Beckman Coulter) at a wavelength of 660 nm [16]. The degree of autoaggregation was determined as the percent decrease of optical density after 60 min using the equation:
OD 0 refers to the initial OD of the organism measured. Sixty min after OD 0 was obtained, the cell suspension was centrifuged at 2000 rpm for 2 min. The OD of the supernatant was measured (OD 60 ). Experiments were carried out in triplicate on two separate occasions [16].
E. meningoseptica isolates CH2B and NCTC 10016 T were examined for their ability to coaggregate with each other as well as with the following bacterial partner strains: Aeromonas hydrophila, A. sobria, A. salmonicida, A. media, Acinetobacter spp., Enterococcus faecalis ATCC 29212, Escherichia coli ATCC 25922, Flavobacterium johnsoniae-like spp. isolates YO12, YO19, YO51, YO60, YO64, Listeria monocytogenes NCTC 4885, L. innocua LMG 13568, Micrococcus luteus, Pseudomonas aeruginosa, Salmonella enterica serovar Arizonae and Staphylococcus aureus ATCC 25923 [11].
For coaggregation assays, bacteria were grown in 20 ml EAOB or TSB, harvested after 36 h, washed and resuspended in sterile distilled H 2 O to an OD of 0.3 at a wavelength of 660 nm. The degree of coaggregation was determined by OD readings of paired isolate suspensions (500 μl of each isolate). The cell mixture was centrifuged at 2000 rpm for 2 min and the OD of the supernatant (600 μl) was measured at a wavelength of 660 nm [16]. The quantitative coaggregation rate of paired isolates was calculated using the equation:
where OD Tot value refers to the initial OD, taken immediately after the relevant strains were paired; and OD S refers to the OD of the supernatant, after the mixture was centrifuged after 60 min [16]. Experiments were carried out in triplicate on two separate occasions. Differences in coaggregation between E. meningoseptica CH2B and E. meningoseptica NCTC 10016 T were determined using one way repeated measures analysis of variance (ANOVA; SigmaStat V3.5). Differences were considered significant if p < 0.05.
The effect of simple sugars, heat and protease treatment on isolate CH2B's ability to coaggregate with L. innocua LMG 13568 and L. monocytogenes NCTC 4885 was investigated.
The ability of simple sugars to reverse E. meningoseptica isolate CH2B coaggregation with Listeria spp. involved filter-sterilized solutions of lactose and galactose, respectively, being added to one of the coaggregating partners at final concentrations of 50 mM [17]. Mixtures were vortexed and tested for coaggregation using the coaggregation assay described above.
The ability of heat treatment to inhibit E. meningoseptica isolate CH2B coaggregation with Listeria spp. was conducted using the method of Kolenbrander et al. [18]. Cells were harvested from O/N EAOB/TSB cultures, washed three times and resuspended in de-ionized water. One of the coaggregating partners was then heated at 80°C for 30 min in a waterbath. Following heat treatment, the OD of each bacterial suspension was adjusted to 0.3 at a wavelength of 660 nm. Heat-treated and untreated cells were combined in reciprocal pairs and their capacity to coaggregate was assessed.
Protease sensitivity of the polymers mediating coaggregation of isolate CH2B with Listeria spp. isolates was tested using a method described by Rickard et al. [19]. Cells were harvested from O/N EAOB/TSB cultures and resuspended in de-ionized water to an OD of 0.3 at a wavelength of 660 nm. Proteinase K was added to the standardized cell suspensions to a final concentration of 2 mg/ml. Incubation at 37°C for 2 h was followed by centrifugation and washing of the pelleted cells three times in de-ionized water. Cells were resuspended and the OD adjusted to 0.3 at 660 nm. Protease-treated and untreated cells were combined and their capacity to coaggregate determined.
Differences in coaggregation between untreated E. meningoseptica CH2B and treated bacteria (E. meningoseptica CH2B, L. innocua, and L. monocytogenes) were determined by paired t-tests (SigmaStat V3.5). Differences were considered significant if p < 0.05.
The standard microtiter plate adherence test [12] was modified to determine the ability of extracellular secretions of various aquaculture, food and/or human pathogens (Aeromonas hydrophila, A. salmonicida, A. sobria, Chryseobacterium spp. isolates CH8, CH15, CH23, CH25 and CH34, E. meningoseptica CH2B, E. coli, Edwardsiella tarda, F. johnsoniae-like isolate YO59, L. innocua, L. monocytogenes, Myroides odoratus MY1, P. aeruginosa, S. enterica serovar Arizonae, and V. mimicus VIB1) to induce enhanced adherence of E. meningoseptica CH2B.
Three-day old cultures of each of the above organisms were centrifuged at 2000 rpm for 10 min and supernatants were filter-sterilised using 0.2 μm filters, in order to obtain cell-free spent medium. E. meningoseptica CH2B cell pellets were washed and re-suspended in phosphate-buffered saline (PBS, pH 7.2) to a turbidity equivalent to a 0.5 M McFarland standard [11]. Ten μl of the standardized suspension was added to microtiter wells containing 100 μl TSB and 90 μl of the filtered supernatant. Controls included standardised isolate CH2B cell suspension added to TSB and respective filtered supernatants in TSB without isolate CH2B, in order to determine a change in adherence abilities and ensure that the change in adherence was due to induction, respectively. Microtiter plates were incubated at room temperature (21°C ± 2°C) for 48 h. An optical density (OD) reading of each well was obtained at 595 nm using an automated microtiterplate reader (Microplate Reader model 680, BioRad Laboratories Inc., Hercules, California). Tests were done in triplicate on three separate occasions and the results averaged [12].
Biofilm formation by E. menigoseptica isolate CH2B was investigated using continuous culture once-through eight channel flow-cell system, while the mixed-species biofilm flow cell study involved L. monocytogenes strain NCTC 4885 together with isolate CH2B. The eightchannel perspex flow cell (channel size 30 × 4.5 × 3 mm), the glass cover-slip covering (no. 1 thickness, 75 mm by 50 mm), and attached silicone tubing (1 × 1.6 mm × 3 mm × 5 m tubing; The Silicone Tube, RSA) was assembled as described previously [11]. Silicone tubing was connected to a reservoir containing 2 l of EAOB/TSB and the flow cell was filled with EAOB/TSB, with a flow rate of 0.25 ml/min being maintained using a multi-channel peristaltic pump (Model 205S, Watson-Marlow, UK) located upstream of the flow cell. A one ml volume of EAOB overnight cultures of isolate CH2B was inoculated into each channel, below the clamps sealing silicone tubes upstream of each channel, using sterile syringes. One ml mixed pure culture inoculations, consisted of 0.5 ml combinations of E. meningoseptica isolate CH2B and L. monocytogenes NCTC 4885. Stagnant conditions were maintained for the first hour to allow attachment, prior to inoculated channels being exposed to flowing EAOB/TSB at a constant flow rate of 0.25 ml/min. Flow cell systems were kept at room temperature (21°C ± 2) throughout the experiments. Each flow cell channel was investigated by transmitted light using a Nikon Eclipse E400 (Nikon, Japan) microscope at 600-fold magnification and after 24 h and 48 h, respectively, to visualize bacterial attachment to a glass surface and biofilm development. Images were documented with a CHU high-performance charge-coupled camera device (model 4912-5010/000).
Biofilm-forming ability of E. meningoseptica E. meningoseptica isolates CH2B and NCTC 10016 T , as well as V. mimicus VIB1 were screened for their adherence to polystyrene microtitre plate wells following 24 h incubation at room temperature (21°C ± 2°C) or 37°C, under static or dynamic conditions in nutrient-rich (TSB) or nutrient-poor (EAOB) media (Table 1).
Isolate CH2B displayed moderate adherence in EAOB at both room temperature and 37°C, respectively, and became strongly adherent when exposed to TSB (Table 1). E. meningoseptica NCTC 10016 T was moderately adherently at room temperature and strongly adherent at 37°C in TSB, but weakly adherent in EAOB. In contrast, V. mimicus displayed strongest adherence in EAOB at room temperature. An increase in temperature to 37°C or alteration of the medium to TSB resulted in weak to moderate adherence for V. mimicus (Table 1). Given the small sample number, none of the physicochemical parameter combinations resulted in statistically significant adherence.
Both E. meningoseptica isolate CH2B and E. meningoseptica NCTC 10016 T appeared to be strongly hydrophilic with BATH indices of 0.77% and 0.36%, respectively. Isolate CH2B was 'salted out' with a 4 M (NH 4 ) 2 SO 4 concentration, confirming its hydrophilicity. a Biofilm formation assay data is the mean of three independent experiments carried out in triplicate ± standard deviation following growth in minimal (EAOB) or rich (TSB) media at 21 or 37°C under dynamic or static conditions, respectively. b Biofilm formation (BF) was classified as non-adherent (N), weakly (W)-, moderately (M)-or strongly (S)-adherent using previously described criteria [12].
Isolate CH2B displayed an autoaggregation index of 37.4% (
Table 2), while that of the type strain E. meningoseptica NCTC 10016 T was 33.1%. Coaggregation occurred to varying degrees between all of the 18 partner strains and E. meningoseptica isolates CH2B or NCTC 10016 T (Table 2), respectively. Isolate CH2B displayed coaggregation indices ranging from 2.5% with E. meningoseptica NCTC 10016 T to 82.2% with S. aureus ATCC 25923. Isolate CH2B had coaggregation indices >40% with 31.8% of the partner strains (Table 2). E. meningoseptica NCTC 10016 T displayed coaggregation indices ranging from 2.5% with E. meningoseptica CH2B to 75.1% with a Micrococcus spp. isolate. Strain NCTC 10016 T had coaggregation indices >40% with 42.1% of the partner strains (Table 2). Although, differences were observed in the coaggregation indices profiles of E. meningoseptica CH2B and E. meningoseptica NCTC 10016 T , these were not statistically significant.
Since coaggregation indices of 70.4% and 77.4% were obtained between isolate CH2B and L. monocytogenes NCTC 4885 and L. innocua LMG 13568 (Table 2), respectively, they were selected for the reversal and inhibition of coaggregation assays following sugars, heat or proteinase K treatments. Sugar reversal experiments with lactose or galactose of either partner increased both the autoaggregation and coaggregation indices (Table 3). Lactose treatment resulted in greater coaggregation of isolate CH2B with both treated and untreated L. innocua LMG 13568 compared with L. monocytogenes NCTC 4885, while this was reversed following galactose treatment, with greater coaggregation being observed with treated and untreated L. monocytogenes NCTC 4885.
Heat treatment of isolate CH2B resulted in a decrease in autoaggregation (Table 3) and coaggregation, respectively. A greater decrease in coaggregation was observed with untreated L. monocytogenes NCTC 4885 than with untreated L. innocua LMG 13568. However, increased coaggregation was observed when the Listeria spp. partner strains were treated with heat (Table 3).
A similar trend was observed with proteinase K treatment of isolate CH2B, i.e., decreased autoaggregation of CH2B as well as coaggregation with the untreated partner strains. Proteinase K treatment of Listeria spp. isolates resulted in increased coaggregation between L. monocytogenes NCTC 4885 and isolate CH2B (Table 3). A greater reduction in the coaggregation indices were observed when heat-or protease-treated isolate CH2B cells were partnered with L. monocytogenes NCTC 4885 than with L. innocua LMG 13568.
Following exposure to cell-free supernatants from the three Aeromonas spp. isolates, Chryseobacterium spp. isolates CH8 and CH25 and V. mimicus, isolate CH2B's adherence decreased 0.48 -1-fold (Figure 1). Increased adherence, ranging from 1.3 -3.58-fold was observed with the remaining cell-free supernatants (Figure 1). Cell-free supernatants from Chryseobacterium spp. isolates CH15 and CH34, P. aeruginosa, L. innocua, and L. monocytogenes increased adhesion 2 -4-fold. A 1.5fold increase in adherence was observed following exposure of isolate CH2B to its own cell-free supernatant (Figure 1).
Adherence of isolate CH2B to glass coverslips was investigated by light microscopy, starting from the surface of the glass slide and scanning several planes interspersed by short distances in order to visualize biofilm architecture and microbial behavior throughout the depth of the individual flow chambers. By 24 h in nutrient-poor (EAOB) medium, isolate CH2B displayed initial widespread attachment to the glass coverslips and microcolonies were observed. After 48 h, majority of the cells were attached at a polar end (Figure 2), and conestructures were observed with chains of cells reaching into the flowing medium. In nutrient-rich (TSB) medium flow cells, cells were attached along their length in microcolonies interspersed with polarly-attached cells (Figure 3a). Microcolonies merged by 48 h to form a thick, complex biofilm structure across entire channel surface (Figure 3b). Distinction between bacterial strains in mixed-culture experiments was made visually by comparing images to that of pure culture, single-species flow cell experiments. Cells differed morphologically with E. meningoseptica CH2B cells being longer, thinner cells, and Listeria spp. shorter and thicker. When co-inoculated in nutrientpoor medium, both isolate CH2B and L. monocytogenes NCTC 4885 cells displayed delayed attachment to the glass surfaces, and attached cells were only observed 48 h following inoculation. Although both CH2B and L. monocytogenes NCTC 4885 cells were able to attach to the glass slides, distinct colonies were formed with no association between the different species (Figure 4). In nutrient-rich medium, cells of both species appeared to be scattered over the surface after 24 h, but by 48 h only a monolayer of isolate CH2B was observed covering the surface of the glass slide.
E. meningoseptica has been identified in infection outbreak associated with municipal water reservoirs, potable water [7] and colonization of tap water in a neonatal intensive care unit [7]. Infections associated with E. meningoseptica have been associated with instrumentation contamination or the internal placement of indwelling medical devices [5,20]. Although its role in infection appears to be linked to biofilm formation and a worse outcome in patients [5], no studies have focused on investigating the factors involved in the adherence of E. meningoseptica to abiotic or biotic surfaces. The presence of E. meningoseptica in various hospital environments involves optimal growth conditions including moist, cool environments or standing water at approximately 21°C [8]. Typically a shift to oligotrophic conditions triggers adhesion and biofilm formation [21], however, the converse was observed for E. meningoseptica CH2B. Unlike V. mimicus isolate VIB1, biofilm formation for E. meningoseptica was optimal in nutrientrich TSB at both 21°C and 37°C, respectively. Lin et al. [5] also observed strong E. meningoseptica isolate-specific biofilm formation in the relatively nutrient-rich Luria-Bertani medium. A similar trend was observed for Hafnia alvei, where higher nutrient concentrations favoured biofilm formation [22]. Myroides odoratus, a related organism, by contrast, displayed strong adherence in both nutrient-rich and poor media at 21°C but was moderately adherent at 37°C in nutrient-rich medium [23]. Biofilm formation by avian faecal commensal E. coli strains was induced by both nutrient-rich and nutrient-poor media [24]. Even under nutrient-poor conditions at both 21 and 37°C, E. meningoseptica CH2B did not lose its ability to adhere, but displayed moderate biofilm-formation. Nutrient-poor conditions at lower temperatures and nutrient-rich medium at 37°C are conditions typically associated with environmental and clinical conditions, respectively. E. meningoseptica adherence occurred preferentially in nutrient-rich medium at both 21°C and 37°C, suggesting that nutrient limitation is not a cue in the switch to a sessile lifestyle for E. meningoseptica. Altering the hydrodynamic conditions appeared to affect the degree of biofilm formation more significantly in nutrient-rich medium and requires further investigation.
Whole cell hydrophobicity, autoaggregation, and coaggregation are important for colonisation and biofilm development in flowing environments [25]. Bacteria behave as hydrophobic particles due to their net negative surface charge and this surface hydrophobicity is usually associated with bacterial adhesiveness, varying from organism to organism, from strain to strain and is influenced by the growth medium, bacterial age and bacterial surface structures [26,27]. Although the general rule has been that adhesiveness increases and decreases with increasing and decreasing hydrophobicity, respectively [28], a number of studies have shown contradictory results where no relationship was found between the bacterial strain's surface hydrophobicity and the extent of initial binding to either a hydrophilic or hydrophobic substrate [14,29]. Flavobacterium johnsoniae-like and F. psychrophilum isolates from fish were hydrophilic by the BATH assay [11,27], as were adhesion-defective mutants of a F. johnsoniae strain displaying poor adherence [26]. Although E. meningoseptica CH2B appeared to be very hydrophilic by both the BATH and SAT assays, it displayed strong adherence.
The hydrophilic nature of the E. meningoseptica isolates might account for cells adhering preferentially along the entire surface of the glass slide rather than to the perspex surfaces in flow cells. Majority of the cells attached by their polar sides to glass in nutrient-poor medium, which could be an attempt to increase surface area for nutrient uptake in nutrient-limited environments, since horizontal attachment was observed in nutrient-rich medium.
According to Ofek and Doyle [20], capsule presence obscures cell hydrophobicity. Coagulase-negative Staphylococcus strains with capsules were more hydrophilic than non-encapsulated strains [30,31]. E. meningoseptica CH2B's hydrophilicity might be explained in part by the presence of a capsule layer (unpublished data). The capsule presence might also account for the autoaggregation index of 37%. Autoaggregation is a 'selfish' mechanism whereby a strain within the biofilm will express polymers to enhance the integration of genetically identical strains into biofilms [32], especially in high shear environments [25]. The high autoaggregation index could thus explain the aggregation of E. meningoseptica CH2B cells in the high shear inflow point of the flow-cell chambers.
Bacterial coaggregation is defined as cell-to-cell adherence of different bacterial species or strains [33]. Coaggregation plays an important role in the development of multi-species biofilms into integrated biological structures, by mediating the juxtaposition of species next to favourable partner species within taxonomically diverse biofilms [34]. The coaggregation profiles of isolate CH2B and strain NCTC 10016 T were not identical and this might be accounted for in part by the environmental and clinical isolation sources of the respective bacteria and diverse selection pressures potentially experienced in their diverse ecological niches. The strongest coaggregation partners with Elizabethkingia spp. isolate CH2B were not other Gram-negative bacteria commonly found in the aquatic environment, i.e., Aeromonas or Flavobacterium spp., but rather organisms important in food spoilage and/or intoxications, i. e., S. aureus, L. innocua, L. monocytogenes, S. enterica, E. faecalis, and P. aeruginosa. A similar trend was observed for F. johnsoniae-like isolates [11]. M. luteus, B. natatoria, Fusobacterium and Prevotella spp. have been identified as bridging organisms in biofilms due to their ability to coaggregate with diverse coaggregating partners [18,35,36]. In the present study, study isolate CH2B displayed high coaggregation indices with 12 of the 19 partner strains, and it is, therefore, not unlikely that it is a possible bridging organism in aquaculture environments.
Although both CH2B and L. monocytogenes NCTC 4885 attached to the glass slide in the mixed-species flow cell experiment, the high coaggregation index displayed by these two bacterial species was not apparent. The microcolonies of the two species appeared to be distinctly separated from one another on the glass surface. Based on induction experiment data, extracellular molecules in Listeria spp. growth medium supernatants, as well as that of Chryseobacterium sp. isolates CH15 and CH34, and P. aeruginosa increased the adherence of E. meningoseptica isolate CH2B to microtiter plate surfaces more than 2-fold. Quorum sensing signaling molecules within the cell-free supernatants could account for the increased adherence to polystyrene microtiter plates. This might also explain the high coaggregation indices observed with Listeria sp. isolates and the increased, albeit separate, adherence observed for both L. monocytogenes and E. meningoseptica CH2B in the mixed-species biofilm flow cell experiments. The 37.4% autoaggregation index, and the increased adherence observed for isolate CH2B following exposure to its own cell-free supernatant, suggests a potential role for quorum sensing in autoaggregation and biofilm formation.
Cell surface components or properties (flagella, pili, adhesin proteins, capsules, and surface charge) influence attachment and coaggregation. Flagella facilitate bacterial motility to specific attachment sites, while changes in cellular physiology affects surface membrane chemistry, surface proteins such as pili and adhesins, synthesis of polysaccharides, and cell aggregation, all of which influence adhesion [37]. Adhesive bacteria have developed various strategies to scaffold or present their adhesins. These include surface appendages and structures that bear adhesins, i.e., flagella, fimbriae, capsules, outer membranes, loosely attached peripheral components, etc. Adhesins may be proteins, polysaccharides, lipids, or teichoic acids [30]. The coaggregation interaction is a highly specific process mediated by the recognition of either complementary lectin (sugar-binding proteins)carbohydrate molecules between the aggregating partners [33]; polysaccharides of capsule or LPS bind to lectins on host-cell surface; protein-protein; hydrophobic moieties of proteins on one cell binding with lipids on another cell; and/or lipid-lipid interactions. Receptors may contain carbohydrate or amino acid residues [30].
In order to investigate the type of adhesin structures present on the E. meningoseptica CH2B surface, inhibition of coaggregation assays were undertaken. Autoaggregation of CH2B cells was inhibited when untreated cells were paired with heat-or protease-treated cells. A similar trend was observed for Acinetobacter calcoaceticus [35]. Proteinase K treatment inhibited biofilm formation by non-typeable Haemophilus influenzae, as well as rapidly detached preformed biofilms [38]. Since both heat and protease treatments of E. meningoseptica CH2B resulted in decreased autoaggregation and coaggregation, heat-and protease-sensitive adhesins (lectins) appear to be localized on the E. meningoseptica cell surface. Attachment of heat-and protease-treated E. meningoseptica cells to untreated L. innocua and L. monocytogenes appears to involve different combinations of receptors, since variations were observed in the decreased coaggregation indices (Table 3). Heat-and protease-treatment of Listeria spp. cells resulted in increased coaggregation indices with untreated CH2B cells, indicating the presence of heat-and proteasestable listerial receptor molecules.
Sugar treatment did not produce a partial or complete inhibition of E. meningoseptica autoaggregation and coaggregation as observed for freshwater/aquatic bacteria [17,35,36] and sewage sludge bacteria [16]. Since the lectin-saccharide interactions are usually very specific, a wider variety of sugars might have to be assayed to yield a reversal of the coaggregation reactions. However, protein-carbohydrate interactions were not reversed by sugars [16]. The increased autoaggregation and coaggregation indices with both lactose and galactose were unexpected. This occurred when either isolate CH2B or the listerial cultures were treated with sugars. It might be speculated that the treatment sugars added to the capsular material enclosing isolate CH2B and intensified the adhesive effect and thus coaggregation. While capsule presence may mask potential adhesins such as fimbriae, it may stabilize the adhesion-receptor interaction. The capsule chemical composition, while primarily polysaccharide may also include protein adhesion molecules. Thus the capsule components may also be receptors for lectins on another bacterium [30]. Thus, in addition to conferring a hydrophilic nature to the cell, the capsule in E. meningoseptica CH2B might play an integral role in the strong adherence ability of this organism.
Factors affecting coaggregation include: adhesin and receptor density and distribution; hydrophobic character of receptor, adhesin or receptor nearest neighbours; medium composition and pH; and chelating agents [30]. Coaggregation among aquatic bacteria is mediated by lectin-saccharide interactions, and these aquatic strains often carry multiple adhesins or receptors or a combination of both, which is also a common feature of coaggregating oral bacteria [36]. Multiple adhesins may also be distributed around the E. meningoseptica cell surface allowing interactions with diverse microorganisms and colonization of diverse substrata. This would allow E. meningoseptica to compete successfully in a microorganism-rich environment.
The present study has shown that an E. meningoseptica isolate CH2B from tilapia possesses the ability to adhere to abiotic surfaces and form biofilms under various environmental conditions. Hydrodynamic flow in clinical or environmental niches may be more rapid than the rate of multiplication and unattached organisms will be eliminated, thus adhesion confers the important ability to colonise substrata [30]. E. meningoseptica CH2B was able to coaggregate with bacterial species important from a food and health perspective. Although E. meningoseptica are mostly described as opportunistic pathogens in both veterinary and human infections, the cause for concern arises from their association with pathogens and spoilage organisms causing great economic losses in the aquaculture and food industries and lethal device-or equipment-associated infections in immuno-compromised humans.
The ability of Elizabethkingia spp. to adhere to biotic and abiotic surfaces and the association with disease requires further study. Quantitative characterization and chemical analysis of the capsular material might provide valuable information regarding the capsule's role in the adherence abilities of E. meningoseptica. Furthermore, an investigation of specific cell-surface molecules mediating strong coaggregation abilities between E. meningoseptica and coaggregating partners may provide valuable information for anti-adhesion therapy which could be applied in aquaculture systems for the eradication of biofilms harbouring pathogenic organisms. The enhanced adherence of E. meningoseptica CH2B induced by cell-free supernatants points to the presence of a quorum sensing system, whose activity might be associated with autoaggregation, biofilm formation and/or the ability to colonise surfaces and initiate infection.
a For partner strains, autoaggregation indices were determined using assay described by Malik et al.[16]
and are indicated within (). b Coaggregation indices represent the means of two independent replicate experiments as described by Malik et al.[16]
and Basson et al.[11]
. c NT refers to not tested.
a Coaggregation indices represent the means of two independent replicate experiments as described by Malik et al.[16]
and Basson et al.[11]
. Figure 1 Microtiter plate adherence of E. meningoseptica isolate CH2B, following exposure to cell-free spent medium supernatants from selected Gram-negative and Gram-positive bacteria, at room temperature (21°C ± 2°C) under static conditions in nutrient-rich (TSB) medium. Bars represent means ± standard deviations for three independent replicate experiments.
This work was funded b a
The authors declare that they have no competing interests.
Authors' contributions AJ participated in designing the experiments, executing them, and performing data analysis. HYC conceived the study, participated in its design, data analysis and coordination and drafted the manuscript. All authors read and approved the final manuscript.
The species sensitivity distribution (SSD) concept is an important probabilistic tool for environmental risk assessment (ERA) and accounts for differences in species sensitivity to different chemicals. The SSD model assumes that the sensitivity of the species included is randomly distributed. If this assumption is violated, indicator values, such as the 50% hazardous concentration, can potentially change dramatically. Fundamental research, however, has discovered and described specific mechanisms and factors influencing toxicity and sensitivity for several model species and chemical combinations. Further knowledge on how these mechanisms and factors relate to toxicologic standard end points would be beneficial for ERA. For instance, little is known about how the processes of toxicity relate to the dynamics of standard toxicity end points and how these may vary across species. In this article, we discuss the relevance of immobilization and mortality as end points for effects of the organophosphate insecticide chlorpyrifos on 14 freshwater arthropods in the context of ERA. For this, we compared the differences in response dynamics during 96 h of exposure with the two end points across species using dose response models and SSDs. The investigated freshwater arthropods vary less in their immobility than in their mortality response. However, differences in observed immobility and mortality were surprisingly large for some species even after 96 h of exposure. As expected immobility was consistently the more sensitive end point and less variable across the tested species and may therefore be considered as the relevant end point for population of SSDs and ERA, although an immobile animal may still potentially recover. This is even more relevant because an immobile animal is unlikely to survive for long periods under field conditions. This and other such considerations relevant to the decision-making process for a particular end point are discussed.
Decades of ecotoxicologic testing have repeatedly showed large differences in the response of species toward toxicants, but they have not resulted in the identification of a ''most sensitive species'' (Cairns 1986), which is now widely accepted as nonexistent, although some indications for generally more sensitive groups exist (Dwyer et al. 2005). Across chemicals, it is primarily the mode and mechanism of action of a toxicant that determines an organism's response to exposure (Thurston et al. 1985;Escher and Hermens 2002;Jager et al. 2007), but even for a single toxic compound, large differences in species sensitivity have been found (Rubach et al. 2010). Differences in sensitivity across species are a source of uncertainty for the process of ERA. In the lowest tier of ERA, this uncertainty is often accounted for by using safety factors to derive threshold values for acceptable environmental concentrations (Van Leeuwen and Vermeire 2007). Despite their importance to the improvement of risk assessment, surprisingly little is known about the underlying mechanisms driving differences in sensitivity. Hence, most higher tier interspecies extrapolations are performed using probabilistic approaches such as the species sensitivity distribution (SSD) concept (Posthuma et al. 2002) or by performing multispecies tests. It is well known that differences in uptake and elimination of a compound into the body or organs cause differences in sensitivity, and these can be decreased when risk assessment is based on internal concentrations (McCarty and Mackay 1993). In addition, numerous studies have indicated that differences in sensitivity can also be explained by physiologic factors, such as differences in target enzyme constitution, detoxification or compensation abilities, e.g., (Heckmann et al. 2008). In this context, the comparison of toxic effects measured with different end points in bioassays can hold useful information for ERA when interpreted with regard to the processes of toxicity, such as toxicokinetics and toxicodynamics. For instance, time dependency of toxicity and differences in the toxicity response for different end points indicate that major differences in the processes of toxicity exist across species (Verhaar et al. 1999). Although suitable methods, such as time-to-event analysis (Newman and McCloskey 1996), time-independent sensitivity values (Mayer et al. 2002), and the dynamic energy budget theory (Kooijman and Bedaux 1996) have been developed, dynamics of effects have been largely ignored in ERA. Often lethal and effective concentration values for different time points are treated equally without distinction, both for practical reasons and for lack of more specific data. The consequences of this ignorance are difficult to estimate at this point but may lead to arbitrary conclusions. For instance, for an organophosphorous compound, such as chlorpyrifos, which affects the nervous system but does not lead to immediate mortality, large differences in effect end points could exist between species, which is also relevant to risk assessment of time-variable exposure scenarios. Despite widespread activities to establish sublethal end points in risk assessment not only for chronic but also for acute toxicity, the literature lacks publications discussing the relevance of particular end points for the ERA of pesticides (Baas et al. 2010). The only exception are endocrine disruptors, for which the debate for the most relevant end point has developed further (Rhind 2009).
This study aimed to evaluate how well two end points, mortality and immobility, reflect the toxicity of the organophosphorous insecticide chlorpyrifos in a variety of freshwater arthropods in terms of their effect dynamics and their variability across species. Experimental data were collected by means of 96-h toxicity tests with a range of exposure concentrations. These data were used to calculate both 50%-effective and lethal concentrations (L(E)C 50 s) with which SSDs per investigated time point and end point were populated. The influence of end point choice on risk indicators, such as L(E)C 50 s and the 5% hazardous concentration (HC 5 ) is discussed. Furthermore, we discuss the context in which toxicity information on different end points can improve understanding of differences in toxic processes among investigated species and how this can lead to more mechanism-based risk assessment.
For the 96-h toxicity experiments, chlorpyrifos (O,O-diethyl O-3,5,6-trichloro-2-pyridyl phosphorothioate, 99% purity, CAS 2921-88-20, lot no. 51205) purchased from Dr. Ehrenstorfer GmbH (Augsburg, Germany) was used. To avoid the use of a solvent carrier, aqueous solutions of chlorpyrifos were prepared using the principle of generator columns (Devoe et al. 1981) as described in the supporting information of (Rubach et al. in press). The effluent of the generator column delivered a stable concentration of parental chlorpyrifos of approximately 1250 lg/l. Before use, each such obtained stock solution was measured by liquid-liquid extraction of subsamples with n-hexane, followed by gas chromatography (GC) and electron capture detection (ECD) to determine its exact concentration of chlorpyrifos, and dosing schemes were subsequently adapted.
Gas Chromatography GC of aqueous stock solutions and test concentrations for the 96-h short-term toxicity testing was performed with an Agilent 6890 N GC equipped with a micro EC detector and 7683 Series B Injector (Agilent Technologies, Inc., Santa Clara, CA, USA). The injection volume was 3 ll at a temperature of 250°C with a split ratio of 10. As stationary phase, a DB5 medium-bore column of 30-m length and with a 0.32-mm internal diameter was used, whereas the mobile phase was created with a constant flow of 3.1 ml helium/min. The oven temperature was set to 200°C in isotherm mode. The temperature of the ECD was 300°C, and N 2 was used as make-up gas with a flow of 30 ml/min. The retention time of chlorpyrifos using this method was 3.8 min. For each test a specific limit of detection (LOD) was calculated (Table 1) due to differences in the sample extraction on the basis of the detection limit of the apparatus (0.1 lg/l) and the highest respective concentration factor used for the lower concentration levels per species.
Test Species and Test Media The 14 freshwater arthropod species used in the present study and their origin, life stage, and sampling date are listed in Table 1. Species identity was determined by trained staff according to established protocols using 6 to 10 individuals subsampled from the catch. Most species were collected at the experimental field station of Alterra, ''The Sinderhoeve'' (Renkum, The Netherlands), where they were sampled from untreated cosms, ditches, and storage systems, but Molanna angustata originated from
Table 1 Test species and methodologic details of 96-h toxicity tests (including LODs) Order Species Life stage Origin a Sampling date Test system b Intended concentration (lg/l) Measured concentration (lg/l) Replication (n) No. of animals/ replicate LOD (lg/l) Anisoptera A. imperator Larva c A 01.03.2007 ***** c 0.9; 1.9; 4; 8.4; 17.6 0.8; 1.6; 3.5; 7.3; 14.3 2 5 0.05 Isopoda A. aquaticus Adult B 06.12.2006 **** g 0.2; 0.5; 1.2; 2.7; 6.2 0.2; 0.6; 1.3; 2.9; 6.1 3 10 0.01 Diptera C. obscuripes Larva A, B 02.02.2007 **** 0.2; 0.5; 1; 2; 4.2 0.3; 0.9; 1.7; 3.3; 3.9 3 10 0.05 Diptera C. dipterum Larva A, B 23.04.2007 **** g 0.03; 0.07; 0.17; 0.41; 1 0.04; 0.06; 0.12; 0.3; 0.8 3 10 0.01 Cladocera D. magna Neonate d C 27.11.2006 * 0.01; 0.03; 0.1; 0.3; 1; 3 0.01; 0.03; 0.1; 0.3; 0.9; 2.5 4 5 0.01 Amphipoda G. pulex Adult A, B 24.11.2006 ***** c,g 0.02; 0.05; 0.11; 0.25; 0.6 0.02; 0.04; 0.1; 0.24; 0.4 2 9 0.01 Trichoptera M. angustata Larva D 09.08.2007 ** g 0.1; 0.4; 1.6; 6.4; 25.6 0.3; 0.5; 1.9; 7.3; 29.1 3 5 0.01 Decapoda N. denticulata Adult/juvenile C 12.09.2008 ***g 5; 13.3; 35.1; 93.1; 246; 653 5; 11.7; 34.2; 88.1; 228; 735 4 5 0.04 Heteroptera N. maculata Adult B 11.09.2008 *** c 1; 2.7; 7.29; 19.7; 53.1 0.8; 2.1; 6.8; 15.6; 50.2 4 5 0.04 Lepidoptera P. stratiotata Larva B 17.06.2007 ** g 0.2; 0.9; 4.1; 18; 82 0.4; 2.6; 5; 16; 23 2 5 0.05 Heteroptera P. minutissima Adult/nymph A, B 01.03.2007 **** 0.9; 1.9; 4; 8.4; 17.6 0.8; 1.8; 3.8; 8.9; 13.5 3 5 0.05 Decapoda Procambarus spec. Juvenile C 20.07.2007 **** g 0.2; 0.9; 4.1; 18.2; 82 0.4; 0.9; 3.2; 14.1; 107 3 5 0.05 Decapoda Procambarus spec. Adult C 20.07.2007 *****c,g 5; 10; 20; 40; 80 4; 7; 16; 30; 42 2 4 0.05 Heteroptera R. linearis e Adult B, E 01.03.2007 ***** c 3.3; 6.5; 13; 26; 52 NA 2 5 NA Megaloptera S. lutaria Larva A 03.08.2007 * g 0.2; 1; 5; 25; 125; 625 0.3; 1; 5; 27; 65; 327 the field and Daphnia magna, Procambarus spec. and Neocaridina denticulata sinensis were cultured at Alterra as reported in the supporting information (Rubach et al. in press). Procambarus spec., here tested in both adult and juvenile stages, is better known as the ''Marmorkrebs'' (marbled crayfish), a parthenogenetic freshwater crayfish species belonging to the Cambaridae showing only female phenotypes, at least under culture conditions (Scholtz et al. 2003;Vogt et al. 2004;Martin et al. 2007). In a recent study, this species has been used in an ecotoxicologic test (Vogt 2007). The species N. denticulata sinensis var. red is a tropical shrimp, also called the ''sherry red shrimp,'' and although particularly popular with hobby aquarists, it has rarely been employed in ecotoxicology as a test species. All test animals, including the cultured species, were transferred into 0.45-lm membrane pressure filtered and 24-h aerated water pumped from the groundwater horizon of The Sinderhoeve and supplied with appropriate food to acclimatize to the test medium for at least 3 days before testing. The same water was used to prepare either the test media or inter dilutions by spiking and homogenizing the filtered and aerated water with the adapted volumes of stock solution after the exact concentration of chlorpyrifos in the stock solution had been determined.
To address differences in sensitivity and species-specific requirements, such as prevention of cannibalism, the specific test design of toxicity experiments varied slightly among tested species (Table 1). Cannibalistic species were either tested in 4.2-l aquaria divided into the necessary number of compartments with inlets of stainless steel gauze, singly in 100-ml screw cap glass beakers or in 600-ml borosilicate beakers, which were divided into four compartments with stainless steel gauze. Expected noncannibalistic species were either tested in 250-ml SCHOTT flasks, 1.5-l WECK beakers, or 600-ml borosilicate beakers. When required for a particular species, these test systems were provided with stainless steel hook-shaped gauze pieces to provide a structural element. To maintain constant temperature, the aquaria, the WECK beakers, and the 600-ml beakers were kept in a water bath, whereas the 100-ml screw-cap glasses, the 250-ml SCHOTT flasks, and the aquaria (for one experiment (Procambarus spec. adults]) were kept in an incubator cabinet (Sanyo MIR 552). All experiments were performed at the same light-todark regime (16:8 h) with an average light intensity of 13 lmol Á s -1 Á m -2 (minimum to maximum 10.5 to 15.5 lmol Á s -1 Á m -2 ). However, species known to be stressed by light were shaded using aluminium foil to decrease stress. The experiments were performed at a temperature of 17°C ± 3°C, an average pH of 7.61 ± 0.41 (measured with electrode pH 323B/set, WTW, Germany), and an average dissolved oxygen level of 8.8 ± 1.8 mg/l (measured with electrode Oxi330/set, WTW, Germany). These parameters were measured at 0, 48, and 96 h in at least one replicate per treatment. To decrease the losses of chlorpyrifos through evaporation, the test vessels were covered with parafilm or cling foil during exposure. However, if atmospheric breathers were tested, beakers were only covered with nylon gauze to prevent escape of the organisms. All experiments were run in a static exposure regime with initial peak dosage at the start of the experiment. After insertion of the test animals (t = 0) with appropriate forceps or low-volume pipettes, water samples were taken from the test systems to determine the measured nominal concentrations. The water samples were extracted with n-hexane (99% pure) in graduated glass tubes by horizontal shaking for 3 min, subsequent layer separation for 10 min, after which approximately 1 ml of the upper (n-hexane) layer was transferred into amber GC vials and capped using lids with Teflon-lined septa. Samples were stored at -20°C until GC-ECD analysis. Water samples were taken after 0.5 (= 0), 48, and 96 h of application of chlorpyrifos. The intended and verified concentrations (t = 0) are listed in Table 1 together with the replication of treatments, number of test animals per replicate, and test system used.
Investigated end points of toxicity in each test were mortality and immobilization at 24, 48, 72, and 96 h of exposure. At these time points, the number of dead and immobile animals were counted in each replicate. For every species, clear criteria were set beforehand to distinguish between immobilization and death. In detail, test animals showing abnormal movement (paralyzed limbs, inability to walk, missing reflexes) compared with control animals after repeated agitation with forceps were classified as immobile. Subsequently, if immobile animals did not show any visible movement within 30 s after repeated agitation, they were classified as dead. To distinguish between death and immobility, immobile specimens of some species (Asellus aquaticus, Chaoborus obscuripes, Cloeon dipterum, D. magna, G. pulex, M. angustata, N. denticulata, Parapoynx stratiotata, Plea minutissima, and Sialis lutaria) were investigated using a binocular microscope, whereas the other species were controlled for effects macroscopically. D. magna was classified in the second step as dead if no heartbeat could be detected within 30 s. Ranatra linearis was classified as dead if no movement was detected after removing it from the water and putting it back upside down on the water surface, because mobile animals immediately turned themselves back, and immobile specimens would show at least limb movements in this position.
For analysis of the effect data, first the total number of dead and immobile animals was calculated per replicate and observation time point, and dead animals were also counted as immobile. All subsequent calculations were performed on basis of the measured initial concentrations of each single replicate (for rationale see later text). The total number of dead and immobile animals per replicate, together with the initial number of test animals, was used to calculate EC 50s and LC 50 s, respectively. The calculation of EC 50 and LC 50 values was performed for every end point and every observation time by means of log-logistic regression using the software GenStat 11th edition (Lawes Agricultural Trust 2009, VSN International Ltd., Oxford, UK) and Equation 1, with y being the fraction of dead or affected test animals (dimensionless), conc being the applied dose in lg/l on basis of the measured concentrations at t = 0, and the parameters a being ln EC 50 , b being slope in l/lg, and c being fraction of background effect, all of which were fitted:
For both mortality and immobility of species exposed to chlorpyrifos, SSDs (Posthuma et al. 2002) were constructed per observation time point. For this the E T X 2.0 program (Van Vlaardingen et al. 2004) was used, which fits a log-normal model to the data. For each SSD, the geometric mean of the log-transformed toxicity data (log HC 50 SSD data), their SD (r' SSD data), and 95% confidence interval (CI SSD data) were used as indicator and uncertainty measures for the variation in sensitivity observed (Aldenberg and Jaworska 2000). Furthermore, the median 5% and 50% hazardous concentrations (HC 5 and HC 50 ) and their confidence limits were calculated. The goodness-of-fit was tested using three different tests for normality: the Anderson-Darling-test, the Kolmogorov-Smirnoff test, and the Cramer van Mises test.
A prerequisite for the correct interpretation of effects of chemicals on biota is the confirmation of intended exposure regimes in experimental studies. Table 1 lists the measured concentrations of chlorpyrifos at the start of the experiments. The intended concentrations were well achieved with an average coefficient of determination of 0.97 ± 0.04, average slope = 0.86 ± 0.22, and average intercept = -0.55 ± 2.77 for linear regression across all experiments of intended and measured concentrations at t = 0. As expected, during the course of the experiment, the measured concentrations of chlorpyrifos in most of the test media decreased (Table 2). The experiments with the different species, however, showed differences in dissipation of chlorpyrifos ranging from 55.2% to 111.7% remaining after 48 h exposure and 21.9% to 120.7% remaining after 96 h of exposure (Table 2). Although no consistent pattern could be found when comparing species or treatments, the interplay of factors, such as animal size, bioconcentration, evaporation, and degradation, could explain these differences. For further calculation of L(E)C 50 values and SSDs, the initial concentrations in the static systems were used because of the short test duration and the relatively long organism recovery time shown for organophosphates (Ashauer et al. 2007).
In general, concentrations of chlorpyrifos in control replicates were lower than the respective LODs, but occasionally chlorpyrifos was measured in single control samples (C. obscuripes, D. magna, N. denticulata, P. stratiotata, Procambarus spec. juveniles, and S. lutaria), which explains the high variability in the intercept reported previously. These exceptions are related to cross-contamination of controls with chlorpyrifos due to its high volatility and associated contamination routes, especially when experiments with high-exposure concentrations were performed. In all the controls showing cross-contaminations, immobilization or mortality was \ 10% and therefore tolerated in the presented study. In the experiment with M. angustata, all control replicates were contaminated on average with 0.143 lg/l chlorpyrifos, and immobility was induced in one to two control animals per replicate (20% to 40%), leading indirectly to cannibalism. As a result, control mortality and increased mortality in the lower concentrations was observed during the experiment. Cannibalism was also observed at the lower concentrations, but not in the intermediate and high concentrations, where no mortality but full immobilisation was induced in all test animals and thus prevented cannibalism. Due to a current lack of data on the effects of chlorpyrifos for this species, it was decided rather to correct the total number of immobile or dead animals for further analysis instead of excluding the species. The data correction was performed by setting both concentrations of the contaminated control replicates and the induced immobility/mortality at 24 h in the first two treatments to zero.
Effects of chlorpyrifos induced in a range of freshwater arthropods by short-term exposure in simple laboratory test systems are presented as concentration response relations for mortality and immobility (Tables 3 and 4). Respective concentration-response parameters are reported (according to Equation 1) next to L(E)C 50 values and their confidence limits. At least the highest test concentrations of chlorpyrifos induced 100% immobility within the 96-h exposure in all species except in C. dipterum, where 93% of all test animals were immobile at the end of the experiment. In P. stratiotata, all initial test animals in the highest concentration were immobile at 72 h of exposure, but subsequent recovery occurred, resulting in 70% immobilization at the end of the experiment. In contrast, at least the highest concentrations induced 100% mortality only in 6 of the 14 experiments (Anax imperator, G. pulex, P. minutissima, Procambarus spec. adults and juveniles, and R. linearis), and for some species no mortality (S. lutaria, M. angustata) or low mortality (A. aquaticus) was observed within the time of exposure, even at the highest concentrations. For the remaining 5 species, we found between 70% and 87% mortality (C. obscuripes, C. dipterum, D. magna, N. denticulata, Notonecta maculata, and P. stratiotata) at their respective highest concentrations. Figure 1 illustrates that for some species the differences between lethal and sublethal effect concentrations are substantial independent of time, whereas others show a relatively good match.
The observed effects on mortality and immobilisation for the tested species is in the range reported in literature, which is up to 3 to 4 orders of magnitude for both lethal and sublethal effects for up to 96 h (Van Wijngaarden et al. 1993;Maltby et al. 2005;Rubach et al. 2010). In addition, L(E)C 50 values for particular species agree well with previous findings, with the exception of C. obscuripes, for which an LC 50 22 times higher (6.6 lg/l, 96-h exposure) was previously determined (Van Wijngaarden et al. 1993). These investigators also observed differences in the concentrations inducing mortality and immobility for some arthropod taxa (A. aquaticus, Proasellus coxalis, G. pulex, C. dipterum, C. horaria, and C. obscuripes) and, again, no mortality was induced within 96 h of exposure for A. aquaticus and also Caenis horaria in their study. Although the experiments of Van Wijngaarden et al. (1993) were either performed in flow-through or semistatic test systems and therefore under constant exposure, the congruence of results with our study shows that for chlorpyrifos exposures up to 96 h the initial (peak) concentrations are equally representative for effects induced within this period of time.
As expected, the effects of chlorpyrifos on both mortality and immobility increase in time in all tested species, which can be seen from the overall decrease in L(E)C 50 values (Fig. 2 and Tables 3, 4). Chlorpyifos mainly acts on the nervous system by inhibiting the enzyme acetylcholinesterase, leading to a synaptic block and therefore inhibiting Taken from (Rubach et al. in press) electric signal transmission. This first leads to sublethal intoxication symptoms, and subsequent death is likely caused by final respiratory failure (Eaton et al. 2008). Hence, as expected, observations of immobility consistently resulted in lower EC 50 values in time compared with their respective LC 50 values, although the difference between these end points decreased during the course of the experiment, especially for D. magna, for which the LC 50 /EC 50 ratios decreased from 128.6 to 4.75 in 72 h (Fig. 2). In general, it is logical that the effect concentrations for immobility and mortality will converge to the same value with time; however, it is evident that this does not occur with the same speed for all the tested species (Fig. 2). For some species, the differences between LC 50 and EC 50 even stayed relatively constant within the 96 h of test duration. However, for the species A. imperator, C. dipterum, G. pulex, Procambarus spec., and R. linearis, a good match between effective and lethal concentrations was observed right from the start of the experiments (estimated LC 50 /EC 50 ratios approximately 1; see also Figs. 1 and 2). In contrast, for a third group of species (N. denticulata, N. maculata, and P. stratiotata), the difference between EC 50 and LC 50 values for a particular species does not necessarily change in time (LC 50 /EC 50 ratios were constantly approximately 2, 2, and 9, respectively). In addition, the extent to which LC 50 and EC 50 values differ for certain time points seems rather speciesspecific, especially for S. lutaria and M. angustata, in which no significant incipient mortality was induced by the applied concentrations but in which immobility was induced at quite low concentrations (Fig. 2). This is interesting in the sense that these species-specific differences in incipient mortality or immobility can be either due to differences in toxicokinetics and/or toxicodynamics. For instance, on one hand, S. lutaria, M. angustata, and A. aquaticus could have the ability to either decrease or regulate uptake and/or elimination of chlorpyrifos, to biotransform chlorpyrifos slower to the chlorpyrifos-oxon, or to detoxify the latter quickly and therefore delay incipient mortality significantly, all of which would relate to differences in toxicokinetics. In contrast, differences in the species responses might be caused by other processes pertaining to toxicodynamcis, e.g., differences in the interaction of chlorpyrifos and acetylcholinesterase (target enzyme) or in the ability to compensate or repair damage. For details on toxicokinetics and toxicodynamics see Ashauer et al. (2006). Rubach et al. (in press) measured uptake and elimination kinetics of 14 C-labeled chlorpyrifos in the same species and indicated that B38% of the variation in sensitivity (EC 50 , immobilisation in 48 h, same data) may be explainable by uptake and B28% by elimination kinetics. Interestingly, S. lutaria, A. aquaticus, and M. angustata, which responded with a remarkable concentration difference between incipient immobility and mortality in this study, show high bioconcentration factors (9625, 3242, and 5331 lg/kg ww , respectively). Because their uptake rates are moderate to high, and because immobility is effectively induced at much lower concentrations, differences in uptake itself can be excluded. More likely are differences in biotransformation rates (either bioactivation or detoxification) or a highly efficient compensatory gene-regulation ability. The most insensitive of the investigated species, N. denticulata, shows high uptake and high elimination rates and therefore a moderate bioconcentration, which partly explains its insensitivity.
Clearly, the extent of variation in observed sensitivity to chlorpyrifos across species highly depends on the end point under consideration, which is already evident from the concentration-response relations but also from the SSDs shown in Fig. 3. The SSDs also indicate by the ''left shift'' of both mortality and immobilisation that effects increase in time. The slopes of the SSDs for immobility do not seem to be significantly different; however, the variability increases slightly in time, as indicated by an increase in r' (Table 5). Nevertheless, the CIs of the SSD (Table 5) show a strong overlap; therefore, this trend cannot be confirmed reliably, and species variation in immobilisation might still be relatively constant in time. The variation in mortality is generally much higher and also relatively constant in time (r', Table 5) until 72 h of exposure, after which a sharp increase in species variability was observed. This is evident from the high r' (Table 5) and low slope of the SSD for 96 h, but only if all available species are included in the SSDs (Fig. 3). This impact on the slope is an artefact caused by the species selection and the close-to-zero mortality in A. aquaticus, M. angustata, and S. lutaria, for which no LC 50 could be calculated and which were thus The species M. angustata is not shown because only 24-and 48h observations were available, and no mortality was observed not included in the SSDs for exposure times B96 h. If these insensitive species (with high LC 50 values) had been included for all time points, the slopes of these mortality SSDs would have been similar in time or even lower than the one for 96 h of exposure, and r' would have indicated an even bigger variation. Therefore, the variability in immobility is generally lower than for mortality; however, both are relatively stable in time if based on the same selection of species. The ratio of HC 50 (mortality)/HC 50 (immobility) derived from Table 5 decreases slowly in time (2.9, 2.4, 2.1, 2.0) when based on the same species selection; similarly, if all 14 species are included in the 96h mortality SSD this ratio is 4.4. This shows that strong differences in time-dependent toxicity exist between species and that the SSD can only account for this if species selection is restricted to similarly reacting groups.
Choice of End Point for ERA Until now, from the perspective of individual sensitivity response, the presented data support the assumption that immobility is the better end point when investigating a neurotoxic substance, such as chlorpyrifos, because paralysis is the first visible symptom, and a species-specific time lag until incipient mortality was found for several species. In addition, also from a population sustainability viewpoint, immobility is almost just as relevant as mortality.
Where on one hand, mortality is unilateral in the sense that a dead specimen cannot become alive again, any immobile or otherwise sublethally affected specimen may become mobile again and thus be able to contribute to a population's sustainability. In contrast, an immobile specimen is likely going to be outcompeted, starved, drifted, or predated quickly under field conditions and also more prone to multiple stress regimes. However, when the risk assessment is based on SSDs rather than on safety factors, including only data based on immobility, illogically does not always yield the most conservative threshold but may deliver a more confident estimate of the HC 5 due to less variation in the species selection. Steep SSDs with higher confidence, as calculated here for immobility, will deliver less conservative HC 5
0.1 1000 10000 72h 24h 48h 72h 0.0 0.0 0.2 0.2 0.4 0.4 0.6 0.6 0.8 0.8 1.0 1.0 1000 1000 EC50, immobility data [µg/L] EC50, immobility data [µg/L] 0.1 1000 10000 72h 24h 48h 72h 0.0 0.0 0.0 0.2 0.2 0.2 0.4 0.4 0.4 0.6 0.6 0.6 0.8 0.8 0.8 1.0 1.0 1.0 0.001 0.001 0.001 0.01 0.01 0.01 0.1 1 1 0.00 0.001 1 .0 .0 0 1 0 1 .1 .1 0 0 1 1 0 10 00 00 1 1 1 0 0 0 10000 0 00000 00000 1 1 1 1 1 0 10 10 00 00 1 1 100 1000 10000 100000 100000 100000 LC50 data [µg/L] LC50 data [µg/L] LC50 data [µg/L] Potentially Affected Fraction Potentially Affected Fraction Potentially Affected Fraction Potentially Affected Fraction Potentially Affected Fraction Potentially Affected Fraction 24h 24h 24h 24h 48h 48h 48h 48h 72h 72h 96h 96h 96h 96h 96h 96h 96h 96h 24h 48h 72h 24h 48h 72h 96h 96h 96h 96h Fig. 3 SSDs for freshwater arthropod species under 24-, 48-, 72-, and 96-h exposure to chlorpyrifos constructed from observations of immobility (upper panel) and mortality (lower panel). For mortality, SSDs were calculated from 24 to 96 h using the same closed data set (minimum number) of species (k = 11, empty symbols) and for 96 h using the maximum number of species available (k = 14, filled symbols) to illustrate the influence of species selection. In Table 5, the SSD indicator values and statistical details are given values than shallower SSDs, such as shown here for mortality after 96 h of exposure. This is especially the case if the lower end of the curve shows a bad fit with the data (as the 96-h mortality SSD for all species in Fig. 3). Such SSDs can, however, be excluded using goodness-of-fit measures, especially the Anderson-Darling test, which is sensitive for the quality of fit in the lower concentration range (Van Vlaardingen et al. 2004). All presented SSDs passed the three performed goodness-of-fit tests with p \ 0.001, except for the 96-h mortality SSD with all species included, which did not pass the Anderson-Darling and the Cramer van Mises tests (both p \ 0.1). The lower confidence limit of the HC 5 (LLHC5) derived from an SSD may serve as a protective threshold in higher-tier risk assessment (Maltby et al. 2009;Brock et al. 2010). This value can differ substantially between different observation times and different end points (see Table 5) depending on the species selection. Although in general a rather conservative threshold, the protectiveness of the LLHC5 highly depends on the incipient time of the effect, the end point included (which correlates to the incipient of effects), the extent to which the species included in the SSD vary in their sensitivity, and other quality criteria as reviewed in Brock et al. (2010). Our results show that the most confident estimates can be derived with an SSD when immobility after sufficient time of exposure is chosen as an end point for the SSD. Herewith, if these ''quality criteria'' are addressed, the convolution caused by inclusion of insensitive species does not touch the usefulness of the SDD as a tool for ERA, especially if the most sensitive group for one chemical is well represented in the SSD (Van Den Brink et al. 2006); however, it also clearly shows that an SSD does not necessarily represent the true existing variation in sensitivity.
Other end points, in addition to mortality and immobility, describing population sustainability, such as reproduction, can be derived from chronic testing but cannot be deduced from short-term tests. To derive a good and conservative proxy for population sustainability, another sublethal end point such as postexposure feeding inhibition should be considered for short-term testing (McWilliam and Baird 2002;Satapornvanit et al. 2009). If a specimen is not able to feed within a given time period, e.g., 24 h after the end of a short-term exposure, it is rather unlikely that it is able to contribute to the population's sustainability. This, however, is yet far from being taken into account in current risk-assessment practices. Another problematic issue for the selection and definition of an appropriate end point for ERA when comparing effects on species are the criteria that must be set for this particular end point. For instance, in this study, transparent species could be observed for heartbeat and thus had a good criterion by which to distinguish death from immobility. Nontransparent and highly-sclerotised species, however, do not provide such clear-cut criteria to determine clinical death. In the present study, the end point criteria for each species were rather well defined, thus minimizing such described difficulties.
To improve ERA, the identification of the best end point for assessing the risk of a certain group of chemicals must be based on its exposure scenario, its functional relevance for the mode or mechanism of action, its toxicity in time or on other previous knowledge, and the ecologic consequences of a given end point.
The presented data set demonstrates that freshwater arthropod species can be highly variable in their dynamic response toward a particular stressor. What exactly causes these differences in sensitivity within such a narrow group of taxa in response to chlorpyrifos, an insecticide designed to affect this particular test group, remains mostly unclear. Hypothetically, the differences in effects among tested species are partly related to differences in bioconcentration, but biotransformation and/or differences in the amount of internally caused damage and/or differences in their abilities to recover or repair the induced damage must also play a major role. However, clear mechanistic explanations remain open, considering the current lack of knowledge on how these processes differ in the tested arthropods. Furthermore, presented findings illustrate the importance of considering an appropriate end point for a protective risk assessment based on knowledge about the mode of action of a particular group of compounds. In general, but surely for neurotoxic compounds, such as organophosphates, immobilization in favour of mortality may be the appropriate end point to use for further risk assessment. Mortality and/or immobility, however, might not be the appropriate end points for other nonneurotoxic compounds, such as growth or molting inhibitors, endocrine disruptors and mutagenic or genotoxic substances. For other modes of action, the most relevant end points for further ecologic risk assessment still need to be defined on basis of the function(s) affected by the mode of action and the relevance of these function(s) in a field scenario. The test-length of short-term toxicity tests for some of those types of compounds is not sufficient, but more realistic approximations of acute risks may be derived from shortterm tests if a postexposure feeding assay is performed after a 96-h exposure. However, sublethal effects are sometimes reversible, meaning that recovery even on individual level is possible. In ERA, probabilistic approaches, such as the SSD, provide useful tools to derive protective threshold values; however they do not necessarily account for true variation in sensitivity. Compared with the mechanistic effect models described, these tools are not based on the processes of toxicity, and, therefore, extrapolation of information from one chemical to the other is difficult. Future research may be able to relate certain species characteristics to these processes and also to modeof-action-specific sensitivity and thus provide a more mechanistic understanding on which to base and evaluate ERA.
a Highest treatment (107 lg/l) induced 100% mortality (denoted 'empty') after 24 h in this species; therefore in A the third highest and in B the second highest treatment are shown
a b For the 96-h mortality data, two SSDs were calculated, maximized and minimized data set, for comparison
Acknowledgments This work was financially supported by
Background: Cancers are some of the leading causes of human deaths worldwide and their relative importance continues to increase. Since an increasing proportion of cancer patients are acquiring resistance to traditional chemotherapeutic agents, it is necessary to search for new compounds that provide suitable specific antiproliferative affects that can be developed as anticancer agents. Propolis from the stingless bee, Trigona laeviceps, is one potential interesting source that is widely available and cultivatable (as bee hives) in Thailand. Methods: Propolis (90 g) was initially extracted by 95% (v/v) ethanol and then solvent partitioned by sequential extractions of the crude ethanolic extract with 40% (v/v) MeOH, CH 2 Cl 2 and hexane. After solvent removal by evaporation, each extract was solvated in DMSO and assayed for antiproliferative activity against five cancer (Chago, KATO-III, SW620, BT474 and Hep-G2) and two normal (HS27 fibroblast and CH-liver) cell lines using the MTT assay. The cell viability (%) and IC 50 values were calculated.
Results: The hexane extract provided the highest in vitro antiproliferative activity against the five tested cancer cell lines and the lowest cytotoxicity against the two normal cell lines. Further fractionation of the hexane fraction by quick column chromatography using eight solvents of increasing polarity for elution revealed the two fractions eluted with 30% and 100% (v/v) CH 2 Cl 2 in hexane (30DCM and 100DCM, respectively) had a higher antiproliferative activity. Further fractionation by size exclusion chromatography lead to four fractions for each of 30DCM and 100DCM, with the highest antiproliferative activity on cancer but not normal cell lines being observed in fraction# 3 of 30DCM (IC 50 value of 4.09 -14.7 μg/ml). Conclusions: T. laeviceps propolis was found to contain compound(s) with antiproliferative activity in vitro on cancer but not normal cell lines in tissue culture. The more enriched propolis fractions typically revealed a higher antiproliferative activity (lower IC 50 value). Overall, propolis from Thailand may have the potential to serve as a template for future anticancer-drug development.
Cancers are some of the major fatal diseases to humans. Chemotherapy is one of the most widely used approaches for the treatment of many cancers, but the long-term use of chemotherapy can lead to drug resistance via several different mechanisms, such as gene mutation, DNA methylation and histone modification. These resistance mechanisms have been reported to play important roles in the resistance of cancers to chemotherapeutic agents [1]. Thus, patients are gradually developing resistance to widely used and standard chemotherapeutic agents, such as 5-fluorouracil, taxol, doxorubicin, cisplatin, campothecin, paclitaxel and topotecan [2]. Due to this resistance to cancer drugs, it is important to find new anticancer agents in order that they can be developed into novel anticancer drugs that can circumvent the existing resistance mechanisms. Herbs and other natural plant products have become interesting sources for this purpose, but animal modified or selected plant products have been largely overlooked.
Propolis, one of the economic natural products from bees, is an interesting source for several bioactivities, such as antimicrobial as well as anti-cancer. Although an animal product it is largely plant based in its chemical origins. It is a sticky resin and varies in colour, including brown, green and red amongst others, based upon the plant exudates that the bees have selectively collected from flower buds, leaf buds and tree barks. These plant resins are mixed with waxes and other bee excretions, including enzymes [3], to form the final propolis product. Although used as a sealing wax for filling cracks and repairing combs, and in some cases embalming wax, the principal use of propolis in the beehive is as a protective barrier against their enemies [4]. As such it has broad antimicrobial activities and has been extensively used in the traditional medicine [5]. In Europe, propolis was accepted as an official drug due to its antibacterial activity during the last 400 years [6]. Furthermore, propolis has been long used as a dietary supplement for disease prevention [7], since it can provide antimicrobial activity, including antiviral activity [8], anti-inflammatory [9], immunomodulatory [10], antitumor [11] and antioxidant effects [12].
The chemical composition of each type of propolis and its associated bioactivities mainly depend on the macro-and micro-geographical regions, due to the differences in the plant resin compositions or available plant species [8,13], and on the bee species, due to the different preference for food and resin plants and foraging distances between bee species [14]. For example, the pollen in the propolis from Apis mellifera in the Preveza region of Northwest Greece was mainly from Pinaceae [14] whilst in the propolis in Brazil it was mainly from two poplar trees, Hyptis di Varicata and Baccharis dracunculifolia [15,16], suggesting quiet different sources of plant resins and volatiles for the propolis production. In addition, the propolis from A. mellifera in Brazil, Chili and Myanmar presented a different in vitro cytotoxicity activity against the PANC-1 human pancreatic cancer cells in a nutrient-deprived medium [16][17][18].
Although propolis is typically a complex mixture of diverse compounds from plant resins, volatiles, pollen and animal enzymes, etc, numbering over 300 characterized components, the reported bioactivities associated with propolis have been found in both the crude and the purified extracts [12,19], suggesting that they may be associated with single compounds rather than complex interactions between compounds, and so amenable to purification. To date, most of the active compounds from purified extracts have been found to be phenolics and polyphenolics [20]. Propolin A and propolin B, which belong to the prenylflavanone group, were the main chemical components that could be isolated using Nuclear Magnetic Resonance (NMR) from Taiwanese propolis [21]. Both components displayed an in vitro antiproliferative effect on human melanoma, C6 glioma and HL-60 cell lines in tissue culture.
Caffeic acid phenethyl ester (CAPE), the main component from A. mellifera propolis in Chili, along with a considerable number of flavonoid compounds, such as p-cumaric acid and ferulic acid, were purified and found to display an in vitro antiproliferative effect on the KB, DU-145 and Caco-2 cell lines in tissue culture [12].
Given that the main active components found to date are plant derived flavonoids, it would seem likely to be better to preserve the polyphenolic fraction of propolis after purification.
Propolis from the stingless bee Trigona laeviceps Smith (Hymenoptera: Apoidae) was used in this research since these bees can be commercially cultivated in a sustainable and potentially ecologically friendly manner in artificial hives (with the potential add on value of certain crop plant pollination), and can provide a lot of propolis per hive. In addition, the bees are widely distributed throughout Thailand. Although Umthong et al. reported that a high cytotoxicity on the SW620 colon cancer cell line was obtained from the crude extract of propolis harvested from T. laeviceps in Thailand [22], it is not recommended to use or consume propolis in the form of a crude extract because it may still contain various adverse bioactivities. In this research, we attempted to purify the ethanol crude extract of T. laeviceps propolis from Samut Songkram, central Thailand, using a bioassay-guided isolation procedure. Solvent partitioning of the propolis, based upon solvent polarity, was performed and human cell lines derived from five different types of tissue cancers were used to screen for any in vitro antiproliferative affect in comparison to the two normal cell lines.
Propolis of T. laeviceps was collected from an apiary in the Samut Songkram province, central Thailand, and was kept in the dark at 4°C until use.
The method of propolis extraction followed that reported by Najafi et al. [23]. Briefly, propolis (90 g) was cut into small pieces and was then extracted by 95% (v/ v) ethanol (400 ml) at 15°C with shaking at 100 rpm for 20 h. The suspension was then clarified of residual propolis solid by centrifugation at 7,000 rpm for 15 min at 20°C. The supernatant was harvested and kept whilst the pellet was re-extracted and then clarified as above except using 100 ml and not 400 ml of 95% (v/v) ethanol. The two ethanolic extracts (supernatants) were pooled together and evaporated in a rotary evaporator (40°C). The obtained residue (crude ethanolic extract) was weighed and stored at -20°C at dark.
The ethanolic extract of the propolis was dissolved in 80% (v/v) methanol until it was not sticky and then an equal volume of hexane was added, stirred (15 mins) and then allowed to phase separate in a separating funnel. The upper hexane phase containing the non-polar compounds was harvested and kept whilst the 80% (v/v) methanol phase containing polar compounds was extracted three more times with hexane in the same manner. The four hexane extracts were pooled, the solvent evaporated in a rotary evaporator (40°C), and the sticky liquid residue weighed to give the yield of the hexane extract. The 80% (v/v) methanol phase was then mixed with an equal volume of H 2 O to increase the partition coefficient (solubility) of polar compounds, and then extracted four times with an equal volume of CH 2 Cl 2 as per hexane above, except that following phase separation the CH 2 Cl 2 phase containing the less polar compounds was the lower phase. The pooled CH 2 Cl 2 phases and the residual 40% (v/v) methanol phase were separately evaporated to remove the respective solvent in a rotary evaporator (40°C), and the residues were each weighed to give the yield of the CH 2 Cl 2 and methanolic extracts, respectively.
Silica gel was packed into a glass funnel (50% of internal volume) connected to a vacuum pump. The selected fractions (see Bioassay guided selection below) were mixed with CH 2 Cl 2 until they were not sticky and were then mixed with silica gel, evaporated to dryness and loaded on top of the silica gel containing column. The column was then eluted by various solvents (2,000 ml each) of different polarities starting from the least polar solvent and increasing, that is from 0:1, 1:9, 3:7, 1:1, 7:3 and 1:0 (v/v) of CH 2 Cl 2 : hexane and then followed by 5% and 10% (v/v) MeOH in CH 2 Cl 2 , respectively. A reduced pressure was used in order to enable the flow rate of the solvent through the column to be obtained, and fractions were collected. Each fraction was solvent evaporated in a rotary evaporator (40°C), and the residue weighed before being dissolved in DMSO to the required concentrations to assay for antiproliferative activity as detailed below.
Further partial purification of each selected fraction was performed by size exclusion chromatography using a Sephadex LH-20 column (10 ml internal volume). Each selected fraction was dissolved in a 1:1 (v/v) ratio of MeOH: CH 2 Cl 2 until it was not sticky and was loaded on top of the column. The column was then eluted with 10 ml of a 1:1 (v/v) ratio of MeOH: CH 2 Cl 2 , collecting 2 ml fractions. The four fractions obtained were then solvent evaporated by a rotary evaporator (40°C), the residue weighed and then dissolved in DMSO to the required concentrations ready to assay for any selective antiproliferation activity as detailed below.
One-dimensional thin layer chromatography (TLC) was used in order to group the obtained components and to determine the purity of fractions. A silica gel plate (10 cm in height) was used as the stationary phase, and the respective extracts or fractions were spotted at the starting line at 0.5 -1 cm intervals. The mobile phase solvents used (one per TLC plate) were 1:0, 3:1 and 1:1 (v/ v) ratio of CH 2 Cl 2 : hexane. When the mobile phase had almost reached the top of the plate, the samples were visualized under U.V. light (254 nm). Alternatively, the gel plate was sprayed by a 5% (v/v) H 2 SO 4 /0.03% (w/v) α-naphtal methanolic solution, dried in an oven or hot plate and visualized under U.V. light (350 nm).
Each of the three crude solvent partitioned extracts (hexane, dichloromethane and methanolic) of propolis, obtained as detailed above, were evaluated for their in vitro antiproliferative activity on five selected cancer and two normal cell lines in tissue culture using the MTT assay as detailed below. The extract which provided the best selective antiproliferative activity, that is the highest activity on the cancer cell lines but not on the normal cell lines, as determined by comparison of the IC 50 values, was selected for further partial purification by quick column silica chromatography. In the same way, each fraction obtained from the quick column chromatography was likewise assayed for selective antiproliferative activity on the cell lines and this was used to select fractions for further size exclusion chromatography (see above).
The five selected cancer cell lines for screening for the in vitro antiproliferative bioactivity were derived from colon (SW620), breast (BT474), hepatic (Hep-G2), lung (Chago), and stomach (Kato-III) tissue cancers. The two normal cell lines used were of liver (CH-liver) and fibroblast (HS-27) origins and were used as comparative controls to check for selective specificity towards cancer cells rather than all dividing cells. All cell lines were cultured in RPMI medium containing 5% (v/v) fetal calf serum (complete medium) at 37°C in a humidified air atmosphere containing 5% (v/v) CO 2 .
Cultured cells (5,000 cells) in 200 μl complete media were transferred into each well of a flat 96 well plate and then incubated at 37°C in a humidified air atmosphere enriched with 5% (v/v) CO 2 for 24 h in order to let the cells attach to the bottom of each well. The cultured cells were then treated with the tested propolis extract (triplicate wells per condition) by the addition of 2 μl of serial dilutions of the propolis extract dissolved in DMSO to give a final concentration of 100, 50, 25, 12.5, 6.25 and 3.125 μg/ml. In addition, 2 μl of DMSO alone was added to another set of cells as the solvent control. The cells were then cultured as above for another 48 h prior to the addition of 10 μl of a 5 mg/ml solution of 3-(4, 5-dimethylthiazol-2-yl)-2, 5-diphenyltetrazolium bromide (MTT)into each well. The incubation was continued for another 4 h before the media was removed. A mixture of DMSO (150 μl) and glycine (25 μl) was added to each well and mixed to ensure cell lysis and dissolving of the formasan crystals, before the absorbance at 540 nm was measured. Three replications of each experiment were performed and the half maximal inhibitory concentration (IC 50 ) of each extract was calculated as detailed below.
The obtained absorbance at 540 nm was used to determine the percentage of cell survival assuming that 100% survival was obtained from the solvent only control and that no differences in metabolic activity existed between surviving cells under the different conditions. Under these assumptions, the percentage survival of the treated cancer and normal cultured cells was calculated according to the formula below:
The mean (± 1 standard deviation (SD)) cell survival (%) was plotted against the corresponding propolis extract concentration and the best fit line was used to derive the estimated IC 50 value from the concentration that could provide a 50% cell survival.
From 90 g of T. laeviceps propolis a yield of 18 g (20%) of crude ethanol extract was obtained as a sticky brownish to dark brown resin with a distinctive smell. When assayed in vitro on the five cancer and two normal cell lines, using the MTT assay, the IC 50 (μg/ml) values for the five cancer cell lines ranged from 1.9-fold lower (SW620 and BT474) to only 1.04-fold lower (Hep-G2) than that for the control HS27 cell line (Table 1).
However, the CH-Liver control cell line was some 1.3fold more sensitive than the HS27 cell line and so only three of the cancer cell lines showed reduced IC 50 (μg/ ml) values compared to the control CH-liver cell line. Over all, it indicates that the crude ethanolic extract of T. laeviceps propolis contained some antiproliferative activity, but with the overlapping variation within the cancer and normal cell lines the existence of any selective specificity for cancer cells was less clear.
After resolvation of the crude ethanolic extract in 80% (v/v) methanol, sequential solvent partitioning with three different polarity solvents was used to further fractionate the propolis extracts. For each of the three solvent extractions the new fractions obtained gave different appearances (Table 2), but the CH 2 Cl 2 extract showed no significant antiproliferative bioactivity on four of the five cancer cell lines tested (Figure 1), whilst the methanol extract was only significantly effective against two of the five cancer cell lines. In contrast, the hexane extract, and thus the low polarity compounds, were the most affective, showing a strong antiproliferative affect on all five cancer cell lines above that seen on the two control cell lines (Figure 1).
Given that the hexane partitioned fraction showed the greatest antiproliferative effect on the five tested cancer cell lines (Figure 1), it was selected for further enrichment by quick column chromatography. The eight fractions obtained from the different elution solvent mixtures revealed seven different appearances (70DCM and 100DCM were similar) (Table 3), but the highest polarity fraction, obtained from elution with 1:9 (v/v) ratio of methanol: CH 2 Cl 2 revealed essentially no antiproliferative activity (Figure 2). In contrast, based on the IC 50 (μg/ml) values, the fraction obtained from the 1:9 (v/v) ratio of CH 2 Cl 2 : hexane revealed a general antiproliferative effect and was not specific to the cancer cell lines. The pure hexane eluted fraction revealed a low cytotoxicity to the control CH-liver cells and a high cytotoxicity on four of the cancer cell lines, the exception being the Chago cell line (66.9 μg/ml compared to 66.2 μg/ml for the control CH-liver cell line). In contrast, the fractions eluted from the 3:9, 1:1 and 7:3 (v/v) ratios of CH 2 Cl 2 : hexane revealed a significant antiproliferation activity on all five cancer cell lines but only a slight affect on the control CH-liver cell line. From these three fractions that which eluted from the 3:7 (v/ v) ratio of CH 2 Cl 2 : hexane (30DCM) was selected for further enrichment by size exclusion chromatography.
In addition, the fraction that eluted in pure CH 2 Cl 2 (100DCM) was selected for further enrichment.
Although the 100DCM extract displayed a high antiproliferation activity on the control normal cell lines, it displayed a very high activity against all five cancer cell lines, and so was selected in case the two activities could be segregated by further enrichment.
As mentioned above, fractions 30DCM and 100DCM were the most active in the MTT based in vitro antiproliferation of cancer cell lines, and so were selected for size exclusion chromatography. For both 30DCM and 100DCM, after size exclusion chromatography, four positive fractions (F1 -F4) were obtained with varying appearances (Table 4). Comparing the color of these fractions after size exclusion chromatography to the color of crude extract in table 2, a lighter color was evident after size fractionation. Table 2) on the five cancer and two normal cell lines in tissue culture, as determined by the MTT assay. Hex = hexane extract, DCM = CH 2 Cl 2 extract and MeOH = methanol extract. Data came from the mean ± 1 SD of percentage of cell viability which was derived from three replicates. Average IC 50 (μg/ml) values of the eight fractions, obtained after quick column chromatography, on the five cancer and two normal cell lines. 100HEX, 10DCM, 30DCM, 50DCM, 70DCM, 100DCM, 5MET and 10MET stand for the fractions eluted in 0:1, 1:9, 3:7, 1:1, 7:3 and 1:0 (v/v) ratios of CH 2 Cl 2 : hexane, and 1:19 and 1:9 (v/v) ratios of MeOH: CH 2 Cl 2 , respectively (see table 3). Data came from the mean ± 1 SD of percentage of cell viability which was derived from three replicates.
Table 4 Characters of isolated active fractions after size exclusion chromatography
These eight fractions (Table 4) revealed different antiproliferation activities on the five cancer cell lines as well as the control cell lines (Table 5). Fractions 1 and 4 from the 30DCM extract revealed no detectable activity on all tested cell lines, whilst fractions 1 and 2 from the 100DCM fraction showed only a weak activity against one or all, respectively, of the five cancer cell lines. Fractions 3 and 4 from the 100DCM extract were among the three most active fractions, but also showed a strong inhibition of the control cell lines. Thus, the antiproliferative affect of the 100DCM fraction on the cancer and normal cell lines was not segregated. In contrast, fractions 2 and 3 from the 30DCM extract showed a moderately and very strong antiproliferative affect on all five cancer cell lines but not the control cell lines (Table 5).
In this research, initial purification steps were performed on the propolis in order to enrich for the selective antiproliferative bioactive fraction, using a cell line antiproliferation assay as a guide. Ethanol was used as the extraction solvent to provide a crude propolis extract as discussed in Orsolic et al. [24]. In addition, Sawaya et al. reported that the soluble compounds could be easily released from the sticky part of propolis by a high percentage of alcohol [25]. Considering the crude ethanol extract of T. laeviceps propolis, not only was a cytotoxic affect on the five selected cancer cell lines seen but also on the two control normal cell types. The IC 50 (μg/ml) between the inhibition of cancer and control cell line proliferation were close (Table 1), preventing useful anti-cancer application. A similar result has been observed in the crude aqueous extract of propolis from A. mellifera in Iran, where although it could inhibit the in vitro growth of some cancer cell lines, it could also stimulate the growth of normal cells [23]. Because there might be compounding effects due to the potential presence of catatonic agents at high bioactive concentrations mixed in with the desired bioactivity compounds, we further enriched the fractions.
Solvent partitioning based upon different solvent polarities was used for further purification and the antiproliferation active compounds were likely to be non-polar or low polar chemicals since the high cytotoxicity was principally observed in the hexane partitioned fraction (Figure 1). This agrees with Marcucci et al. who reported that the main chemical components in propolis were phenolic compounds of a low polarity [26]. Considering the physical appearance of the hexane-partitioned fraction, which was a light brown sticky liquid (Table 2), this contrasts to the crude ethanolic extract of the propolis as a rather dark solid. The change of extract characters after solvent partition based enrichment may represent the removal of some resin and wax from the original propolis. The difference in the observed IC 50 (μg/ml) values for the antiproliferation activity between the five cancer cell lines (< 25) and the two normal cell lines (> 30) was evident (Figure 1), but not different enough to likely be clinically useful. Thus, the hexane-partitioned fraction was then further fractionated by quick column chromatography with elution based upon solvents of increasing polarity. In terms of the IC 50 values for antiproliferation activity, the 30DCM fraction was the one of the most active fractions against the five cancer cell lines but uniquely showed in contrast a very poor activity against the two normal (control) cell lines, and thus showed a good selective antiproliferation activity. That it eluted from a 3:7 (v/v) ratio of CH 2 Cl 2 : hexane is still consistent with the notion that the active compounds are of low polarity. The physical appearance of the 30DCM fraction as a brown solid (Table 3) again supported the removal of other components including resin and wax. Regardless, the IC 50 (μg/ml) value for the CH-liver control cell line was increased, whilst that for the five cancer cell lines was significantly decreased (Figure 2), although the variation in response between the cell lines was marked.
The further enrichment of fraction 30DCM by size exclusion chromatography yielded one fraction, 30DCM-F3, which provided the highest cytotoxicity on cancer cell lines but the lowest cytotoxicity on normal cells (Table 5).
Furthermore, considering the TLC separation pattern, the more purification steps that were performed the better the observed separation and migration of compounds was (data not shown). Overall, the bioassayguided isolation would appear to be suitable to use in order to obtain the purified active compounds (Figure 3), as reported before [18], but the other fractions need to be screened as well, and the fractions processed to pure compounds for confirmation. However, considering the results presented here, it is possible that the antiproliferation activity in each or some of the fractions is derived from a combination of compounds. Certainly, Orsolic et al. reported a synergistic antitumor effect by the water-soluble derivatives of propolis from A. mellifera in Croatia [27]. In the future, the fractional inhibitory concentration index (FICI) method should be performed in order to determine the level, if any, of synergy and antagonism. In addition, the purification to homogeneity and analysis of the chemical structures of each bioactive component should be performed in order to investigate which exact compounds in Thai propolis are responsible for the antiproliferation, and to act as the template for future drug design.
Propolis of the stingless bee (Trigona laeviceps) from Thailand was tested for antiproliferative activity. Five cancer cell lines (Chago, Kato-III, SW620, BT474 and Hep-G2) and two normal cell lines (CH-liver cells and HS27 fibroblast cells) were selected for this purpose. The crude ethanolic extract displayed a good antiprolferative activity. For example, the IC 50 was 19.9 and 29.1 μg/ml for the SW620 cancer cell line and CH-liver cells, respectively. After solvent partitioning, the hexane fractions revealed the highest antiproliferative activity against the five cancer cell lines and the lowest cytotoxic activity on the normal cell lines. For example, IC 50 values of 16.4 and 32.4 μg/ml for the SW620 cancer cell line and CH-liver cells, respectively. The hexane fraction part was, therefore, purified in the next step by quick column and size exclusion chromatography. Considering the IC 50 , it was obvious that the more purified fractions were, the higher the antiproliferative activity was achieved. Two fractions that eluted with 30% and 100% (v/v) CH 2 Cl 2 in hexane (30DCM and 100DCM, respectively) are potential sources of new antiproliferative compounds. Thus, T. laeviceps propolis from Thailand contains some bioactive compounds that are not only effective in antiproliferative activity on cancer cell lines, but also nontoxic to normal cell lines. In the future, it is possible that the bioactive compounds could serve as a template for anticancer-drug developing program.
Figure 3 The decrease in the in vitro IC 50 values (μg/ml) on five cancer cell lines, plus the CH-liver and HS27 fibroblast normal cell lines as a comparative reference control, with increasing purification stages of the propolis extract from T. laeviceps. Data came from the mean ± 1 SD of percentage of cell viability which was derived from three replicates.
We thank the
The authors declare that they have no competing interests.
Authors' contributions SU collected the samples, prepared the extracts and carried out the experiments. PP contributed to the study design, analyzed and interpreted the data. SP carried out some of the experiments. CC designed and supervised the experiments and contributed in drafting the manuscript. All authors read the manuscript, contributed in correcting it and approved its final version.
The pre-publication history for this paper can be accessed here
We examined the differential associations of each parent's height and BMI with fetal growth, and examined the pattern of the associations through gestation. Data are from 557 term pregnancies in the Pune Maternal Nutrition Study. Size and conditional growth outcomes from 17 to 29 weeks to birth were derived from ultrasound and birth measures of head circumference, abdominal circumference, femur length and placental volume (at 17 weeks only). Parental height was positively associated with fetal head circumference and femur length. The associations with paternal height were detectible earlier in gestation (17-29 weeks) compared to the associations with maternal height. Fetuses of mothers with a higher BMI had a smaller mean head circumference at 17 weeks, but caught up to have larger head circumference at birth. Maternal but not paternal BMI, and paternal but not maternal height, were positively associated with placental volume. The opposing associations of placenta and fetal head growth with maternal BMI at 17 weeks could indicate prioritisation of early placental development, possibly as a strategy to facilitate growth in late gestation. This study has highlighted how the pattern of parental-fetal associations varies over gestation. Further follow-up will determine whether and how these variations in fetal/placental development relate to health in later life.
Low birth weight and size are related to the risk of cardiovascular disease and type 2 diabetes later in life [1,2]. Part of this relationship is thought to be explained by developmental programming, which suggests that environmental exposures occurring during developmental periods cause physiological adaptations that alter the long term propensity for disease [3,4]. In © 2010 Elsevier Ireland Ltd. ⁎ Corresponding author. MRC Epidemiology Resource Centre, University of Southampton, Southampton General Hospital, Southampton SO16 6YD, UK. Tel.: + 44 2380 777624. chdf@mrc.soton.ac.uk. This document was posted here by permission of the publisher. At the time of deposit, it included all changes made during peer review, copyediting, and publishing. The U.S. National Library of Medicine is responsible for all links within the document and for incorporating any publisher-supplied amendments or retractions issued subsequently. The published journal article, guaranteed to be such by Elsevier, is available for free, on ScienceDirect.
this sense, fetal growth is a marker of a disturbed prenatal environment. Investigating environmental and genetic determinants of fetal growth may enhance our understanding of the associations between birth size and later life health. And from a public health perspective, intrauterine effects due to modifiable maternal factors, such as diet or adiposity, are crucial as they provide additional opportunities for intervention. This has particular relevance in a developing country such as India where the prevalence of low birth weight is as high as 30% [5].
Maternal and paternal height and body mass index (BMI) capture information on the genetic potential of the offspring, the shared environment of the parents, and the historical and present nutritional condition of the mother. The latter is reflected onto the fetus' own nutritional status in what we refer to as intrauterine effects. There is evidence that parental anthropometry is strongly associated with birth size. For example, maternal BMI and height explain a large proportion of the geographical variation in birth weight, length and head circumference [6], and paternal height is associated with birth weight [7][8][9] and length [8], independent of maternal height. One way of disentangling intrauterine effects from the inherited component of size and growth is to compare the size of the association between the mother and offspring with the size of the association between the father and offspring [10,11]. If intrauterine effects are present, then we would expect the maternal associations to be stronger than the paternal association. Previous studies showed that the height of both parents correlate approximately equally with newborn length, while maternal BMI correlates more strongly than paternal BMI with measures of newborn soft tissue mass [6].
A limitation of birth anthropometry as a proxy for fetal growth is that two babies of a similar size at birth may have had radically different growth trajectories in-utero [12]. The absence of an association between a maternal exposure and birth size therefore, does not necessarily exclude an effect acting during a period of fetal development. Serial fetal ultrasound alongside data from birth should provide a better marker of growth restriction. Fetal growth is characterised by substantial centile crossing in individual growth trajectories suggesting it is a highly adaptive process [13]. Studies to date have tended to examine cross sectional associations at each time of measurement [14] or modeled the trajectory using multilevel methods [15,16]. An alternative which may capture this adaptive process more directly is to examine growth conditional on earlier size, this approach allows growth at different periods of gestation to be examined independently of earlier size and growth, and thus an examination of the onset and pattern of associations over the course of gestation.
Data from mother-father-fetus/offspring trios from the Pune Maternal Nutrition study were used to: [1] examine the association between parental height and BMI and fetal head circumference, abdominal circumference and femur length, [2] assess the relative influence from the mother and father, and [3] explore the timing and behavior of these associations over gestation. In an a posteriori analysis, we also examined the relationship with placental volume at 17 weeks gestation.
Data are from the Pune Maternal Nutrition Study (PMNS), a rural Indian community-based prospective study. The cohort was established from a house-to house survey of all married women of childbearing age (15-40 years) living in 6 villages located 40-50 km from Pune City (n = 2675). Enrolment took place between 1994 and 1996; 2466 women (92%) agreed to participate and 1102 became pregnant. The majority of women were vegetarian and had below recommended intakes of energy and protein. Further details are elsewhere [17].
2.2 Data collection 2.2.1 Parental data-Each woman's height and weight were measured by health workers every three months until pregnancy occurred; the last set of measurements prior to conception was used as pre-pregnant values. Paternal height and weight were also measured within 3 months of the confirmation of the woman's pregnancy. BMI was calculated as wt/ht 2 . Several potential confounding variables were examined. Information on household socioeconomic status was collected using a standardised questionnaire [18], this derives a composite score based on the occupation and education of the head of the household, caste, type of housing, and family ownership of animals, land and material possessions (5 level ordinal scale). Maternal parity and religion were recorded. None of the women were smokers, but it was recorded if there were any other smokers in the household (Yes/No).
Fetal ultrasonography and birth measures-Each woman was visited monthly by a trained health worker to record the date of the last menstrual period (LMP). Women who reported a missed period underwent an ultrasound examination 15-18 weeks after their LMP to confirm pregnancy. A further ultrasound scan was scheduled for 28 ± 2 weeks gestation. The median gestational age at each examination was 17.1 (IQR: 16.6, 18.0) and 29.4 (IQR: 28.6, 30.1) weeks.
On each visit, one of two trained sonologists (MCC and ASK) obtained measures of fetal head circumference (HC), biparietal diameter (BPD), abdominal circumference (AC) and femur length (FL). Scans were carried out using a portable machine fitted with a curvilinear array 5 MHz transducer (ALOKA SSD 500, version 8.1, Osaka, Japan). Placental volume was measured at the first ultrasound scan only, with a linear-array 3.5 MHz transducer with a foot length of 14 cm and a 12.5 cm field of vision, using a modified planimetric technique [19]. The inter-observer variation in the extracted fetal measurements was excellent (0.004-0.04%). Full technical details on the fetal measurements [20] and the placental measurements [21] have been reported elsewhere.
Health workers performed detailed anthropometry of the babies within 72 hours of birth using standardised techniques. Birthweight was measured using a Salter spring balance; crown-heel length was measured using a portable Pedobaby Babymeter (ETS J.M.B., Brussels, Belgium); occipito-frontal head circumference and abdominal circumference were measured using a fibre glass tape (CMS Instruments, London UK), the latter immediately above the umbilical cord insertion, in expiration.
For the purpose of this analysis, we used LMP dates to estimate gestation -this is important in a study of fetal growth, because ultrasound dating assigns a gestational age based on fetal size, thus erasing variation in growth during early pregnancy. However, to minimize errors from erroneous LMP dates, fetuses whose gestational age estimated by ultrasound differed by more than 2 weeks from the LMP estimate were excluded (n = 144). Sonographic gestational age was determined on the first visit, using a prediction equation based on fetal BPD, AC and FL [22].
From 1102 pregnancies, in addition to those excluded due to gestational dating discrepancies, 288 were excluded due to spontaneous abortions, fetal anomalies on ultrasound, multiple pregnancies, medical terminations or pregnancies detected later than 20 weeks gestation (Fig. 1). Babies born pre-term (< 37 weeks) were excluded because they may show different parent-offspring relationships due to underlying morbidity. One baby whose mother had gestational diabetes was also excluded. The final sample comprised 478 babies with complete parental size, ultrasound and birth data (Fig. 1).
Exposures and outcomes were standardised to a z-score so that we could compare the relative associations from the mother and father, and across fetal components. Parental BMI was log transformed before standardisation. Growth models were developed to describe the relationship between the fetal variables (HC, AC and FL) and gestational age using the method described by Royston (1995) [23]. These were used to compute z-scores at scan 1 and 2, and are reported elsewhere [13]. At birth, z-scores for head circumference, abdominal circumference and crown-heel length were estimated using ordinary regression accounting for gestational age and stratifying by sex, a similar method was used for placental volume at the 1st scan.
2.6 Analysis 2.6.1 Conditional growth variables-Our analysis strategy was to assess associations of parental height and BMI with fetal size at scan 1 (17 weeks), then determine the association with growth from scan 1 to scan 2 (17-29 weeks) independent of size at 17 weeks, and move forward again to assess associations with growth from 29 weeks to birth independent of size at 17 and 29 weeks. To do this, we created 2 conditional growth variables for each fetal component -size at 29 weeks conditional on size at 17 weeks (29| 17 wks), and size at birth conditional on size at 29 and 17 weeks (Birth| 29 and 17 wks), by regressing each size measurement (z-score) on earlier size (stratified by sex) and keeping the standardised residuals. The conditional z-score is a measure of growth velocity between 2 time points (e.g. 17-29 weeks) on a theoretical distribution which compares each fetus against other cohort members of the same size at time point 1 (e.g. 17 weeks). It can be interpreted as growth above or below that expected given earlier size. The set of variables (size at 17 weeks; 29| 17 wks; Birth| 29 and 17 wks) thus contain information on whether a fetus grew quickly or slowly for the intervals 0-17, 17-29 and 29-birth relative to other fetuses of the same growth history. Importantly, by construction, the set of conditional variables are uncorrelated, and so any association between an exposure and an interval is independent of its association with other intervals.
2.6.2 Statistical tests-The outcome variables were the growth set for HC, AC, FL and placental volume at 17 weeks. We had no data on FL or leg length at birth so used birth length as a proxy for birth femur length; these have been shown to be closely related [24]. We also estimated the association with overall size at 29 weeks and birth. Multivariable regression was used to examine the independent associations between parental height and BMI and the fetal outcomes. We pooled sexes as there was only one sex interaction out of 36. We adjusted for socio-economic status (SES), smoking and religion (Hindu/Buddhist, Muslim). Although there was little suggestion of confounding, we kept them in the models along with maternal age and parity because they increased the precision of the exposure estimates. Non-linear associations were assessed using plots and Wald tests of quadratic terms. We tested for a difference between the size of the mother-offspring association and the father-offspring association by reparameterising the model [25]. STATA v10 was used for all analyses.
2.6.3 Misattributed paternity-False paternity will dilute the paternal offspring relationship and bias the comparisons of the mother-offspring and father-offspring relationship in favour of showing larger maternal effects. We performed a sensitivity analysis to examine the effect of non-paternity using a previously described method [26]. This method corrects the estimated coefficients for the mother and putative father by modifying the covariance matrix under the assumption that the non-biological father's BMI is unrelated to the offspring's BMI but is related to the mother's BMI to a similar magnitude of the biological father's BMI. To try to include the true non-paternity rate we simulated rates up to 15%.
Mothers were short, light and thin (Table 1) -65% were underweight (BMI < 18.5 kg/m 2 , WHO guidelines). Most were younger than 22 years and one-third were primiparous. Fathers were also short and thin (Table 1) -39% were underweight. The correlations between maternal and paternal height and between maternal and paternal BMI were 0.19 and 0.13 respectively. The mean birth weight was 2597 g and 26% had a low birth weight (Table 1).
Maternal height was not related to fetal HC at 17 weeks (Fig. 2). There were positive associations with conditional growth between 17 and 29 wks, and between 29 wks and birth although the evidence was weak (p = 0.08 for both) (Fig. 2). At birth, an SD increase in maternal height (5 cm) was associated with a 0.09 SD increase in birth HC (Fig. 2).
There was no evidence for an association between paternal height and HC size at 17 weeks. Paternal height was positively related to growth in fetal HC between 17 and 29 weeks -an SD increase in the father's height (6 cm) was associated with a 0.11 SD increase in the part of HC growth from 17 to 29 weeks not associated with HC size at 17 weeks. There was no evidence for an additional effect of paternal height from 29 weeks beyond that which occurred earlier, however, paternal height was related to unconditional HC size at birth (Fig. 2).
There was no evidence of an association between maternal height and any AC growth or size parameters. There was a very weak suggestion of a positive association between paternal height and AC at 17 weeks. Paternal height was not related to AC in any of the other independent growth intervals or to AC size at 29 weeks or birth (Fig. 2).
There were no associations between maternal or paternal height and femur length (FL) at 17 weeks, or conditional growth from 17 to 29 weeks (Fig. 2). Both maternal and paternal height were positively related to length at birth conditional on earlier FL, and with overall length at birth. An SD increase in maternal and paternal height was associated with a 0.22 SD and 0.15 SD increase in the part of birth length not associated with femoral growth up to 29 weeks. Unlike maternal height, paternal height was related to unconditional FL size at 29 weeks (β (per SD) = 0.1SD, p = 0.024; data not shown in Fig. 2).
Despite a suggestion that the associations with paternal height were apparent earlier in gestation, there was no evidence for a difference between maternal-offspring and paternaloffspring relationships for any period of gestation or for any fetal component (Fig. 2).
Pre-pregnancy maternal BMI was negatively associated with HC size at 17 weeks (β (per SD) = -0.11 SD; 95% CI: -0.19, -0.02) and positively associated with conditional HC growth from 29 weeks to birth (β (per SD) = -0.09 SD; 95% CI: 0.0, 0.18) (Fig. 3). There was no evidence for an association between maternal BMI and fetal AC or FL size or growth (Fig. 3).
Paternal BMI did not appear to be related to any fetal growth or size component (Fig. 3), and despite there being some evidence for a maternal BMI-offspring relationship, tests of a difference between the parent's coefficients did not support differential parent effects.
Wills et al. Page 5 Published as: Early Hum Dev. 2010 September ; 86(9): 535-540.
Sponsored Document Sponsored Document
The median (IQR) placental volume at 17 weeks gestation was 148.6 ml (120.5, 179.2). Maternal BMI ((β (per SD increase) = 0.14 SD; p = 0.003) and paternal height ((β (per SD increase) = 0.1 SD; p = 0.02) were associated with placental volume at 17 weeks. There was no evidence for a relationship with maternal height or paternal BMI (Fig. 4). Placental volume was more strongly related to maternal BMI than paternal BMI (difference in β = 0.15, p = 0.034), there was no evidence that these associations differed for parental height (p for difference = 0.146).
Increasing the rate of non-paternity drew the putative father's association towards the mother's. For example, with non-paternity assumed at 15% the association between paternal height and growth in length from 29 weeks to birth changed from 0.13SD (per SD increase in father's height) to 0.16SD, the mother's coefficient reduced by only 0.004SD from 0.23SD. However, the effect of non-paternity was very small when the putative father's association was close to zero. Thus the patterns where there is a suggestion of an association for the mother but no clear evidence for the father, for example the associations with placental size, were unaffected by rates of non-paternity.
To the best of our knowledge, this is the first study to examine parental height and BMI associations with ultrasound measures of fetal growth in a rural Indian population with a high prevalence of low birth weight. Our results suggest that associations between paternal height and fetal head circumference and femur length are detectible earlier in gestation (17-29 weeks) compared to the same associations with maternal height. Fetuses of mothers with a higher BMI had smaller head circumferences at 17 weeks, but head growth caught up to become larger in late gestation. Paternal BMI showed no significant associations with any component of fetal growth. Maternal but not paternal BMI, was positively related to placental size at 17 weeks.
Our study has several important strengths. First, we used serial fetal biometry to characterise prenatal growth as opposed to birth weight. Second, ultrasound measurements were made by only 2 sonographers and intra and inter-observer reliability was excellent. Third, we were able to obtain accurate LMP dates since health workers visited women every month prior to pregnancy. Fourth, the maternal BMI data are likely to be an accurate representation of prepregnant BMI as data were collected within a 3 month window prior to pregnancy. Fifth, a sensitivity analysis showed that non-paternity was unlikely to have affected our main findings. And last, the use of conditional growth variables as a model of fetal growth allowed an exploration of the onset and pattern of parent-fetal associations.
The main limitation is a lack of power, a bigger study may have exposed some of the trends in the timing of associations over gestation that we reported, however intergenerational data with fetal ultrasound from populations in rural India are rare. We were also restricted to investigating growth intervals at relatively fixed time points. However, while more densely spaced measures may have offered a deeper insight, it does increase signal-noise ratio in the growth variable.
Maternal height was positively related to conditional growth in the last gestational interval (29 weeks -birth) for head circumference and femur/birth length. A study of American singletons also found positive associations between maternal height and head circumference at 31 weeks, and with femur length at 25 weeks [14]. Paternal height was positively related to conditional head circumference growth from 17 to 29 weeks and size of the femur at 29 weeks. The associations with father's height were thus evident earlier in gestation compared to the associations with maternal height. In an Australian study, paternal height was positively related to fetal femur length at 24 weeks, while maternal height was negatively related [27]. We could find no other studies to compare this differential pattern.
Positive correlations of maternal and paternal height with newborn skeletal measurements like length and head circumference are well described [6]. Our new finding that correlations between paternal height and fetal size and growth occurred earlier in gestation compared to associations with maternal height, may be related to the strong correlation between paternal height, but not maternal height, and placental volume at 17 weeks gestation. It suggests that paternally inherited genes influencing skeletal growth are expressed throughout gestation, while those from the mother are expressed in late gestation, and may be related to larger placental size. These differences could be mediated by epigenetic mechanisms; it is known that a number of genes that promote fetal growth are paternally imprinted [28,29]. There is insufficient data like ours to know if the associations observed are specific to our population or seen in other populations. Taken together, the maternal and paternal findings would be consistent with the 'selfish gene' theory, which suggests that paternally inherited genes promote fetal growth, regardless of maternal nutritional status [30].
For maternal BMI, there was a positive association with HC growth in the 29 week to birth interval, consistent with findings from Goldenberg et al. (1993) [14], who found positive associations with HC at 31, 36 weeks and birth. The associations with maternal BMI at 17 weeks were negative for all fetal components in our study. In particular, in this population with a high prevalence of chronic maternal undernutrition, women with lower BMI had fetuses with larger head circumference at 17 weeks. A comparison of our finding with the American study [14] is not possible because the coefficients were not reported, however, Blake et al.
[27] did report negative associations between femur length and both maternal and paternal BMI, although the evidence was weak. It is unlikely that our results are due to gestational dating errors related to maternal BMI, as there was no relationship between the discrepancy in ultrasound and LMP dated gestational age and maternal BMI (p = 0.6), unlike that found in another study [31]. Pre-pregnancy BMI is a marker of energy balance, nutrition and adiposity. Interestingly, maternal BMI was positively associated with placental volume at 17 weeks. An explanation for this differing placental-fetal response could be that better nutrition (higher maternal BMI) stimulates placental development at the expense of the fetus in early pregnancy, possibly as a strategy to enable greater fetal growth in late gestation. A larger placenta, whilst requiring more energy itself, has a larger surface area, which allows increased nutrient transfer to the fetus.
The different patterning of associations reflected by the mother and father's BMI on placental volume and the suggestion of a different pattern of associations on the fetal componentsparticularly head circumference, suggest some direct and potentially modifiable effects of maternal nutritional status on placental and fetal growth. Higher maternal BMI may promote placental growth in early-mid pregnancy, and this may enhance fetal growth in late pregnancy, although this interpretation of our findings is speculative. We do not yet know the implications of variations in placental and fetal growth for future health in the child. Recent studies in the Helsinki birth cohort suggest that placental size and shape at birth, newborn size, and ratios of placental to newborn size, predict adult hypertension, and that these associations are conditioned by maternal nutritional status [32]. In further follow-up of this cohort, we will be able to explore associations of placental measurements and fetal growth patterns with cardiometabolic risk factors in the children.
We recommend that epidemiological studies exploit the use of ultrasound to characterise variations in fetal and placental growth, and their timing during gestation. Such measures will allow more carefully designed analyses into questions related to the developmental origins of disease, and offer insight beyond that possible with birth weight alone. Wills et al. Page 10 Published as: Early Hum Dev. 2010 September ; 86(9): 535-540. Sponsored Document Sponsored Document Sponsored Document Wills et al. Page 12 Published as: Early Hum Dev. 2010 September ; 86(9): 535-540. Sponsored Document Sponsored Document Sponsored Document Wills et al. Page 13 Published as: Early Hum Dev. 2010 September ; 86(9): 535-540. Sponsored Document Sponsored Document Sponsored Document Sponsored Document Sponsored Document Sponsored Document Wills et al. Page 14
Table 1
Characteristics of the study sample.
n Mean (SD) Mother Age (years) a 557 21 (19, 23) Height (m) 557 1.52 (0.05) Weight (kg) 553 41.8 (5.1) BMI (kg/m 2 ) a 553 17.9 (16.7, 19.1) Nulliparous (n, %) 557 177 (31.8) Father Age (years) a 222 28 (25, 31) Height (cm) 532 1.65 (0.06) Weight (kg) 537 52.8 (7.9) BMI (kg m 2 ) a 532 19.0 (17.6, 20.7) Baby (at birth) Female (n, %) 557 259 (46.5) Weight (g) 519 2658 (361) Length (cm) 538 47.7 (2.0) Head circumference (cm) a 539 33 (33.2, 34.0) Abdominal circumference (cm) a 539 28.6 (27.5, 29.8) Gestational age (weeks) 557 39.5 (1.2) Religion (n, %) Hindu 557 538 (96.6) Muslim 16 (2.9) Buddhist 3 (0.5) Household Smokers (n, %) b 557 141 (25.3) a Median (IQR). b All smokers were male.
Published as: Early Hum Dev. 2010 September ; 86(9): 535-540.
Published as: Early Hum Dev. 2010 September ; 86(9): 535-540.Sponsored DocumentSponsored Document Sponsored Document
Published as: Early Hum Dev. 2010 September ; 86(9): 535-540.
We thank the women and families of the PMNS, and the late
Published as:
Sponsored Document Sponsored Document
In the global effort to eliminate lymphatic filariasis (LF), rapid field-applicable tests are useful tools that will allow on-site testing to be performed in remote places and the results to be obtained rapidly. Exclusive reliance on the few existing tests may jeopardize the progress of the LF elimination program, thus the introduction of other rapid tests would be useful to address this issue. Two new rapid immunochromatographic IgG4 cassette tests have been produced, namely WB rapid and panLF rapid, for detection of bancroftian filariasis and all three species of lymphatic filaria respectively. WB rapid was developed using BmSXP recombinant antigen, while PanLF rapid was developed using BmR1 and BmSXP recombinant antigens. A total of 165 WB rapid and 276 panLF rapid tests respectively were evaluated at USM and the rest were couriered to another university in Malaysia (98 WB rapid, 129 panLF rapid) and to universities in Indonesia (56 WB rapid, 62 panLF rapid), Japan (152 of each test) and India (18 of each test) where each of the tests underwent independent evaluations in a blinded manner. The average sensitivities of WB rapid and panLF rapid were found to be 97.6% (94%-100%) and 96.5% (94%-100%) respectively; while their average specificities were both 99.6% (99%-100%). Thus this study demonstrated that both the IgG4 rapid tests were highly sensitive and specific, and would be useful additional tests to facilitate the global drive to eliminate this disease.
Diagnostic tools are an essential component for the success of the Global Program for Elimination of Lymphatic Filariasis (GPELF). Thus far, the established diagnostic tests that are commercially available for bancroftian filariasis are two antigen detection tests namely NOW Filariasis Test [1] and Og4C3-ELISA (Trop Bio, Pty. Ltd., Australia); and for brugian filariasis is the Brugia Rapid test [2]. A laboratory-based Bm14-ELISA has also been extensively employed in studies in Egypt [3,4]. In addition
PCR-based assays for both brugian and bancroftian filariasis are also promising tools for the GPELF which can be employed for monitoring infections in both human and vector [5,6]. LF mainly affects the poor who reside in areas which are remote and/or without adequate health and laboratory facilities. Thus diagnostic tools in the format of rapid tests, particularly those based on immunochromatography technology, are most suitable to be employed for the GPELF, since they allow easy on-site testing, followed by rapid, simple reading and interpretation of results. These would avoid potential logistical challenges for sample storage and transportation, as well as more serious problems such as sample mix-up due to unclear/unreadable labels and sample degradation that may occur if collection and performance of tests are not conducted at the same or nearby locations. For such a major global program which needs to be sustained for a prolonged period, availability of a panel of rapid tests would help ensure smooth progress of the program and avoid potential problems such as supply interruption and changes/variations in test performance. Two new rapid immunochromatographic cassette tests based on detection of anti-filarial IgG4 antibody are now commercially available namely WB rapid and panLF rapid. The aim of this study is to perform a multicentre study to validate the sensitivities and specificities of the tests.
The test kits were acquired by the senior author from the manufacturer. A proportion of the tests were validated at USM, and the rest of the tests were couriered to the other four participating laboratories. The WB rapid test consists of two lines namely a test line and a control line, with the former comprising BmSXP recombinant antigen.
The panLF rapid test consists of three lines namely two test lines, one comprising BmSXP and the other BmR1 recombinant antigens; and a control line. Goat antimouse IgG antibody is employed as the control line for both tests. These lines are invisible in an unused test and are coloured red after performance of the test. Serum/ plasma and whole blood may be employed as test samples.
For serum samples, the test is performed by delivering 25 ul serum sample into the square bottom well. When the sample front reaches the blue line on the cassette window, two drops of buffer are added to a top oval well to release the conjugate solution (monoclonal anti-human IgG4 conjugated to colloidal gold). This is followed by pulling a plastic tab at the bottom of the cassette and adding a drop of buffer into the square bottom well, and by 15 minutes, the results can be read. For both tests, appearance of only the red control line denotes a negative result. For WB rapid test, a positive result is demonstrated when two red lines (a test and a control line) are seen. For panLF rapid test, a test is interpreted as positive when either three red lines (two test lines and a control line) or two red lines (a test and a control line) are observed.
Each participating institutions employed serum samples from their serum bank, which were obtained according to the ethical requirements of the respective organizations.
With regard to the samples tested in Japan, the sera from W. bancrofti patients were collected in Sri Lanka, while the normal sera were from Japanese. The tests were performed in a blinded manner and the results were collected from each centre by e-mail attachments.
Table 1 shows the number of the tests and the results obtained at each institution. WB rapid test displayed an average sensitivity of 97.6% (239/245), ranging from 94% to 100%. The average overall sensitivity of panLF rapid test was 96.5% (390/404), ranging from 94% to 100%; the sensitivity for detection of W. bancrofti infection was 96.0% (217/226) [94% to 100%] while that for detection of brugian filariasis was 97.2% (173/178) [92% to 100%]. The specificities of both tests were evaluated with serum samples from quite a large variety of other infections, which included helminthes, protozoan, bacterial and viral infections. The results showed that the tests were either 99% or 100% specific, with an average specificity of 99.6%. Thus, at all the evaluation centres, the sensitivites and specificities of both tests were consistently high.
The mf+ samples with circulating filarial antigen (CFA), as determined by Og4C3 assay in samples from Sri Lanka (n = 41) and India (n = 18), had CFA values > 512 and >1000 respectively. The Sri Lankan mf-samples had CFA values > 512; with 62/63 (98.4%) and 59/63 (96.7%) samples positive for WB rapid and panLF rapid respectively. In addition, the two rapid tests were also tested with samples from 22 amicrofilaraemic, CFA+ individuals (cryptic infections) from India which had a wider range of CFA values (100-4786). In general both rapid tests tested positive with sera which had CFA units greater than 200, and they tested negative with sera which had CFA units below this value. Since CFA may remain positive for sometime after death of adult worms, some of the mf-, low CFA+ individuals may no longer be actively infected. On the other hand these may also be individuals with reproductively immature worms.
In the pre-certification phase of the elimination program and in the surveillance activities post-elimination, a highly sensitive test, as displayed by an antibody-based diagnostic tool is essential since the level of infection, if any, is very low. Therefore, although a rapid antigen detection test is already available for bancroftian filariasis, an antibody detection assay would probably be more useful in the screening of young children as required in the precertification phase of GPELF. Antigen detection assays depend on presence of developmentally mature worms while antibody assays could potentially detect exposure to infective larvae by children. Thus WB rapid would be helpful to address this diagnostic requirement. For detection of all species of lymphatic filariasis, a rapid test such as panLF rapid would be very useful in several kinds of situations, namely testing in areas where there are mixed bancroftian and brugian filaria infections, in areas where the infecting species is not known or not confirmed, and for screening of immigrant workers in countries such as Malaysia which has more than 1.3 million workers from filarial endemic countries. These workers may pose a threat to the achievement of the disease elimination or they may be a source of resurgence of the disease in the future.
BmR1 is a recombinant antigen derived from Bm17DIII gene [GenBank: AF225296] and employed in a rapid test called Brugia Rapid. It has been shown to be highly sensitive (>95%) and specific (≥ 99%) for detection of B. malayi and B. timori infections in laboratory evaluations [7][8][9][10] and field studies [11][12][13][14]. In a field study in Malaysia which is a low endemic area, Brugia Rapid detected about ten times more positive cases than parasitological diagnosis, while in the high endemic area of Indonesia, the increase in detection was about three times [11,12]. Follow-up post-treatment studies of mirofilaraemic individuals showed that the titres of IgG4 antibodies to BmR1 decreased post-treatment. In Malaysia which is a low endemic area, it took approximately 6 months to 2 years post-treatment for the assay to become negative [15,16]. In a study involving a paediatric population in Kerala, ultrasonography ('filarial dance sign' or FDS) identified adult worms in 7 out of the 39 (18%) amicrofilaraemic children who were Brugia Rapid positive, thus providing definitive evidence that the rapid test detected active infection. This was comparable to the observation of FDS in 6 out of 32 (19%) microfilaraemic children [17].
BmSXP is a recombinant antigen derived from SXP1 gene [GenBank no: M98813], the clone was isolated from a B. malayi adult male worm cDNA library with sera of bancroftian filariasis patients [18]. A rapid flow-through IgG immunofiltration test using WbSXP recombinant antigen has been developed and a sensitivity of 91% (30/33) was recorded for detection of W. bancrofti infection [9].
In a recent study, BmSXP was found to be more sensitive (95%) in detecting W. bancrofti infection as compared to BmR1 (14%). On the other hand BmR1 was more sensitive than BmSXP in detecting B. malayi infection (98% and 84% respectively) [19]. Since BmR1 and BmSXP recombinant antigen cross-reacts with bancroftian and brugian filaria infection sera respectively, the panLF rapid test is not useful for species identification. However in the context of GPELF or for screening of foreign workers, this does not pose a problem. Cross-reactivities with Loa-loa and Onchocerca infection sera were observed with both rapid tests, thus they are not useful in areas co-endemic with these infections. However the tests may be employed in the vast lymphatic filariasis endemic areas in the world, particularly in Asia, which do not overlap with endemic areas for non-lymphatic filariasis. Since LF endemic areas are also prevalent for infections with soil-transmitted helminthes and intestinal protozoa, the high specificities shown by both rapid tests with respect to non-filarial infections would allow the tests to be used with high confidence in these areas.
In conclusion, the present multicentre evaluation study conducted in five institutions (located in four different countries) clearly demonstrated the high sensitivities and specificities of WB rapid and panLF rapid tests. Thus these tests should be employed in further field studies and would merit consideration as potential tools to assist in the GPELF.
*Other infections: ascariasis, trichuriasis, hookworm, strongyloidiasis, toxocariasis, toxoplasmosis, typhoid, cysticercosis, schistosomiasis, malaria, dengue, amoebiasis WB rapid: sensitivity : 97.6% (239/245); specificity : 99.6% (243/244) panLF rapid: overall sensitivity : 96.5% (390/404); sensitivity for Wb detection: 96.0% (217/226); sensitivity for Bm/Bt detection: 97.2% (173/178); specificity: 99.6% (232/233)
(page number not for citation purposes)
The funding for this study was provided by Malaysian govt. IRPA grant:
RN, with the assistance of RAR, developed WB rapid and panLF rapid tests
RN -conceive, design and supervise the study, participated in the evaluation at USM, analysed the results, wrote the first draft of the manuscript. RAR -performed the evaluation and participated in the analysis of the results at USM.
IM & KE-supervised and participated in the evaluation at Aichi Medical University, edited the manuscript. RB -supervised and participated in the evaluation at the Institute of Life Sciences, edited the manuscript. RM -supervised and participated in the evaluation at University of Malaya, edited the manuscript. ST -participated in the evaluation at University of Indonesia, edited the manuscript.
WMV -supervised and participated in the serum sample collection in Sri Lanka, edited the manuscript.
All authors read and approved the final manuscript
Polymorphisms within the open reading frame as well as the promoter and regulatory regions can influence the amount of CCR5 expressed on the cell surface and hence an individual's susceptibility to HIV-1. In this study we characterize CCR5 genes within the South African African (SAA) and Caucasian (SAC) populations by sequencing a 9.2 kb continuous region encompassing the CCR5 open reading frame (ORF), its two promoters and the 3′ untranslated region. Full length CCR5 sequences were obtained for 70 individuals (35 SAA and 35 SAC) and sequences were analyzed for the presence of single-nucleotide polymorphisms (SNPs), indels and intragenic haplotypes. A novel SNP (+258G/C) within the ORF leading to a non-synonomous amino acid (Trp → Cys) change was detected in one Caucasian individual. Results demonstrate a high degree of genetic variation: 68 SNP positions, four indels, as well as the Δ32 deletion mutant, were detected. Seven complex putative haplotypes spanning the length of the sequenced region have been identified. These haplotypes appear to be extensions of haplotypes previously described within CCR5. Two haplotypes, SAA-HHE and SAC-HHE were found in high frequency in the SAA and SAC population groups studied (20.0% and 18.6%, respectively) and share four SNP positions suggesting an evolutionary link between the two haplotypes. Only one of the identified haplotypes, SAA/C-HHC, is common to both study populations but the haplotype frequency differs markedly between the two groups (8.6% in SAA and 52.9% in SAC). The two population groups show differences in both haplotype arrangement as well as SNP profile.
The human coreceptor, CCR5, acts as the principal coreceptor required for macrophage-tropic (R5) human immunodeficiency virus type 1 (HIV-1) virions to gain entry to a cell (Deng et al., 1996;Dragic et al., 1996). Shortly after the role of CCR5 was discovered, the CCR5Δ32 © 2010 Elsevier B.V. This document may be redistributed and reused, subject to certain conditions.
mutant and its association with protection to HIV-1 infection in individuals homozygous for this allele, was found (Samson et al., 1996). This discovery provided the first genetic evidence of protection to HIV-1 infection and prompted further studies of the gene and how its naturally occurring mutations may influence the outcome of HIV-1 exposure and infection.
The CCR5 gene is composed of four exons and two introns, where exons 2A and 2B are not interrupted by an intron (Mummidi et al., 1997). Exon 3 contains an intronless open reading frame (ORF). Two CCR5 promoters have been described, a weak upstream promoter (P U or P2) and a stronger downstream promoter (P D or P1) (Mummidi et al., 1997). Cell surface expression of CCR5 is highly variable, even in individuals homozygous for the wild type ORF region. This may be explained by differences in the promoter region of CCR5 which result in differing CCR5 expression levels (Menten et al., 2002). To date, several mutations and singlenucleotide polymorphisms (SNPs) in the HIV-1 coreceptor gene, CCR5, have been found to be important genetic factors capable of influencing susceptibility to HIV-1 infection or affecting the rate of disease progression.
Striking ethnic or population differences in SNP frequencies of CCR5 exist. The most studied polymorphism exhibiting this is the CCR5Δ32 mutant. The CCR5Δ32 allele occurs at a variable frequency of 4-15% in Caucasian populations, with an average of 10% in Europe (reviewed in Galvani and Novembre, 2005) and yet is rarely found in Asian or African populations. In a South African context, Petersen et al. (2001) identified seven novel mutations within the CCR5 ORF in African and Coloured populations, however polymorphisms within the promoter region have not been studied. Thus, a detailed descriptive study of SNPs and/or haplotypes within the CCR5 receptor gene was carried out in two South African populations, South African Africans (SAA) and South African Caucasians (SAC), with the aim of providing a baseline study of the prevalence of polymorphisms that exist in CCR5 within these populations and to determine whether all previously defined haplotypes are represented within these populations.
Characterization of the CCR5 gene was carried out on 70 healthy, HIV-1 uninfected adult volunteers, 35 were SAA and 35 were SAC. This study was approved by the University of the Witwatersrand Committee for Research on Human Subjects, and informed written consent was obtained from all participants.
Genomic DNA was extracted from blood samples anticoagulated with ethylenediaminetetraacetic acid (EDTA) using QIAamp DNA Mini Kit (QIAGEN, Dusseldorf, Germany). A ∼9.2 kb continuous region encompassing the CCR5 open reading frame (ORF), its two promoters and the 3′ untranslated region (UTR) was polymerase chain reaction (PCR) amplified in five overlapping sections using Expand High Fidelity PCR System (Roche, Mannheim, Germany). PCR and sequencing primers were designed using PRIMER DESIGNER for Windows (v. 2.0) (Supplementary Table S1) using the published sequences for CCR5 (GenBank accession: U95626, AF017632 (Moriuchi et al., 1997), AF031236 and AF031237 (Mummidi et al., 1997)) as reference sequences.
All sequencing reactions were carried out using BigDye Terminator version 3.1 chemistry (Applied Biosystems, Foster City, CA, USA). Amplified fragments were sequenced using the automated 3100 Genetic Analyzer (Applied Biosystems).
Sequence data was assembled and analyzed for the presence of SNPs and indels using SEQUENCHER software version 4.5 (Gene Codes Corporation, Ann Arbor, MI, USA). Assembled sequences were aligned with each other and the published GenBank sequence, U95626, using SEQUENCHER, to identify polymorphisms. The GenBank NCBI SNP database (dbSNP) was searched for all reported SNPs in the CCR5 gene to determine whether polymorphisms detected in this study had been previously reported.
The CCR5 numbering system used in this study is as described by Mummidi et al. (2000) where the first nucleotide of the translational start site is designated as +1 and the nucleotide immediately upstream from that is -1. A composite of the reference sequences AF031236 and AF031237 (Mummidi et al., 1997) was used as a basis for determining SNP positions as these sequences appeared to be closer to the wild type (WT) or more 'ancestral' gene. It must be noted however that when all Homo sapiens reference sequences used in this study were aligned, a number of differences between them were noted, including base insertions or deletions (indels), which would affect the SNP position values. Also, AF031236 and AF031237 do not encompass the entire region sequenced in our study. Thus, using the sequences flanking the various SNPs may be a more reliable means of identifying SNP positions (Supplementary Table S2).
Once polymorphisms within CCR5 were identified, it was necessary to determine which nucleotide to deem as the WT nucleotide. Generally, the most prevalent nucleotide in our combined populations was considered to be the WT or ancestral nucleotide/allele. In addition, to identify the WT nucleotide where it was not apparent which nucleotide/allele was most prevalent, the human CCR5 sequences were aligned with those of the chimpanzee found on the sequence available for Pan troglodytes chromosome 3 (GenBank accession number: NW_001232822.1) and the Pan troglodytes CCR5 sequence (GenBank accession numbers: NM_001009046 and AF005663).
Analysis of the sequence data generated for CCR5 revealed certain obvious patterns wherein the presence of a polymorphism at one position was consistently associated with polymorphisms at one or more other positions. These associations were identified as putative intragene haplotypes. The HAPLOTYPER software which uses a Bayesian algorithm for haplotypes inference (Niu et al., 2002) was also used to infer haplotypes for CCR5.
The frequencies of putative intragenic haplotypes were calculated by counting the number of alleles harbouring the haplotypes and dividing by the total number of alleles. Counting of the haplotypes was irrespective of the presence of additional SNPs not forming part of the haplotypes in question.
Five SAA and one SAC individual appeared to have a previously unidentified indel downstream from the open reading frame (+2772). Characterization of the putative indel as well as verification of Intron 2 indels was carried out by TA cloning of PCR amplicons into the pCR ® 4-TOPO ® cloning vector using the TOPO-TA Cloning Kit (Invitrogen, Carlsbad, CA, USA). Recombinant plasmids were screened for the allele with the putative indel by sequencing. Sequences were aligned with reference sequences, using SEQUENCHER, in order to characterize the indel.
All polymorphic loci detected within the characterized CCR5 gene region were tested for deviation from Hardy-Weinberg equilibrium using the conventional Monte Carlo exact test of Guo and Thompson (1992) implemented through the computer program TFPGA (Miller, 1997). The two population groups were tested independently.
To test whether the SNPs forming part of the putative intragenic haplotypes were in complete or strong linkage, linkage disequilibrium between every two SNP combination in each haplotype was estimated using the method described by Lewontin (1964) where the linkage disequilibrium coefficient D was calculated (D ij = HF ijp i p j ). D was subsequently normalized (D′) or standardized by the maximum value it can take (D max ) using the formula D′ ij = D ij / D max where HF ij is the frequency of the haplotypes carrying SNPs i and j, p i and p j are the frequencies of SNPs i and j, respectively and D max is either min [p i p j , (1
D′ values are defined in the range [-1, 1] with a value of '1' representing perfect disequilibrium. The statistical significance of the linkage disequilibrium between each of the SNP pairs was evaluated by the approximate chi-square described by Liau et al. (1984).
Fisher exact tests were performed using the Simple Interactive Statistical Analysis software (Uitenbroek, D. G., Binomial. SISA. 1997. http://www.quantitativeskills.com/sisa/distributions/binomial.htm. [1 January 2004]) to test whether there was any significant difference in SNP frequencies between this and other studies.
Assembled sequences of the CCR5 gene including promoter, coding and 3′ UTR regions, from 70 HIV-1 uninfected individuals were analyzed for DNA polymorphisms, SNPs and indels. Across the entire 9.2 kb region sequenced, 68 SNPs were identified. The positions and nucleotide (nt) changes are indicated in Fig. 1. The identified polymorphisms were found across the entire sequenced region with the exclusion of exon 2B, a small region spanning 54 nucleotides. The majority of the polymorphisms were located in the intron and UTR of the gene and only six were located in the ORF (Fig. 1). With regards to the two study populations, 60 and 37 polymorphisms were found in the SAA and SAC populations, respectively. Of the 68 identified SNPs, 46 have been previously described in the GenBank dbSNP database and by Petersen et al. (2001). Their corresponding accession numbers, where available, are shown in Table 1. To the best of our knowledge, with comparison to the GenBank dbSNP database and literature reports, 24 polymorphisms are newly identified and have been designated as newly identified (NI) in Table 1. These NI polymorphisms were found in both population groups. Newly identified polymorphisms are also distributed across the entire gene although the majority are located in the 5′ and 3′ UTRs. Most NI SNPs were found to be rare polymorphisms present in only one individual. Exceptions to this were the NI polymorphisms, -4223C/T, -3886C/T, -2454G/A and -451C/T, which were detected in higher numbers (three or four individuals each, all of which were heterozygous for the polymorphisms) in either/both populations.
The alignment of reference sequences (GenBank accession numbers: U95626, AF031236, AF031237, AF017632 and NT_022517.17) did not demonstrate 100% homology at many of the SNP positions detected, as well as at other potential polymorphic positions not detected in this study. Also, it was not always apparent which the most predominant base at certain positions was. For instance, with the -2554G/T SNP, 79 (56.4%) and 61 (43.6%) alleles in this study contained a G and T nucleotide, respectively. Although the G allele was more frequent overall, in the SAC population the major nucleotide was a T (55.7%) and in SAA individuals the major alleles was a G (68.6%). Caution was necessary in the selection of a WT nucleotide as the population sizes used in this study were of a size where bias could be introduced and the apparent WT allele (most frequent) may not correspond to the ancestral allele. Thus, reference Homo sapiens sequences were aligned with Pan troglodytes sequences. Due to the low mutation rate since the human-chimpanzee divergence, the human allele almost always corresponds to the allele present in chimpanzees (Cargill et al., 1999). Where an allele which was obviously the major allele in Homo sapiens (this study and sequences used as references) was not in agreement with that of the Pan troglodytes sequences at the same position, the former was selected as WT. This was the case for the indels, CTAT/-, AG/-and ACAA/G where published sequences for Pan troglodytes, some of the Homo sapiens reference sequences, as well as data for SAC would indicate the minor allele contained nucleotide insertions at these indel positions, yet overall the most frequent alleles detected in this study indicated that the minor allele for all 3 indel positions was the allele containing nucleotide deletions at those positions. The same was observed for SNPs -2852A/G and -113G/T, where the Pan troglodytes base at that position was the equivalent of the minor allele/nucleotide in humans. The minor allele frequencies of the CCR5 SNPs and indels in both the study populations are shown in Table 1.
Of the polymorphisms detected within the CCR5 ORF, one has not been reported previously (Table 1). This novel SNP (+258G/C) within the ORF leads to a non-synonymous amino acid change (Trp → Cys) at codon 86 and was detected in one SAC individual heterozygous for that mutation. Four mutations in the open reading frame were detected in the SAA population: +225T/C; +319C/T; +673C/T and +1004C/T (S75S; L107F; R225X and A335 V, respectively). One individual was found to be heterozygous for previously described mutations at both codon 107 and 225 (Petersen et al., 2001). In a study conducted by Petersen et al. (2001), these two mutations were reported as occurring simultaneously in an individual and at low frequencies in SAA and South African Coloureds. The codon 335 amino acid substitution mutation was detected within the SAA population at a frequency of 0.071 (n = 5) and only one individual in the SAC population was found to harbour this mutation (Table 1). Previous reports looking at African American (Ansari-Lari et al., 1997;Carrington et al., 1997;Carrington et al., 1999) and SAA (Petersen et al., 2001) populations found the mutation present at a frequency of approximately 3% and 2%, respectively. Although representation of this SNP appears higher (7.1%) in our study, this did not differ statistically from frequencies reported in these studies (P > 0.05).
Several polymorphisms were found to be restricted to either the SAA or the SAC population group (frequencies highlighted in grey in Table 1). The SNPs, -3894T/C, -3261G/A, -2132C/ T, -1686A/C, -1464A/G, -113G/T and +1752G/A, are all restricted to the SAA population at a frequency of >18% and all form part of the putative haplotypes, SAA-HHA and SAA-HHD, identified in this study. In addition, the SNPs, -4745C/T, +1843G/A and +1846G/A, were detected at a reasonably high frequency of 12.9% in SAA individuals but not in SAC individuals. There were no SNPs found exclusively in the SAC population which were also present at relatively high frequencies (i.e., at a frequency >10%) in that population. Where the SNP frequency is low in one population, absence in the other population cannot be used to state that that particular SNP is only prevalent in one population due to the sample size used in this study.
Five indels were detected across the entire CCR5 gene (Fig. 1). Four of the five detected indels have been previously described (Dean et al., 1996;Samson et al., 1996;Mummidi et al., 1997). The Δ32 deletion indel was found in 5/35 (14.3%) SAC individuals, all of which were heterozygous for this mutation, but not in SAA individuals. In a previous South African study, the CCR5Δ32 allele was detected at a frequency of 9.4% and 0.1% in SAC and SAA individuals, respectively (Williamson et al., 2000). Comparison of CCR5Δ32 allele frequencies observed in SAC populations in the two studies showed no significant difference between them (P = 0.68). The other three previously reported deletion indels (CTAT/-, AG/-and ACAA/G) appear to be in very strong linkage disequilibrium (D′ = 1.0, P < 0.0005 for all three indel associations in both population groups).
Two indels are located within Intron 2 of the CCR5 gene. The indel located at position -362 has been reported differently. In a report characterizing the CCR5 gene this is shown to be a CAA indel (Mummidi et al., 1997). Within the dbSNP database there are two polymorphism reports for that location: accession number rs41515644 reports an A/G SNP at that position, whereas accession number rs71615644 reports that the four nucleotide sequence, ACAA, is substituted with a single guanine nucleotide. Within our study group, only the latter polymorphism was observed. The Intron 2 PCR amplicons from two individuals heterozygous for the ACAA indel were cloned and sequenced. Sequencing demonstrated that on the alleles where there was a ACAA deletion, there would be a G substitution at that point. Thus, in our report, the deletion and base substitution have been treated as a single polymorphism. It is possible though that the ACAA/G indel may have arisen as two separate events which became evolutionarily linked, i.e. an A to G substitution and a CAA deletion immediately downstream from the substitution.
The indel downstream from the open reading frame (Fig. 1) consists of a single guanine insertion. Alleles containing the indel have a string of nine guanine bases in that region, whereas the WT alleles have eight guanine nucleotides. The exact position of the single base insertion within the eight consecutive guanine bases of the WT sequence cannot be precisely determined. Thus, the position of +2772, at the end of the eight guanine bases has been selected.
Individuals within the SAA and SAC populations were assigned to previously described haplogroups (Gonzalez et al., 1999) based on SNPs at positions -2733, -2554, -2459, -2135, -2132, -2086 and -1835 as well as the presence of CCR5Δ32 (Fig. 2A). One SAA individual was found to be heterozygous for a haplotype allele which could not be classed into any of the haplotypes defined by Martin et al. (1998) or Gonzalez et al. (1999). Similar trends in haplotype frequency to that reported in a larger study conducted by Gonzalez et al. (1999Gonzalez et al. ( , 2001) ) were observed. In SAC, HHA appears to be underrepresented (4.3% vs. 10% reported in Caucasians (Gonzalez et al., 2001)) and HHC as overrepresented (55.7% vs. 35% in Caucasians (Gonzalez et al., 2001)) and in SAA, HHF appears underrepresented (15.7% vs. 24% reported in African non-pygmies (Gonzalez et al., 2001)) and HHG*1 overrepresented at 4.3% (2% reported in African non-pygmies (Gonzalez et al., 2001)). Fisher exact test shows significant difference in HHC frequencies (P = 0.005) but no significant difference in the HHF, HHG*1 and HHA frequencies between the two studies (P > 0.05). The overrepresentation of HHC haplotype frequency in the SAC population in this study could potentially be attributed to differences in Caucasian population ancestry in the two studies. In the larger Gonzalez et al. (2001) study, their Caucasian study group (n = 959) is comprised of HIV-1 uninfected individuals from Finland, France and Poland and European American individuals of mixed infection status (i.e. both HIV-positive and HIV-negative individuals) (Gonzalez et al., 2001). Another possible explanation would be the presence of a greater amount of admixture within the Caucasian study group in the Gonzalez et al. (2001) study. These previously defined haplotypes however are located in a relatively small region of the CCR5 gene (898 bp of the regulatory region of CCR5, in addition to presence/absence of CCR5Δ32 in ORF and CCR2 V64I upstream on the same chromosome). This study has identified putative haplotypes which extend over the entire gene in both directions.
Seven complex putative haplotypes spanning the length of the sequenced region have been identified (Fig. 2B). These haplotypes appear to be extensions of haplotypes previously described within CCR5 (Gonzalez et al., 1999) (HHA, HHC, HHD, HHE and HHG*2). Haplotypes were named by prefixing the root haplotype name with the population within which it was found. Thus, a distinction can be made when haplotypes with the same root differ in SNP composition between study population groups (e.g. SAA-HHE and SAC-HHE which are both rooted on the HHE haplotype but differ between SAA and SAC individuals bearing the HHE haplotype). Where haplotypes were found to be identical in both study populations, the prefix SAA/C-was used. All the haplotypes described in this study occurred at a frequency greater than 5% and one was found at a frequency of 52.9% (SAA/C-HHC in SAC individuals). Five predominant putative haplotypes were identified in the SAA population whereas only three were identified in the SAC population (Fig. 2B). Only one haplotype appears to be shared by both study populations. This haplotype, SAA/C-HHC, is comprised of three indels and eight SNPs and is the most frequent haplotype in SAC individuals (52.9%) and the least frequent in SAA individuals (8.6%) (Fig. 2B).
Linkage disequilibrium analysis between every two SNP combination in haplotypes identified in this study demonstrated strong linkage disequilibrium between SNPs with a statistical significance greater than 95%.
All HHA, HHF, HHG*2 and HHD haplotypes were found to be associated with further SNPs, forming haplotypes SAA-HAA, SAA-HHF, SAC-HHG*2 and SAA-HHD, respectively (Fig. 2B). The majority of HHC and HHE alleles are associated with further SNPs forming SAA/ C-HHC, SAC-HHE and SAA-HHE. SAC-HHE and SAA-HHE are two putative haplotypes rooted on the HHE haplotype but differ at the inclusion of an additional SNP (-4358A/G) in SAC-HHE. The corresponding haplotypes occur at similar frequencies in the two populations (SAC-HHE: 18.6% in SAC and SAA-HHE: 20% in SAA).
All 13 alleles (SAA) classed as HHD can also be classed as SAA-HHD. Hence, a further three polymorphisms -4088T/C; -3894T/C and -3261G/A) can be said to be associated with the HHD haplotype (D′ = 1.0, P < 0.0005, for all SNP associations within the haplotype). No HHD haplotypes were found in the SAC population. HHF haplotypes appear to be linked to the +2919T/G SNP forming SAA-HHF.
In the SAC population 37/39 alleles classed as haplotype HHC, could also be classed as the HHC extended haplotype, SAA/C-HHC, by far the most predominant haplotype in that population (Fig. 2B). The remaining 2 HHC-bearing alleles occurred in two individuals homozygous for ten of the eleven polymorphism sites comprising SAA/C-HHC and were heterozygous for the +2077T/G SNP. Six out of eight HHC alleles in the SAA population exhibit the SAA/C-HHC polymorphism pattern, one allele lacks the SNP at position +2077 and the other lacks the SNP at -3458.
The HHA haplotype, which comprises WT bases at all SNPs positions used in the Gonzalez et al. (1999) classification system, is present at a frequency of 24.3% and 4.3% in the SAA and SAC populations, respectively (Fig. 2A). Although this may appear to imply that individuals harbouring the HHA haplotype are WT across the entire CCR5 gene, this is not the case as the HHA haplotype appears to be associated with different SNPs in the extended haplotypes in the different populations (Fig. 2B). In SAA individuals, HHA is associated with SNPs: -1686A/
Picton et al. Page 7 Published as: Infect Genet Evol. 2010 May ; 10(4): 487-494.
Sponsored Document Sponsored Document C, -1464A/G; -113G/T and +1752G/A forming SAA-HHA, whereas in SAC HHA alleles demonstrate no association with those SNPs but are instead linked in a haplotype to C/T SNPs at positions -3886; -1060 and +1823, all of which are polymorphisms not detected within the SAA study group. It must be noted, however, that in the SAC population this is a rare haplotype (4.3%) and so has not been shown as one of the predominant haplotypes in Fig. 2B and caution must be taken in assuming this is a true association.
The haplotype HHG can be subdivided into HHG*1 (alleles not containing Δ32 deletion in ORF) and HHG*2 (alleles containing Δ32 deletion in ORF). All HHG*2 haplotypes detected in this study (n = 5) were found to contain an additional three SNPs (-5268G/A, -4257A/C and +2919T/G), forming SAA-HHG*2 (Fig. 2B). The two SAC-HHG*1 alleles were identical to SAA-HHG*2 with one individual lacking the +2919T/G polymorphism. In contrast, the three SAA-HHG*1 alleles only had the additional polymorphisms, -5268G/A and -2852A/ G, in common with the SAA-HHG*2 haplotype.
No significant deviations from Hardy-Weinberg equilibrium were noted for any of the indels or SNP loci detected in this study in both the SAA and SAC population groups.
In this study we have characterized polymorphisms (SNPs and indels) and intragenic haplotypes found within CCR5 for two South African populations, SAA and SAC. This provides a baseline study for the CCR5 polymorphism and haplotype profiles within these two populations. Previously unreported polymorphisms have been identified and previously defined haplotypes within the CCR5 gene have been expanded upon.
There exists greater genetic diversity and low levels of linkage disequilibrium within African populations in comparison to European-originating populations (Tishkoff and Verrelli, 2003;Tishkoff et al., 2009). Hence, it is not unexpected to have found a greater number of polymorphisms in SAA in comparison to SAC individuals, as also observed in a recent study reported by Paximadis et al. (2009). Full length sequencing of the CCR5 gene allows for identification of SNPs which would normally not be detected. Although a number of NI SNPs were identified in our study population, these may have been missed in previous studies which look at a smaller portion of the gene or which use other means of identifying specific polymorphisms. Also, most of the NI polymorphisms identified in this study were detected exclusively in the SAA population. Owing to the sample sizes in this study, detection of a SNP in only one of the two study population groups cannot be used to conclude that that SNP is absent in the other population, but it can be used as an indication of overall prevalence and diversity.
At SNP positions where the major allele (WT) in our study and that of other human reference sequences differed from that of the chimpanzee sequence, the chimpanzee sequence was found to correspond to the minor allele in humans (CTAT/-, AG/-and ACAA/G indels as well as -2852A/G and -113G/T SNPs). In a study characterizing the SNPs in 106 human genes, Cargill et al. (1999) noted that in a significant fraction of cases, the minor chimpanzee allele had become the major human allele and hence the minor human allele was in fact the older allele. This has also been observed in a study reporting variants in the CCL3 and CCL3L genes which code for CCR5 ligands (Paximadis et al., 2009).
The codon 335 mutation resulting in an alanine to valine (A335V) substitution was previously reported by Ansari-Lari et al. (1997) and has been found to be present at a higher frequency in African American populations in comparison to Caucasians (Ansari-Lari et al., 1997;Carrington et al., 1997;Carrington et al., 1999). Although this mutation occurs in the ORF and could be thought to potentially affect protein structure and/or function, in a disease association study, this mutation has been found to have no effect on the rate of progression to AIDS (Carrington et al., 1997). In a study conducted in South African populations, the A335V mutation was detected in African and Coloured populations but not in Caucasians (Petersen et al., 2001). Our study indicates that this mutation is in fact present in the SAC population but as a very rare polymorphism (only 1/35 individuals harboured this allele). In SAA individuals, this mutation was detected at a much higher frequency than that reported elsewhere (Ansari-Lari et al., 1997;Carrington et al., 1997Carrington et al., , 1999;;Petersen et al., 2001). However, comparison of A335V mutation frequencies in SAA individuals from this study to that of healthy SAA individuals in another South African study (Petersen et al., 2001) indicated no significant difference between observed frequencies (7.1% vs. 1.6%; P = 0.099). This and the Petersen et al. (2001) study were conducted in two widely separated geographical regions within South Africa, the Gauteng and Western Cape provinces, respectively.
A novel non-synonymous mutation has been detected within the ORF at codon 86. It is unclear whether this tryptophan to cysteine amino acid substitution will have an impact on chemokinereceptor function. Amino acid alignment of chemokine receptors, CCR5, CCR2B, CCR1, CCR3, CCR4 and CXCR4 (Carrington et al., 1997) shows that this mutation occurs within a highly conserved region (second transmembrane region) between the receptors at a point where all aligned proteins contain a tryptophan residue. This high level of conservation implies that this region is important to the structure or function of the protein and hence indicates that the significance of this novel mutation warrants further study. Both tryptophan and cysteine residues are hydrophobic molecules but cysteine is considerably smaller than tryptophan. Also, tryptophan residues positioned near lipid bilayers, as with residue 86, tend to form hydrogen bonds with the lipid head groups (Schiffer et al., 1992) whereas cysteine residues are likely to form disulphide bonds. Thus, it is possible that this amino acid change may have an impact on the folding of the peptide chain and hence its function as a receptor.
Several SNPs located in the CCR5 promoter have been previously reported to affect the expression of CCR5. One such polymorphism is the -2459G/A polymorphism located within the downstream promoter (P1). This polymorphism has been linked to differences in CCR5 expression levels on CD14 + monocytes (Salkowitz et al., 2003) and has known association with the rate of progression to AIDS (McDermott et al., 1998). Individuals homozogous for the -2459G allele exhibit lower CCR5 receptor density in CD14 + monocytes (Salkowitz et al., 2003) and have been linked to slower disease progression (McDermott et al., 1998;Clegg et al., 2000;Knudsen et al., 2001). Both WT and mutant alleles have been found to be present at high frequencies in all racial groups, with reported frequencies of 43% and 57% in African and Caucasian populations, respectively (McDermott et al., 1998). In this study, the -2459A allele was detected at a frequency of 42.9% and 40.0% in SAA and SAC population groups, respectively. Fisher exact analysis of the frequencies observed in the two studies has shown no statistical difference between the African population (P = 1.0) but there was a statistical difference between the Caucasian populations (P = 0.0083). This is not unexpected as the WT alleles, -2459G and -2135T, which are in very strong linkage disequilibrium with each other (Gonzalez et al., 1999;Clegg et al., 2000), form part of the SAA/C-HHC haplotype which was present at a higher than expected frequency in the SAC population. Thus, it follows that the minor/mutant alleles at those positions will also be underrepresented. These two SNPs form part of the HHE, HHF and HHG haplotypes defined by Gonzalez et al. (1999) and the predominant putative haplotypes, SAC-HHE and SAC-HHG*2 described in the SAC population.
Although studies looking at individual polymorphisms on susceptibility to HIV-1 infection and the rate at which individuals progress to AIDS, do provide useful information, it is not always possible to pinpoint the cause of the observed effect of a particular polymorphism when studying the regulatory region of the gene as different combinations of polymorphisms may be in linkage disequilibrium forming haplotypes. Thus, it is important to look at haplotypes and their prevalence across the breadth of a gene.
When examining the effects of CCR5 haplotypes on HIV-1 disease in different population groups, Gonzalez et al. (1999) observed that haplotype diversity is greatest in African populations. This was also reflected in our study where five major (frequency >5%) haplotypes were observed in SAA individuals, whereas only three were found in SAC individuals.
In this report there are a number of polymorphisms where the minor allele in the SAA population has been shown to be the predominant or 'major' allele in the SAC population. This is most evident with the SNPs comprising the SAA/C-HHC haplotype. While SAA/C-HHC, and hence its associated polymorphisms, is by far the most prevalent haplotype detected in the SAC population (52.9%), in the SAA population this haplotype is much less prevalent (8.6%). The most prevalent haplotype in the SAA population is SAA-HHA which is an extension of the HHA haplotype reported to be the ancestral CCR5 haplotype (Mummidi et al., 2000). This is likely to be due to different evolutionary pressures being exerted on the two populations or a genetic bottleneck where a significant number of members of one the populations was unable to reproduce.
Previous reports have defined haplotypes within the CCR5 gene (Martin et al., 1998;Gonzalez et al., 1999). Martin et al. (1998) described 10 haplotypes, CCR5P1-P10, comprising 10 SNP positions within the region starting at Exon 1 and ending in Exon 2B of the gene. The nine haplotypes described by Gonzalez et al. (1999) comprise seven SNP positions within a similar region but extending slightly into Intron 2, the presence/absence of CCR5Δ32 in addition to the CCR2-64I mutation. In this study we have expanded upon this and have linked previously defined haplotypes to SNPs both upstream and downstream of these regions forming haplotypes which extend over a larger region of the gene with SNPs linked in a haplotype being as much as ∼8.1 kb apart (SNP positions -5248 and +2919 in Hap-Δ32).
A better understanding of the role played by host genes in response to human immunodeficiency virus (HIV-1) exposure will contribute towards a better understanding of the protective immunity to HIV-1 and of the disease process in HIV-1-infected individuals. CCR5 is increasingly being shown to play a critical and central role in HIV-1 infection and to date a number of genetic mutations within the gene have been found to positively or negatively influence an individual's susceptibility and rate of disease progression. Thus, studies such as these which provide valuable new information regarding the genetic diversity within this gene, are important to the further understanding of the impact of CCR5 expression on host susceptibility to HIV-1. Picton et al. Page 13 Published as: Infect Genet Evol. 2010 May ; 10(4): 487-494. Sponsored Document Sponsored Document Sponsored Document Picton et al. Page 14 Published as: Infect Genet Evol. 2010 May ; 10(4): 487-494. Sponsored Document Sponsored Document Sponsored Document Sponsored Document Sponsored Document Sponsored Document Picton et al. Page 15
Table 1
Frequencies of identified polymorphisms within the SAA and SAC study populations.
Location on gene SNP position Base change (wt/mut) Accession number a n (population frequency) b SAA c SAC c 5′ UTR (2762 bp) -5268 G/A rs3136535 3 (0.043) 6 (0.086) -5266 G/A rs6776227 3 (0.043) 0 -5214 T/C NI 1 (0.014) 0 -5080 T/A rs41429449 4 (0.057) 0 -5072 C/T rs35078594 4 (0.057) 0 -4897 G/A NI 0 1 (0.014) -4808 G/A NI 2 (0.029) 6 (0.086) -4745 C/T rs3136536 9 (0.129) 0 -4630 T/C NI 0 1 (0.014) -4358 A/G rs7637813 2 (0.029) 16 (0.229) -4257 A/C rs41490645 0 6 (0.086) -4223 C/T NI 4 (0.057) 0 -4088 T/C rs41499550 14 (0.200) 0 -3949 A/G NI 0 1 (0.014) -3899 A/C rs72622924 10 (0.143) 42 (0.600) -3894 T/C rs41395049 14 (0.200) 0 -3886 C/T NI 0 3 (0.043) -3868 CTAT/-rs10577983 10 (0.143) 40 (0.557) -3833 C/T NI 1 (0.014) 0 -3458 G/T rs2734225 9 (0.129) 39 (0.557) -3432 T/C NI 1 (0.014) 0 -3261 G/A rs41475349 14 (0.200) 0 -2852 A/G rs2227010 19 (0.271) 25 (0.357) -2823 T/A NI 1 (0.014) 0 Exon 1 (57 bp) -2733 A/G rs2856758 3 (0.043) 7 (0.100) Intron 1 (501 bp) -2577 T/G NI 1 (0.014) 0 -2554 G/T rs2734648 22 (0.314) 39 (0.557) -2459 G/A rs1799987 30 (0.429) 28 (0.400) -2454 G/A NI 3 (0.043) 1 (0.014) Exon 2A (235 bp) -2150 A/G NI 1 (0.014) 0 -2135 T/C rs1799988 30 (0.429) 28 (0.400) -2132 C/T rs41469351 13 (0.186) 0 -2086 A/G rs1800023 9 (0.129) 39 (0.557) -2048 C/G rs41355345 0 1 (0.014) Published as: Infect Genet Evol. 2010 May ; 10(4): 487-494. Sponsored Document Sponsored Document Sponsored Document Picton et al. Page 16 Location on gene SNP position Base change (wt/mut) Accession number a n (population frequency) b SAA c SAC c Intron 2 (1903 bp) -1835 C/T rs1800024 11 (0.157) 3 (0.043) -1686 A/C rs9282632 17 (0.243) 0 -1464 A/G rs3181037 17 (0.243) 0 -1193 C/T NI 1 (0.014) 0 -1130 AG/-rs3054375 9 (0.129) 39 (0.557) -1060 C/T rs2856762 0 3 (0.043) -976 C/T rs2254089 9 (0.129) 39 (0.557) -975 G/A rs41395249 6 (0.086) 0 -730 A/T NI 2 (0.029) 0 -651 C/T rs2856764 9 (0.129) 39 (0.557) -451 C/T NI 3 (0.043) 1 (0.014) -444 G/A rs2856765/rs35046662 9 (0.129) 39 (0.557) -362 ACAA/G rs71619644 9 (0.129) 39 (0.557) -113 G/T rs3176763 17 (0.243) 0 -112 G/A rs41352147 1 (0.014) 0 Exon 3/ORF (1059 bp) +225 T/C rs1800941 1 (0.014) 0 +258 G/C NI 0 1 (0.014) +319 C/T Petersen et al. (2001) 1 (0.014) 0 +554 Δ32 rs333 0 5 (0.071) +673 C/T Petersen et al. (2001) 1 (0.014) 0 +1004 C/T rs1800944 5 (0.071) 1 (0.014) 3′ UTR (2651 bp) +1253 A/G NI 1 (0.014) 0 +1752 G/A rs41495153 18 (0.257) 0 +1810 G/A NI 1 (0.014) 0 +1823 C/T rs17765882 0 3 (0.043) +1843 G/A rs41418945 9 (0.129) 0 +1846 G/A rs41466044 9 (0.129) 0 +2066 G/A NI 2 (0.029) 0 +2077 G/T rs1800874 9 (0.129) 37 (0.529) +2225 T/C rs41535253 4 (0.057) 0 +2293 A/G rs41526948 0 1 (0.014) +2381 A/G NI 1 (0.014) 0 +2435 T/A NI 1 (0.014) 0 +2458 A/C rs3188094 6 (0.086) 0 +2676 C/A rs41442546 0 6 (0.086) +2772 G insertion NI 5 (0.071) 1 (0.014) Published as: Infect Genet Evol. 2010 May ; 10(4): 487-494. Sponsored Document Sponsored Document Sponsored Document Picton et al. Page 17 Location on gene SNP position Base change (wt/mut) Accession number a n (population frequency) b SAA c SAC c +2838 C/G rs41512547 3 (0.043) 0 +2919 T/G rs746492 28 (0.400) 27 (0.386) +3132 T/G NI 0 1 (0.014) a Accession numbers of SNPs detected in this study which have been previously reported in the SNP database (dbSNP) or reference to report not in database are listed here; NI indicates newly identified polymorphisms not found in dbSNP. b Frequency was calculated for both populations using total number of alleles, i.e., n = 70. c Grey shading highlights polymorphisms which were found to be restricted to either the SAA or the SAC population group.
Published as: Infect Genet Evol. 2010 May ; 10(4): 487-494.
Published as: Infect Genet Evol. 2010 May ; 10(4): 487-494.Sponsored DocumentSponsored Document Sponsored Document
Published as: Infect Genet Evol. 2010 May ; 10(4): 487-494.Sponsored Document Sponsored Document
This study was supported by grants from the
Page 10
Published as: Infect Genet Evol. 2010 May ; 10(4):
Sponsored Document
Refer to Web version on PubMed Central for supplementary material.
Refer to Web version on PubMed Central for supplementary material.
Directorate of Health, 2005). Thus, 20% of boys aged 13-14 years stated that they smoke cigarettes daily or occasionally, while the corresponding figure for use of snus were 29%. The figures for boys aged 12-13 years were 12% (cigarettes) and 15% (snus: Norwegian Directorate of Health, 2005). Other research reports among Norwegian adolescents indicate that fewer adolescents start to smoke, while at the same time, there is an increase in the use of snus (M. Lund & Lindbak, 2007;Øverland, Hetland, & Aarø, 2008). Furthermore, the use of snus is becoming more widespread among experienced smokers, indicating that snus may be used as an alternative nicotine delivery system either as a temporary substitute for cigarettes or as a quitting product (M. Lund & Lindbak, 2007;K. E. Lund, Tefre, Amundsen, & Nordlund 2008).
This development has raised concern for the Norwegian health authorities. The possible effect of snus at the population level has also led to a debate about whether the ban on sale of snus in the European Union has resulted in a loss or gain for public health (Scientific Committee on Emerging and Newly-Identified Health Risks [SCENIHR], 2008). On one hand, it could be argued that snus may attract new nicotine users who would otherwise not have started to use tobacco at all. Furthermore, snus may lead to later uptake of cigarettes among those who would not have started to smoke if not for their experience as snus users (the gateway hypothesis; Tomar, Fox, & Severson, 2009). On the other hand, compared with smoking cigarettes, the use of snus is found to be considerably less harmful (Royal College of Physicians [RCP], 2007; SCENIHR, 2008). It can thus be argued that snus is a better alternative for those who otherwise would have started to smoke cigarettes (a possible immunization effect) and that snus could be a quitting alternative for cigarette users who are not able or willing to quit (see K. E. Lund, 2009).
Given that one would like either to promote or to prevent the use of snus, for example, by targeting young people by means of persuasive communications, it is crucial to have information about the determinants of use of snus. However, few studies have provided insight into why adolescents start to use snus, and it might thus be worthwhile to explore the processes underlying the decision to use snus. One theoretical perspective that has been widely applied to explore the cognitive and
The use of snus (low-nitrosamine smokeless tobacco, Swedish type) is increasing in the United States (Alpert, Koh, & Conolly, 2008) as well as in Northern Europe (Gilljam & Lund, 2009). For example, in Norway, the use of snus is more prevalent than that of cigarettes in some segments of young men (Norwegian motivational underpinnings of addictive behaviors is derived from the concept of expectancies. Expectancies are defined as individuals' beliefs that a specific action will lead to specific consequences (Bandura, 1986). While perceived positive expectancies are anticipated to reinforce the particular behavior, negative expectancies are believed to restrain involvement in the behavior. In research on addictive behaviors, expectancies have traditionally been related to the use of a particular substance (see Brandon, Herzog, Irvin, & Gwaltney, 2004), including smoking (Brandon, Juliano, & Copeland, 1999). With regard to cigarette smoking, it has been documented that smokers have established a number of strong specific expectancies, and several versions of the Smoking Consequences Questionnaire (SCQ) have been developed assessing the role of smoking outcome expectancies (Brandon & Baker, 1991;Copeland, Brandon, & Quinn, 1995;Lewis-Esquerre, Rodrigue, & Kahler, 2005). In a recent study, Juliano and Brandon (2004) modified the SCQ-Adult (Copeland et al., 1995) by incorporating a number of different outcomes expected to be associated with the following products containing nicotine in addition to cigarettes: nicotine chewing gum, nicotine patch, and nicotine nasal spray, so-called nicotine replacement therapy products (NRTs). At a conceptual level, the range of items related to expected outcomes included five distinct expectancy scales: (a) "negative affect reduction" (the expectation that the particular nicotine product would help overcome unwanted affective states), (b) "craving reduction" (the expectation that the product would help control cravings), (c) "weight control" (the expectation that the product would be an efficient weight watcher), (d) "health risks" (the expectation that the particular product would be a risk to one's health), and (e) "quitting facilitation" (the expectation that the product would help during an attempt to quit smoking). The latter scale concerned only the NRTs.
The five expectancy scales showed high internal consistencies in terms of Cronbach's coefficient alpha, and it was also found that the higher the level of expectations of NRTs, the higher the level of immediate plans to quit smoking (Juliano & Brandon, 2004). The role of expectancies in the use of snus has yet to be established. In this study, we wanted to utilize the ideas of Juliano and Brandon in relation to the use of snus and test the factor structure of the five dimensional model using confirmatory factor analysis (CFA). We also wanted to investigate the role of the expectancy dimensions in the prediction of intentions to use snus applying a full structural equation model (SEM). Thus, we extended the work of Juliano and Brandon by using a CFA to investigate the hypothesized underlying structure of the expectancy items in relation to use of snus and by applying SEM analysis to estimate the predictive power of the expectancy factors for intentions to use snus. Finally, drawing on research of the Theory of planned behavior (TPB; Ajzen, 1991), we extended previous research on expectancies and nicotine products by examining the role of current (snus) behavior in the context of intentions to use snus. This research has shown that past behavior typically predicts behavioral intentions above the three TPB components: attitude, subjective norms, and perceived behavioral control (see Norman & Conner, 2006). In the present study, we had two categories available in terms of whether the respondents used snus "daily" versus "sometimes." Thus, we wanted to explore the possibility that there was a direct effect of "current (snus) behavior" on intentions in additions to the expectancy variables.
The present study had two objectives. First, we wanted to test how well the hypothesized five-factor solution fitted the data in terms of a set of expectancy items derived from Juliano and Brandon (2004) related to the use of snus using a CFA. Second, we wanted to explore the role of the snus expectancies in predicting intentions to use snus the next six months applying a full SEM. We also included current behavior in the analysis to investigate a possible direct effect on intentions. Furthermore, we controlled for the predictive role of age, sex, and smoking behavior.
The data stem from a questionnaire survey among first-year students at the University in Bergen and at the Norwegian School of Economics and Business Administration in 2004. Altogether, 858 students responded to the questionnaire, which constituted 25.6% of 3,344 registered first-year students who were invited to participate in the study. Among the respondents, 151 (17.6% of the total sample) were snus users. A relatively low percentage (6.6%) of the sample contained missing data on all the expectancy variables, and we thus decided to remove them from further analysis. Thus, the respondents in the present study consisted of 141 snus users, with a mean age of 20.9 years (SD = 2.1) and 71% were male. Thirty-eight percent were daily users, and 62% were occasional users of snus. The occasional users were on an average using 0.4 (SD =.5) boxes of snus per week, while the daily users were on an average using 2.2 (SD = 1.33) boxes of snus per week. The mean debut age of snus use was 17 years (SD = 2.6), while the mean age for becoming a regular user of snus was 18.3 years (SD = 3.3). Furthermore, 31% reported to have tried quitting using snus. Fifty-two percent reported that they were also smoking cigarettes (32% on a daily basis and 68% less frequently), and 12% reported to be former smokers.
Only the respondents who were using snus were asked to respond to the questions concerning snus expectancies. Participation was voluntary, and the project was approved by the National Committees for Research Ethics in Norway and reported to the Norwegian Social Science Data Services.
Thirteen expectancy items derived from the study of Juliano and Brandon (2004) were selected for the use of snus (see Table 1) and measured on a 5-point probability scale using response categories ranging from very unlikely (1) to very likely (5). The expectancy model involved five latent factors: negative affect reduction (four items, e.g., "snus helps me to relax"), craving reduction (two items, e.g., "snus satisfies my nicotine cravings"), weight control (three items, e.g., "snus keeps me from overeating"), health risks (two items, e.g., "snus is hazardous to my health"), and quitting facilitation (two items, e.g., "snus makes quitting smoking easier"; see Table 1).
Current snus behavior was measured in terms of (1) daily use and (0) occasional use. Intention to use snus was assessed with the following items "I expect to use snus the next six months" and "I intend to use snus the next six months." A 5-point probability scale with response categories ranging from 1 (very unlikely) to 5 (very likely) was used. The two intention items correlated strongly (r = .91) and were subsequently added to constitute a sum score of snus intentions.
The analyses were conducted in three steps applying SEM methodology in AMOS 17.0. First, we analyzed the dataset for missing values, skewness, and kurtosis. If maximum one negative error variance was detected, the variance was set to 0. Second, a CFA was used to test the hypothesized factor structure using maximum likelihood estimation. Various goodness-of-fit indices in the terms of c 2 , comparative fit index (CFI), and root mean squared error of approximation (RMSEA) were reported in the analysis (Mulaik et al., 1989). Model fit criteria have been a subject for discussions and have typically been set to CFI ≥.90 with >.95 representing a good fit and RMSEA ≤.08 with <.05 interpreted as a good fit (see Marsh, 2007 for a discussion). We also investigated the internal consistency of the individual scales using Cronbach's coefficient alpha (SPSS 17.0). Third, a full SEM analysis was applied by combining the measurement model with a structural model. The expectancy scales hypothesized to predict intentions to use snus the next six months were included in the model, and current use of snus, age, gender, and smoking behavior (smoker vs. nonsmoker) were included as control predictors.
Examination of the expectancy measures showed acceptable kurtosis and skewness (±2; Kline, 2005). The results of the CFAs are presented in Table 1 with a number of goodness-of-fit indices; c 2 (64) = 105.2, p = .00, CFI = 0.95; RMSEA = 0. 07, 90% CI = 0.04-0.09. The different indices showed that the hypothesized five-factor model provided a moderate fit to the data (Bentler, 1990;Kline, 2005). Internal consistencies (Cronbach's coefficient alpha) were satisfactory for all the five scales ranging from .66 to .87 (Table 1).
In the full structural model, age, sex, smoking behavior, and snus intensity were included in addition to the expectancy variables, and all the observed and latent variables were allowed to correlate; c 2 (96) = 179.24, p = .00, CFI = 0.91; RMSEA = 0.08, 90% CI = 0.06-0.09. Post-hoc modifications were performed in order to develop a better fitting model. As a result, four covariances were included. The residuals between expectancies of the use of snus to prevent overeating and expectancies of snus to be hazardous to health were allowed to correlate. As they both tap into concerns for healthiness and well-being, it seems reasonable that they are related. Furthermore, the residuals between expectancies of the use of snus to satisfy smoking urge and expectancies of the use of snus to increase the chance of quitting smoking were correlated as they both are associated with the use of snus as an alternative delivery source of nicotine. Also, the residuals between expectancies of the use of snus to keep weight and expectancies of snus to make quitting smoking easier were allowed to correlate, suggesting that both items tap into health concerns. Finally, the residual related to expectancies of the use of snus to satisfy nicotine craving and sex was allowed to correlate, indicating that there are gender differences concerning the role of snus as a craving device. The model fit was significantly improved with the addition of these four paths; Dc 2 (4) = 35.76, p = .00; c 2 (92) = 143.48, p = .00, CFI = 0.95, RMSEA = 0.06, 90% CI = 0.04-0.08. The results presented in Table 2 shows that (negative) expectancies of health risks using snus (b = -.22,
Table 1. Factor Loadings, Cronbach's Coefficient Alpha (bold) for the Five-Factor Solution Expectancy items Negative affect Craving reduction Weight control Health risks Quitting facilitation Snus . . . Negative affect reduction .85 . . . helps me to relax .67 . . . helps me deal with anger .71 . . . calms me down when I feel nervous .87 . . . helps me reduce or handle tension .87 Craving reduction .77 . . . satisfies my nicotine cravings .76 . . . satisfies my urge to smoke .83 Weight control .87 . . . keeps me from overeating .81 . . . keeps my weight down .96 . . . keeps me from eating more than I should .76 Health risks .66 . . . is hazardous to my health .50 . . . increases the risk of cancer 1 Quitting facilitation .81 . . . increases my chances of quitting smoking .97 . . . makes my quitting smoking easier .71
Note. Goodness-of-fit statistics: c 2 (64) = 105.2, p = .00, CFI = 0.95, RMSEA = 0.07. CFI = comparative fit index; RMSEA = root mean squared error of approximation. p =.01) and current snus behavior (b = .18, p = .05) turned out to be the only two significant predictors of intentions to use snus. Thus, the lower the level of expectation that snus was harmful to one's health, the stronger the intentions to use snus in the next six months. In addition, daily snus users demonstrated stronger intentions to use snus in the future than occasional users. Furthermore, a chi-square difference test revealed that entering current behavior significantly improved the model Dc 2 (11) = 32.7, p < .00. The full model was able to explain 27% of the variance in intentions.
To our knowledge, this is the first study to explore expectancies associated with the use of snus and to examine their role in the prediction of intention to use snus. The study was based on the ideas and results from a previous study by Juliano and Brandon (2004) on expectancies related to cigarettes and different NRT products: nicotine gum, nicotine spray, and nicotine patch. On conceptual grounds, they introduced five different expectancy concepts related to different NRT products: (a) reduce negative affect, (b) fulfill a craving for nicotine, (c) facilitate quitting smoking, (d) health risks, and (e) effective weight control. The present study extended these results to the area of snus in two directions. First, we used a CFA to test the dimensionality of 13 expectancy items in relation to another nicotine product, namely snus. The five-factor model showed an acceptable fit to the data. The internal consistency of the five scales was good, but craving reduction, quitting facilitation, and health risks might benefit from adding more items to the scales. This indicates that the hypothesized five-factor model was fully applicable in another behavioral area. Second, we combined a measurement model with a structural model to identify the predictive ability of the expectancy scales in the formation of intentions to use snus in the next six months, thus providing increased insight into the cognitive and motivational underpinnings of the behavior. The expectancy scales, sex, age, snus intensity, and smoking behavior were able to account for 27% of the variance in intentions to use snus. Of the expectancy measures, only expectancies of health risks of using snus turned out to be a significant predictor of intentions so that the lower the expectation that snus constitutes a health risk, the stronger the intentions to continue to use snus. Furthermore, the study showed that current snus behavior significantly explained variance in intentions in the direction of stronger intentions among those of the respondents who were daily users of snus.
The predictive role of health risks parallels the predictive power of perceived health risks in the area of quitting smoking (Rise & Kovac, 2009;Wetter et al., 1994). Although it is documented that the health hazard associated with the use of snus is considerably less than for cigarette smoking, the use of snus is not without risks (Lee & Hamling, 2009;Levy et al., 2004;RCP, 2007;SCENIHR, 2008). Given that health authorities want to encourage adolescents quit using snus, the results suggest that it might be useful to target expectancies related to health risks in persuasive communications.
In contrast to previous expectancy research on NRTs where intentions to quit smoking were correlated with expectancies of NRTs to facilitate quitting (Juliano & Brandon, 2004), expectancies of snus to help quit smoking did not predict snus intentions in the present study. For the prediction of intention to use snus, expectancies of snus to facilitate quitting smoking are primarily relevant among current smokers. Thus, difference in smoking experience might explain the lack of effect of the predictor. This should be addressed in future studies.
The direct effect of current snus behavior beyond the effects of the expectancy components on the intention formation process may be explained in two different ways. First, it may be that the intention measure is partly a self-prediction, that is, snus users do not make a decision whether or not to continue to use snus but rather make a likelihood judgment of what they are going to do in the specified period based on a simple extrapolation from recent performances: "If I have used snus on a daily basis before, I will probably do it in the next 6 months" (see Rise, Åstrøm, & Sutton, 1998). Second, it may be that central predictors of intentions are left out of the equation (see Conner & Armitage, 1998). For example, aspects of social influence along the line proposed by social expectancy scales in the area of alcohol in terms of social facilitation (Goldman, Greenbaum, & Darkes, 1997) may represent a potential predictor candidate. In a similar vein, a recent study among Norwegian adolescents showed that young men in Norway perceive snus as trendy and attractive (Wiium, Aarø, & Hetland, 2009).
It may be argued that 27% explained variance is a relatively low figure as compared with those analyses obtained using the TPB. In this context, a meta-analysis showed that the three theoretical components, attitude, subjective norm, and perceived behavioral control accounted for an average of 39% of the variance in intentions across 185 studies (Armitage & Conner, 2001). The difference in explained variance may partly be explained by the fact that we did not adhere to the principle of compatibility as proposed by the TPB (Ajzen, 2002). While the target variable specifies the time aspect (e.g., "I intend to use snus in the next six months"), the expectancy measures lack the
Table 2. Standardized Regression Coefficients on Intentions to Use Snus the Next Six Months From the Expectancy Variables "Negative Affect," "Weight Control," "Quitting Facilitation," "Health Risks," "Craving Reduction," "Current Behaviour," "Age," "Sex," and "Smoking Behaviour" (N = 141) Model Predictors b p values Negative affect .10 .39 Weight Control -.15 .13 Quitting facilitation .10 .33 Health risks -.22 .01 Craving reduction .18 .33 Current behaviour .18 .05 Age .00 .99 Sex -.16 .07 Smoking behaviour .02 .80 Explained variance .27
Note. Goodness-of-fit statistics: c 2 (92) = 143.48, p = .00, CFI = 0.95, RMSEA = 0.06. CFI = comparative fit index; RMSEA = root mean squared error of approximation.
specification of time and are thus phrased in more global terms (e.g., "snus is harmful for me"). Nevertheless, this is a common way of operationalizing expectancies in current expectancy research.
Because the response rate in the present study was low as well as the fact that snus users as a group were underrepresented in this study (see K. E. Lund et al., 2008), extrapolation of the results to snus users in general should be done with caution. Nevertheless, generalizations based on underlying processes in terms of associations between variables have been found to be less vulnerable to sampling procedures than those of prevalence (Aaberge & Laake, 1984).
We would like to thank
This work was supported by the
None declared.
Genetic defects in amino acid metabolism are major causes of newborn diseases that often lead to abnormal development and function of the central nervous system. Their direct impact on cardiac development and function has rarely been investigated. Recently, the authors have established that a mitochondrial targeted 2C-type ser/ thr protein phosphatase, PP2Cm, is the endogenous phosphatase of the branched-chain alpha keto acid-dehydrogenase complex (BCKD) and functions as a key regulator in branched-chain amino acid catabolism and homeostasis. Genetic inactivation of PP2Cm in mice leads to significant elevation in plasma concentrations of branched-chain amino acids and branched-chain keto acids at levels similar to those associated with intermediate mild forms of maple syrup urine disease. In addition to neuronal tissues, PP2Cm is highly expressed in cardiac muscle, and its expression is diminished in a heart under pathologic stresses. Whereas phenotypic features of heart failure are seen in PP2Cmdeficient zebra fish embryos, cardiac function in PP2Cmnull mice is compromised at a young age and deteriorates faster by mechanical overload. These observations suggest that the catabolism of branched-chain amino acids also has physiologic significance in maintaining normal cardiac function. Defects in PP2Cm-mediated catabolism of branched-chain amino acids can be a potential novel mechanism not only for maple syrup urine disease but also for congenital heart diseases and heart failure.
Amino acids are key nutrient molecules essential for cell growth, survival, and normal function. In addition to providing building blocks for protein synthesis, many amino acids are an essential ingredient for biosynthesis of other molecules as the sources of nitrogen and carbon [8]. Some amino acids, including branched-chain amino acids, also have been shown to possess a potent signaling function to regulate global growth and metabolism [2,13,14,17,18,22,23]. Therefore, the impact of amino acid metabolism on embryonic development and human congenital diseases has long been recognized.
Maple syrup urine disease, one of the most common genetic disorders caused by defects in amino acid metabolism, is a target of mandatory newborn screening in most of the United States and throughout the world [5,10,20]. The clinical manifestations of classical maple syrup urine disease are mostly neuropathologic, including seizure and mental retardation [7]. However, recent studies from animal models of maple syrup urine disease raise concerns about the potential adverse impact of branched-chain amino acid metabolic defects on cardiac development and function. These concerns are the focus of the discussion in this review.
Branched-chain amino acids (BCAAs) including leucine, isoleucine, and valine are essential amino acids that must H. Sun Á G. Lu Á S. Ren Á J. Chen Á Y. Wang (&) Division of Molecular Medicine, Departments of Anesthesiology, Medicine and Physiology, Molecular Biology Institute, Cardiovascular Research Laboratories, David Geffen School of Medicine, CSH, Room BH 569, 650 Charles E. Young Drive, Los Angeles, CA 90095, USA e-mail: yibinwang@mednet.ucla.edu be acquired from external food. In addition to their abundant presence in protein peptides as key hydrophobic building blocks, BCAAs also serve as significant sources for biosynthesis of sterol, keto bodies, and glucose [3]. Among the BCAAs, particularly leucine has potent signaling activity to promote protein synthesis, cellular metabolism, and cell growth in a mammalian target of rapamycin (mTOR)-dependent manner [4,15,19].
Although BCAAs are necessary for normal growth and function at cellular and organism levels, an excess amount of free BCAAs also can be pathologic. A high plasma level of BCAAs is the benchmark and cause of maple syrup urine disease, a potentially life-threatening disorder affecting 1 of 180,000 newborn babies on the average in the general population, with a much higher prevalence among certain Amish, Mennonite, and Jewish communities [5,20].
Because BCAAs are essential amino acids with no biosynthesis pathways in animal cells, their homeostasis is determined largely by catabolic activities in a number of organs, particularly the liver [9]. The first step in the catabolism of BCAAs is carried out in brain, muscle, and many nonhepatic tissues by the branched-chain aminotransferase (BCAT) to convert BCAAs into branched-chain alpha keto acids (BCKAs) [9]. The BCKAs are decarboxylated by the branched-chain alpha keto acid dehydrogenase (BCKD) complex in the liver as well as other tissues and eventually degraded into acetyl-coenzyme A (CoA) or succinyl-CoA to fuel the TCA cycle. As BCKD mediates this rate-limiting step in BCAA catabolism, its activity dictates the steadystate levels of BCAA and BCKA. It thus is a target of multiple regulatory mechanisms, including cyclic adenosine monophosphate (cAMP)-mediated induction of RNA transcription and phosphorylation-mediated inhibition/activation of its enzymatic activity.
The BCKD complex is genetically linked with the pyruvate dehydrogenase complex (PDH) because they share extensive homology in their subunit composition and regulation. Like PDH, BCKD holoenzyme contains 24 copies of catalytic subunits E2/E3 and an equal number of regulatory subunits E1a and E1b. At low BCAA levels, E1a is hyperphosphorylated by BCKD kinase, leading to lower BCKD activity and reduced loss of BCAA. At a high BCAA level, E1a is dephosphorylated by BCKD phosphatase, leading to induced BCKD activity and the removal of excess BCAA. Therefore, BCKD phosphorylation/ dephosphorylation is critical to BCAA homeostasis [9].
Based on the importance of BCKD phosphorylation in its regulation, BCKD phosphatase has been well established as a key regulator in BCAA catabolism. However, its molecular identity has remained elusive for decades. Recently, through genome scanning, we discovered a mitochondrial targeted 2C-type ser/thr protein phosphatase that we named PP2C in mitochondria (PP2Cm) [11].
Through extensive proteomic and biochemical analysis, we established that PP2Cm is the missing BCKD phosphatase responsible for BCAA-induced dephosphorylation and activation of BCKD activity [12]. This conclusion is based on the following lines of evidence: (1) BCKD subunits, including E2 and E1a, are specifically associated with PP2Cm in vitro and in vivo; (2) PP2Cm expression effectively dephosphorylates E1a Ser-293 phosphorylation, a key residue in BCKD enzymatic regulation; (3) PP2Cmdeficient cells have elevated basal E1a phosphorylation, which remains high at BCKA treatment; (4) PP2Cm-deficient mice have a significantly higher plasma level of BCAA and BCKA with impaired BCAA clearance after a high dose of BCAA ingestion; and (5) PP2Cm-deficient newborn mice have significantly higher mortality under a high-protein diet challenge.
All these data establish that PP2Cm is the endogenous BCKD phosphatase with an essential function in BCAA catabolism. The analysis with PP2Cm KO mice indicates that a defect in PP2Cm is a potential novel mechanism for the intermediate/inducible forms of maple syrup urine disease.
The discovery of PP2Cm as the BCKD phosphatase and a key regulator in BCAA metabolism as well as the establishment of PP2Cm-deficient zebra fish and mouse models offers an opportunity to investigate the impact of BCAA regulation in intact animals. In both zebra fish embryos and adult mice, PP2Cm is highly expressed in both cardiac muscle and the central nervous system (Figs. 1, 2). Its expression in the heart is dynamically regulated by stress, as measured by both mRNA and protein levels, with significantly reduced levels in hypertrophic and failing hearts [11].
It is not clear whether loss of PP2Cm is a contributing factor to cardiac pathology or simply a phenomenon associated with the diseased hearts. To investigate this question, we analyzed the cardiac performance in PP2Cm-deficient zebra fish and mice. Using morpholingoes specifically targeted to zPP2Cm translation initiation codon (ATG-MO), we demonstrated that. PP2Cm-deficient fish embryos displayed a dose-dependent loss of cardiac contractility and that most of them did not survive beyond the early embryonic stage [11]. Accelerated heart failure developed in PP2Cmdeficient mice after mechanical overload induced by transaortic constriction (Fig. 3). These evidences suggest that loss of PP2Cm is not a mere consequence but rather a significant contributor to the pathogenesis of heart failure. These studies for the first time implicated BCAA catabolism also as an important aspect of cardiac pathophysiology and showed that defects in BCAA homeostasis can have a significant adverse impact on cardiac function and disease progression.
The underlying mechanisms for the clinical manifestations of maple syrup urine disease still are not well established, and the adverse impact on glutamine transport and the induction of reactive-oxygen species (ROS) have been implicated [1,6,24]. In our own studies with cardiac tissue or cultured cardiomyocytes, we also explored a number of these possible mechanisms.
1. We examined the cytotoxic effect of the accumulated BCAA/BCKA on cadiomyocyte survival, demonstrating that PP2Cm-deficient mice have a significantly elevated ROS level and are more susceptible to calcium-induced permeability transition pore opening in mitochondria [11]. Whether and how this effect is caused by accumulated BCAA or BCKA remains to be investigated further. However, a clear implication of the observed cytotoxicity is more cell death. Indeed, we observed that PP2Cm deficiency can lead to more apoptotic cell death in both developing embryos and cultured myocytes. 2. Metabolic effects of accumulated BCAA/BCKA can be pathologic in the heart. It is known that BCAA/ BCKA can inhibit pyruvate and fatty acid transport and utilization [16,19]. Because cardiac tissue has a constant high demand for pyruvate and fatty acid as its main fuel source, chronically elevated BCAA/BCKA can potentially block normal bioenergenic homeostasis 3. Our findings showed that PP2Cm mediated direct modification of mitochondrial function in cellular bioenergenics and survival. Although BCKD is the only substrate identified for PP2Cm to date, it is possible that other uncharacterized substrates in mitochondria contribute to PP2Cm-mediated signaling. This is supported by our observation that the isolated 2.5ng 5ng ATG-MO ATG-MO HR=100.6±3.0 FS%=46.2% HR=95.2±7.5 FS%=33.3% HR=86.7±5.5 FS%=23.5% Line Scanning Control A B Fig. 2 Loss of fPP2Cm expression leads to heart failure in zebra fish embryo. a Video frames of cmlc-GFP transgenic fish embryo recorded under ultraviolet illumination after treatment with different concentrations of PP2Cm morpholigo. b M-mode image of zebra fish hearts from linescanning analysis as shown in A as well as heart rate (HR) and percentile fractional shortening (FS%) measurements. Adapted from Lu et al. [11] with permission HR: 545, %EF: 60.59, %FS: 32.17 HR: 545, %EF: 37.85, %FS: 18.18 HR: 533, %EF: 65.69, %FS:35.88 HR: 545, %EF: 26.61, %FS: 12.50 Wildtype PP2Cm KO A B Pre-TAC Post-TAC Fig. 3 Expression and function of PP2Cm in the heart. a Expression of PP2Cm in the heart illustrated by positive Lac-Z staining in a PP2Cm ?/lacZ heart. b Representative echocardiogram of a wild type and a PP2Cm KO heart before and after 8 weeks of pressure overload induced by TA
PP2Cm-deficient mitochondria have abnormal sensitivity to calcium-induced permeability transition pore opening even in the absence of BCAA or BCKA treatment [11]. The molecular basis of this observation remains unclear, but it may imply that branched-amino acid metabolism is functionally coupled with mitochondrial inner membrane permeability, which has a key role in mitochondrial calcium, respiration coupling, ROS, and cellular viability.
Branched-chain amino acids are important nutrient molecules with a potent signaling effect. Free BCAAs and their catabolic intermediates, BCKAs, are tightly maintained in animals by a highly regulated catabolic pathway. Defects in catabolism of BCAAs cause maple syrup urine disease, one of the most common metabolic disorders in the human population.
Although most clinical features of maple syrup urine disease are neurologic, our recent findings from cellular studies and animal models suggest that missing a key regulator in the catabolism of BCAAs also can cause a significant impairment in cardiac function (Fig. 4). The implication of this observation remains to be further established. However, a recent study based on metabolic profiling of peripheral blood has demonstrated a link between abnormal metabolism of BCAAs and coronary diseases [21]. As the complete molecular components of BCAA catabolic pathways are being discovered, better diagnosis will be available to identify patients with nonclassic, intermediate, or inducible maple syrup urine disease based on sequencing evidence. The insights learned from our study argue strongly for a better understanding of the role that the catabolism of BCAAs plays in the heart and its potential impact on cardiac pathology in addition to its damage to the central nervous system.
Acknowledgments This work was supported in part by
The KAshinhou Tool for Ecotoxicity (KATE) system, including ecotoxicity quantitative structure-activity relationship (QSAR) models, was developed by the Japanese National Institute for Environmental Studies (NIES) using the database of aquatic toxicity results gathered by the Japanese Ministry of the Environment and the US EPA fathead minnow database. In this system chemicals can be entered according to their one-dimensional structures and classified by substructure. The QSAR equations for predicting the toxicity of a chemical compound assume a linear correlation between its log P value and its aquatic toxicity. KATE uses a structural domain called C-judgement, defined by the substructures of specified functional groups in the QSAR models. Internal validation by the leaveone-out method confirms that the QSAR equations, with r 2 40.7, RMSE 0.5, and n45, give acceptable q 2 values. Such external validation indicates that a group of chemicals with an in-domain of KATE C-judgements exhibits a lower root mean square error (RMSE). These findings demonstrate that the KATE system has the potential to enable chemicals to be categorised as potential hazards.
Quantitative structure-activity relationships (QSARs) are potential tools for predicting the activity and properties of chemicals, including their physicochemical attributes, health effects, ecotoxicity and biological activity. QSAR models can estimate and predict such activity and can thus be used to categorise chemicals in terms of their potentially hazardous nature. A recent review has demonstrated that acute aquatic toxicity [1] can be predicted using QSAR and describes the available databases of ecotoxicity data.
Prediction of toxicity by QSAR does not require lengthy experiments, nor the use of animals, plants or cells. QSAR models have therefore been utilised for the assessment of new and existing chemicals for conformity with regulatory requirements in countries within the Organisation for Economic Co-operation and Development (OECD) [2]. In Japan, under the Chemical Substances Control Law (CSCL), the Ministry of the Environment (MoE) is responsible for evaluating the adverse effects of chemicals on ecosystems, and uses tests involving aquatic organisms such as Oryzias latipes (fishes) or Daphnia magna (daphnia), in addition to algae data available from the MoE website [3]. The Japanese National Institute for Environmental Studies (NIES) was established to apply QSAR models to acute ecotoxicity, and has developed a QSAR prediction system using the MoE ecotoxicity database. This system, published in March 2009, is known as the KAshinhou Tool for Ecotoxicity (KATE) [4].
The present paper focuses on the theoretical and methodological aspects of the KATE system, and QSAR equations classified by chemical substructure are introduced. We shall then present the cross-validation ('leave-one-out') results, and the toxicities calculated by KATE, and by alternative systems such as TIssue MEtabolism Simulator (TIMES) [5,6] (developed by Zlatarov at the Laboratory of Mathematical Chemistry, Bourgas University, Bulgaria), and by ECOSAR TM [7] (developed by the US Environmental Protection Agency (EPA)) using the same end-point data set as that in KATE. The validity of KATE will be discussed using the applicability domain, log P, and C-judgements.
End-point KATE uses experimental data on chemical substances to predict aquatic toxicity. The end-points of interest are the 96-hour median lethal concentration (LC 50 ) in fish after acute toxicity tests, and the 48-hour median effective concentration (EC 50 ) in daphnia obtained after acute immobilisation tests. Training sets for QSAR development were derived from the results of ecotoxicity tests (Oryzias latipes LC 50 and Daphnia magna EC 50 ) obtained by the MoE [3], as well as the results of acute toxicity tests from the US EPA fathead minnow (Pimephales promelas) database [8,9]. In the KATE system, the 96-hour LC 50 data for Oryzias latipes and fathead minnow were combined to reinforce the number of reference datasets. The QSAR equations in KATE for the fish and daphnia end-points were designed using 535 and 258 chemicals, respectively.
Chemical substances can be classified according to the substructures that give rise to specific chemical properties (Appendix 1 of the supplementary material which is available on the Supplementary Content tab of the article's online page at http://dx.doi.org/10.1080/ 1062936X.2010.501815). The rules for daphnia and fish end-points are identical, except for the following five classes: amines aromatic or phenols1, amines aromatic or phenols3, amines aromatic or phenols4, amines aromatic or phenols5, and primary amines. According to KATE, the toxicity of a chemical containing amino functional groups might be different in daphnia from its toxic behaviour in fish.
Forty-four classes are proposed for each end-point of KATE QSAR models. Table 1 shows the QSAR class name, and the detailed class features are listed in Appendix 2 of the supplementary material (available online). The chemicals in the KATE unclassified class were not categorised within any of the rules in Appendix 2. Additional classification rules or fragment definitions are required in further studies to reduce the number of chemicals described as unclassified. It should be noted that the concept of unclassified within KATE does not always include reactive chemicals, and thus differs from the reactive unspecified category in the TIMES software. Ã1 C: an equation is generated by calculated Clog P. N: a member of the Neutral organics class. Note: n, RMSE, r 2 and q 2 denote the number of chemicals in a class, the root mean square error, the squared correlation coefficient, and the leave-one-out version of the squared correlation coefficient, respectively. The log P range shows minimum and maximum log P values.
Neutral organics is an aggregate of the chemicals in defined classes in the KATE system. It comprises the classes: nitriles aliphatic, ketones, alcohols or ethers aliphatic, phosphates, hydrocarbons aliphatic, ethers aliphatic and ethers aromatic. In the OECD Environment Monograph [10], neutral organic compounds of minimal toxicity were divided into the groups: aliphatic alcohols, aliphatic ketones, aliphatic ethers and alkoxyethers, aliphatic halogenated hydrocarbons, saturated alkanes and halogenated benzenes. Some of the neutral organics compounds defined in the OECD monograph were categorised differently from those in KATE.
The QSAR equations in the KATE model express the correlation between the octanol/water partition coefficient (log P) of a compound and its aquatic toxicity, using simple linear regression analysis. Measured log P values were used to derive the QSAR equations, except for the equations labelled C in Tables 1 and 2. In cases where experimental log P values were not available, an equation was constructed from the calculated Clog P value obtained by the Daylight toolkit [11]. The LC 50 and EC 50 values in the equation were expressed in terms of the common logarithm of the inverse of millimoles per litre (mmol L À1 , or mM). The equations and the statistical information obtained are shown in Tables 1 and 2. Where there were fewer than three sets of reference data within one class, QSAR prediction could not be performed. In such cases the class name was the only information obtained from KATE, and the label NO-QSAR is indicated in Tables 1 and 2. The equation for a class named pyrethroids was not constructed, since the log P values in the reference data were gathered in higher ranges [6.1, 6.5].
KATE offers two 'judgements' to verify whether or not a predicted chemical substance falls within the applicability domain of a QSAR class. The first is the log P judgement, based on the log P range defined by the reference chemical data of the class concerned. This has been categorised as a descriptor domain [12,13]. The interpolated log P range for each class is listed in Tables 1 and 2.
The second is the C-judgement, which is categorised as a structural domain and is defined by the substructures shown in Appendix 3 of the supplementary material (available online). The substructures are based on functional groups having similar concepts to those used by Schultz et al. [13], rather than on atom-centred fragments [12,14]. Schultz et al. applied the structural domain to one QSAR equation for aromatic compounds, and the out-of-domain revealed well-known electrophoric mechanisms in the structural space(s) [13]. In the KATE system the classification rules (described in Section 2.2) play a role in constructing such structural space(s). The definition of the applicability domain of C-judgement depends on whether all the substructures of the chemical under test are found in reference chemicals in the class, or secondly, whether all substructures in the test chemical are present in reference chemicals in either neutral organics or the class concerned. The first of these definitions is stricter than the second. The reliability of the log P and C-judgements is assessed later in Section 4 (Results and discussion). In the KATE system, the input is simplified molecular input line entry specification (SMILES) and log P (if available) for toxicity prediction, and the output is the calculated toxicity concentration (LC 50 or EC 50 ), the QSAR class found for the predicted chemical, and the domain judgements. If the measured log P of a chemical is not available, the calculated log P according to the SMILES information (KOWWIN or C log P) is adopted.
First, leave-one-out cross validations were examined for training sets used in the QSAR equations of KATE. Secondly, external validations were performed using test set compounds not included in the KATE training sets due to lack of measured log P values. The 287 fish 96-hour LC 50 and 98 daphnia 48-hour EC 50 from the Japan MoE, along with the US EPA fathead minnow database, were used for comparison of the calculated toxicity by the KATE software version published in March 2009, TIMES v. 2.25, and ECOSAR v. 0.99 h (1999).
It is worth mentioning that the end-points of the data calculated by KATE were not identical to those calculated by TIMES and ECOSAR. Fish (mixed with Oryzias latipes and fathead minnow acute toxicity tests) 96-hour LC 50 and daphnia 48-hour EC 50 (KATE), Pimephales promelas 96-hour LC 50 and daphnia 48-hour EC 50 (TIMES), and fish 96-hour LC 50 and daphnia 48-hour LC 50 (ECOSAR) were therefore adopted. The input of KATE and ECOSAR were SMILES strings, and calculated log P by KOWWIN. In TIMES, only the lists of SMILES strings were used as input values, and quantum chemical calculations were performed using MOPAC AM1 Hamiltonian, using the 'precise' option, without taking other conformers into account.
The QSAR equations were validated by the leave-one-out method obtained from the KATE system. The complete list of results is given in Appendix 4 of the supplementary material (available online). The statistical data are displayed in Tables 1 and 2. The criterion proposed by Hulzebos and Posthumus [16] was evaluated, in which the estimations from models should not deviate from the experimental value by a factor of 10 or above. For fish, 575 of the 628 chemicals met the acceptable criteria, and for daphnia 241 of 290 did so. (In this instance the 628 and 290 chemicals involved some degree of duplication.) Using the QSAR equations in the KATE system, more than 80% of chemicals were predicted within a factor of 10. The classes with less than a 0.7 squared correlation coefficient (r 2 50.7), and/or more than 0.5 RMSE, tended to increase the number of chemical substances in the unacceptable group. For example, the fish hydrocarbons aromatic class had 43 reference data, r 2 ¼ 0.826, RMSE ¼ 0.368, and only one unacceptable chemical. In other words, 98% of the chemicals were classed as acceptable. On the other hand, the fish dinitrobenzene class contained 12 reference data, r 2 ¼ 0.331, RMSE ¼ 0.669, and three unacceptable chemicals. In this case, 75% of the chemicals were thus acceptable.
As shown in Tables 1 and 2, each of the classes with r 2 ! 0.7, RMSE 0.5, and n45, e.g., the fish hydrocarbon aromatic class, had a sufficiently high q 2 . Such classes showed QSAR equations similar to those of neutral organics. Thus the toxicity of such classes could be explained mainly by the narcotic effect of the chemicals. However, the daphnia amines aromatic or phenols4 and amines aromatic or phenols5 groups had a larger intercept b in the QSAR equations than neutral organics with a small log P value (see Figure 1). These classes can be explained in terms of polar narcosis or narcosis II [17]. Narcosis II is known to be more toxic than baseline toxicity, i.e., than neutral organics, non-polar narcosis, narcosis I, or less inert, as explained by Verhaar et al. [18].
In some cases the q 2 values were much smaller than those of r 2 . QSAR equations based on fewer than six reference data require a greater number of reference chemicals.
Tables 3 and 4 list the statistical data of the TIMES, ECOSAR, and KATE with or without the applicability domains. The complete results are given in Appendix 5 of the supplementary material. First, we will focus on the TIMES, ECOSAR, and all the KATE results, without considering any applicability domains. In fish, the determination coefficient, r 2 , and RMSE using KATE (r 2 ¼ 0.868 and RMSE ¼ 0.658) were larger and smaller, respectively, than those using TIMES (r 2 ¼ 0.751 and RMSE ¼ 0.935) and than by ECOSAR (r 2 ¼ 0.790 and RMSE ¼ 0.869). For daphnia, RMSE using KATE (0.993) was smaller than that using TIMES (1.404) and ECOSAR (1.364). However, r 2 using KATE (0.662) showed no noticeable advantage over that by TIMES (0.668) or ECOSAR (0.699). Since reference data for the daphnia end-point (258 chemicals) numbered only half of those for fish (535 chemicals), the reference data for each QSAR equation for daphnia would therefore be less satisfactory for predicting toxicity. The addition of reference data and a change in the classification rules can recover the values of the statistical data. Notes:
Each chemical is identified by one QSAR class.
When a chemical is found to belong to more than one QSAR class, all the estimated data are adopted. If only the name of the class is available, such data are omitted.
Both in-domain and out-of-domain data for log P and C-judgements are included.
In-domain of log P-judgement.
In-domain of C-judgement is defined as all substructures of a test chemical being found in reference chemicals in the class.
In-domain of C-judgement defined as all substructures of a test chemical being in reference chemicals in either Neutral organics or the class.
The number of compounds that can be predicted.
The total number of the predicted values by using the training sets. Some chemicals belong to more than one class, and thus Predicted is larger than Chemicals. r 2 , RMSE, Under and Over were calculated based on the Predicted number.
Fractions (%) of the underestimated chemicals. Underestimation is defined as [calculated log(1/LC 50 ) -measured log(1/LC 50 )]5À1.
Fractions (%) of the overestimated chemicals. Overestimation is defined as [calculated log(1/LC 50 ) -measured log(1/LC 50 )]41. Notes: As in Table 3.
A fraction of log(1/LC 50 ) with an underestimation of less than À1 indicated that, compared with KATE, TIMES and ECOSAR tended to underestimate the toxicities of both fish and daphnia. On the other hand, a fraction of log(1/LC 50 ) showing an overestimation of more than 1 indicated that, compared with TIMES, ECOSAR and KATE tended to overestimate toxicity in both fish and daphnia. Considering these underand over-estimation fractions, we find that KATE gives a higher predictive ability in acute Oryzias latipes and Daphnia magna toxicity tests than does TIMES or ECOSAR. If the alert: Out of domain, in TIMES, and the applicable log P range in ECOSAR are considered rigidly, the correlation between measured and calculated toxicity is improved in TIMES and ECOSAR. Secondly, in fish, the RMSE of one of any in-domains was smaller than if domains were not considered. However, the r 2 in-domain of log P showed no particular improvement. For daphnia, r 2 and RMSE for one of any in-domains were larger and smaller, respectively, than those without considering domains. In the present study, either the descriptor and/or structural domains were related to the reduction of RMSE and the fraction of underestimated chemicals, especially if both domains were considered simultaneously. Additionally, the stricter structural domain C(1) (shown in Tables 3 and 4) demonstrated better predictive performance than the structural domain C(2). The systematic study of the domain based on the atom-centred fragment (ACF) approach by Kuhne et al. [14] showed that the ACF varied with respect to its size in terms of the path length, and the ACF match mode was specified in terms of degree of strictness. They also demonstrated a clear relationship between predictive performance and the levels of the ACF definition and match mode [14]. Even though the definition of substructures for the domain are different, the improvement by using C-judgement is similar in concept to that using the ACF approach. Thus, the log P range of the equation and C-judgement are useful for assessing the applicability of the QSAR results.
We have reported on the KATE system, encompassing a full list of classifications of the QSAR equations and KATE validations. In the KATE system chemicals are classified by their substructure. The QSAR equations express the correlation between log P and log(1/LC 50 ) or log(1/EC 50 ) of a chemical by simple linear regression analyses. The classes of QSAR equations are characterised by fragments of chemicals, except for the neutral organics class. The descriptor and structure domains, log P and C-judgements, in KATE were also introduced.
The cross-validation of the KATE system showed that QSAR equations with higher r 2 and lower RMSE with n45 gave a reliably higher q 2 than the other QSAR equations in KATE, meaning they had better predictive ability. A comparison of KATE, TIMES, and ECOSAR revealed that KATE was more accurate, due to end-point dependence. The use of log P and the C-judgement improved the statistical data. Thus the KATE system is a powerful tool for predicting acute toxicity in Oryzias latipes and Daphnia magna when the log P and C-judgement can be confirmed. Also, KATE has the potential to be useful in risk assessment.
The next topics in QSAR development will be to consider the reactivity of chemicals, and to include multi-regression analysis. The quantum chemical parameters, such as partial charges, are candidates for additional descriptors. Other ways of significantly increasing the reliability of toxicity prediction will be to improve the classification of the substructures, increase the reference data in a QSAR equation, and to refine the C-judgement.
Ã1 C:
KATE was researched and developed by the
E7820 is an orally active inhibitor of α 2 -integrin mRNA expression, currently tested in phases I and II. We aimed to evaluate what levels of inhibition of integrin expression are needed to achieve tumor stasis in mice, and to compare this to the level of inhibition achieved in humans. Tumor growth inhibition was measured in mice bearing a pancreatic KP-1 tumor, dosed at 12.5-200 mg/kg over 21 days. In the phase I study, E7820 was administered daily for 28 days over a range of 0-200 mg, followed by a 7-day washout period. PK-PD models were developed in NONMEM. α 2 -Integrin expression measured on platelets, corresponding to tumor stasis at t=21 in 50% and 90% of the mice (I int,50 , I int,90 ) were calculated. It was evaluated if these levels of inhibition could be achieved in patients at tolerable doses. One hundred nineteen α 2 -Integrin measurements and 210 tumor size measurements were available from mice. The relationship between PK and α 2 -integrin expression was modeled using an indirect-effect model, subsequently linked to an exponential tumor growth model. I inh,50 and I inh,90 were 14.7% (RSE 7%) and 17.9% (RSE 8%). Four hundred sixty two α 2 -integrin measurements were available from 29 patients. Using the schedule of 100 mg qd (MTD), α 2 -integrin expression was inhibited more strongly than the I int,50 and I int,90 in greater than 95% and greater than 50% of patients, respectively. Moderate inhibition of α 2 -integrin expression corresponded to tumor stasis in mice, and similar levels could be reached in patients with the dose level of 100 mg qd.
The investigational anti-cancer drug E7820 is an orally active angiogenesis inhibitor, and has shown anti-tumor activity in several tumor models in mice, the effect of which was mediated through the inhibition of expression of α 2integrin (1,2). Integrins are receptors that mediate attachment between cells, and between cells and the extracellular matrix (3). They also play an important role in cell signaling. Many types of integrins have been identified, multiple types may be expressed on the cell surface simultaneously. Integrins are heterodimers, consisting of α (alpha) and β (beta) subunits. In mammals, 18 α and 8 β subunits have been characterized, with varying functions related to cell attachment and cell signaling. It has been shown in preclinical experiments that suppression of integrin α 2 by E7820 played a crucial role in inhibition of endothelium tube formation (1,2). E7820 was evaluated in a phase I dose escalation study in patients with malignant solid tumors or lymphomas, the clinical results of which have been reported previously (4) and for which a population PK model was presented before (5). Since α 2integrin is expressed both on tumor cells and on platelets, it has been hypothesized that α 2 -integrin expression on platelets may be an easily evaluable biomarker for tumor growth inhibition in response to treatment with E7820 (1,2).
The inhibition of α 2 -integrin measured on platelets provides a measure of pharmacological target modulation, and therefore would theoretically provide a better predictor of activity than measures of E7820 plasma exposure. α 2 -Integrin is therefore currently being evaluated in phase I studies as possible biomarker (4,6). In this article, we present a modeling and simulation analysis that integrates data from preclinical experiments and a phase I clinical trial to evaluate the expected efficacy of the dose levels that were studies in phase I. Similar analyses have been presented earlier for the development of everolimus (7) and gefitinib (8). The aim of the analysis was to establish levels of integrin inhibition in humans which corresponded to inhibition levels associated with tumor stasis in mice. Therefore, we first aimed to establish what level of α 2 -integrin inhibition corresponds to tumor stasis in mice. This was done by establishing PK-PD models for the preclinical experiments, describing the relationships between E7820 plasma exposure, the inhibition of α 2 -integrin expression, and tumor growth inhibition. Subsequently, we investigated if the dose regimen proposed from the phase I study based on acceptable toxicity would result in sufficient inhibition of α 2 -integrin expression to expect antitumor activity in patients. For this, a model was constructed describing effects of exposure to E7820 on α 2 -integrin expression. This model was based on the human PK model and measurements of α 2 -integrin during the phase I trial. Finally, it was evaluated which clinical regimens were capable of achieving inhibition of α 2 integrin expression in humans similar to that which led to tumor stasis in mice. It was thus assumed that a similar level of α 2 -integrin expression would be required for tumor growth inhibition in mice and humans, for tumors that are sensitive to angiogenesis inhibition mediated by α 2 -integrin.
The PK profile of E7820 in female KSN Slc mice aged 6, 7, or 8 weeks was elucidated after single IV (25 mg/kg, n=3), single oral (25, 50, and 100 mg/kg, n=4), and seven repeated oral administrations with about 12 h intervals (50 mg/kg, n= 4). Concentrations of E7820 in plasma were obtained using a validated liquid chromatography method with UV detection, with precision <15% over the entire concentration range. Mean PK parameters CL, V, and F were calculated by noncompartmental analysis. The tumor growth experiments in mice have been described before (1). Briefly, a human pancreatic carcinoma cell line (KP-1, 5×10 6 cells/head) was transplanted subcutaneously into 7-week old female nude mice. Administration of E7820 was started at doses of 12.5, 25, 50, 100, or 200 mg/kg, or vehicle, 1 week after the transplantation. E7820 was administered orally by gavage, twice a day for 3 weeks. PK samples were collected twice weekly. Blood was withdrawn once a week from the eye of anesthesized mice in PBS containing 0.004% sodium citrate and diluted at 1:100. Diluted blood samples were directly stained with FITC-conjugated anti-integrin α2 Ab and expression levels on platelets were analyzed by flow cytometry. The longest diameter of the tumor and body weight was measured twice a week by direct measurement of the tumor diameters with calipers.
The statistical data analysis and simulations were performed with non-linear mixed-effects modeling using NON-MEM, version VI, level 2.0 (Icon Development Solutions, Ellicott City, MD, USA) with gfortran (http://gcc.gnu.org/ fortran/) as Fortran compiler, and Piraña (9) as modeling environment. The first-order conditional estimation method with interaction (FOCE-I) was used throughout the model building. Standard errors for model parameters were estimated using the covariance step in NONMEM. Model evaluation and selection was based on objective function value (p level of 0.01, ΔOFV = 6.63 was considered a significant improvement in fit), successful convergence, goodness-of-fit plots and visual predictive checks (10). R (version 2.9.0, http://www.cran.r-project.org) was used for the analysis of simulated data, and the generation of plots. Modeling and simulation were performed according to a pre-specified analysis plan, and construction of the PK-PD model was performed sequentially. First, the effect of E7820 plasma concentrations was correlated with the inhibition of α 2integrin expression. Next, a tumor growth model based on the unperturbed tumor growth experiment was fitted. The parameter estimates for the α 2 -integrin model were then fixed, and the model predicted expression levels were used to drive the different tumor growth models that were evaluated.
Longitudinal description of the α 2 -integrin expression level on platelets was modeled as an indirect response model (11) with an inhibitory effect on input rate (k in ) of E7820 plasma concentration, described by a linear equation or by the Hill equation:
in which I describes the relative integrin expression, k in and k out are the input and output rates of the indirect-effect model. The plasma concentration of E7820 was defined as c p , and E max,I , IC 50 and γ were the maximal drug effect, the plasma concentration at half maximal effect, and the shape exponent of the Hill equation describing the drug effect on the α 2 -integrin expression rate, respectively. An inhibitory effect of plasma exposure on k in was considered a mechanistically more plausible model than a stimulatory effect on k out as the pharmacological effect of E7820 is mediated through the inhibition of mRNA expression. α 2 -Integrin expression at baseline (I 0 )was estimated (12). To allow for hysteresis in drug effects on α 2 -integrin, the inclusion of one or more effect compartments was tested. It was also assessed if development of tolerance to E7820 could be shown.
Two prerequisites for the tumor growth model were established: the model should be able to (a) capture the unperturbed growth, and (b) capture the inhibitory effect of decreased α 2 -integrin expression on tumor growth/shrinkage rate, over the entire dose range of E7820. Several tumor growth models were evaluated, including exponential models, Gompertz models (13), and a tumor growth model introduced by Simeone et al. (14). Linear and E max relationships were evaluated for their ability to correlate the effect of inhibition of α 2 -integrin expression to one of the relevant growth rate parameters in the tumor growth models. For both the model for α 2 -integrin and the tumor growth model, it was evaluated if the available data supported the estimation of between subject variability (BSV) of the parameters. For both models, exponential residual error model were used.
Using stochastic simulations from the preclinical PK-PD model, it was investigated what level of inhibition of expression of α 2 -integrin correlated with tumor stasis. Achieving tumor stasis at t=21 days (end of treatment period in preclinical experiments) in 50% or 90% of mice were defined as targets. The inhibition of α 2 -integrin expression (I int,av ) at steady state was chosen as the measure of inhibition of α 2 -integrin expression required to achieve these goals. The simulations were performed for dose levels varying from 50 to 200 mg/kg with 20-mg/kg intervals, and each simulations was performed for 500 mice. In each simulation, the percentage of mice having tumor sizes less than or equal to their baseline tumor sizes was recorded, along with the I inh,av . From the results of the simulations, relationships between dose level, α 2 -integrin inhibition and tumor growth inhibition at t=21 were plotted, and the I int,av values at which 50% and 90% of mice achieved tumor stasis were calculated (I inh,90 and I inh,50 ). To account for uncertainty in model parameter estimates, and to calculate relative standard errors (RSE) for these targets, these simulations were repeated 250 times with parameters sampled from the variance-covariance matrix obtained for the model.
E7820 was administered to patients daily for 28 days, followed by a washout period of 7 days prior to starting subsequent cycles. Blood samples were collected from 1 up to 9 treatment cycles. Bioanalysis of PK was performed using the same method as used for the preclinical samples. Blood samples for α 2 -integrin expression were determined from blood collected at predose, and at 6 h post-drug administration on day 1, predose on day 28 of cycle 1 and at 24 h post-drug administration. Beyond cycle 1 and for each subsequent cycle, blood was collected on day 28 predose. Diluted blood samples were directly stained with FITCconjugated anti-integrin α2 Ab and expression levels on platelets were analyzed by flow cytometry, and expressed in molecules of equivalent soluble fluorochrome.
Previously, a population PK model was constructed from data from the phase I trial (Table I) (5). This model was used here as well, however, to avoid difficulties arising from incorporation of multiple dosing in the turnover absorption model, the absorption model was simplified to a first-order absorption model. The empirical Bayesian estimates (EBEs) for oral clearance (CL/F) and volume of distribution (V/F), and typical parameter values for ka obtained from this PK model and the observed data were used to drive the PD model for α 2 -integrin expression. Within-subject variability in PK parameters was disregarded in the PD analysis. For the clinical data, a similar indirect-effect model for the inhibition of α 2 -integrin expression was evaluated as was used for the preclinical data. It was evaluated if the available data supported the estimation of BSV on the PD parameters. Dosing regimens of 50, 70, 100, and 200 mg qd were then simulated using the developed model to investigate the expected inhibition of α 2 -integrin expression during continuous treatment at these levels. These simulations were performed in 1,000 patients, to account for inter individual variability in response to E7820. The relative α 2 -integrin inhibition at steady state was calculated (two cycles), which was compared to the I inh,av that was required for tumor stasis calculated from the preclinical data.
In the PK experiments in mice, elimination half-life was 0.63 h and a low plasma clearance (CL=7.85 mL min -1 kg -1 ) and low volume of distribution (V=0.684 L/kg) after IV dosing of E7820 were observed. The observed PK profiles after IV administration were monophasic, and a linear increase of AUC was observed from 25 to 100 mg/kg after a single oral dose. As the absorption was rapid with a t max at ∼1 h, an absorption rate (ka) of 3 h -1 was used throughout. Plasma protein binding of E7820 was concentration independent from 5 to 100 mg/L, and was 89% (range 88.0-90.2%). Mean bioavailability after oral administration of 25 mg/kg was 61.3%. E7820 showed anti-tumor activity at doses of 50, 100, and 200 mg/kg in the tumor growth and α 2integrin expression experiments.
The exposure-response relationship between E7820 plasma concentration and the inhibition of input rate of the turnover model was significantly better described using an E max -equation than when using a linear relationship. An E max value of 1.1 (RSE 36%) was estimated, but was not significantly different from 1. Likely, due to the fact that α 2integrin expression was inhibited maximally by only 30%, a sigmoidal E max model could not be estimated. Therefore, the
Table I. Characteristics of Clinical Phase I Study (4) Value (range) Total number of patients 36 Dose level 10 mg 3 20 mg 4 40 mg 3 70 mg 3 100 mg 17 200 mg 6 Sex Male 20 Female 16 Race Caucasian 31 Black 1 Hispanic 4 Weight (kg) 70.9 (43.2-113.6) Height (cm) 167.3 (150.5-182.9) BSA (m 2 ) 1.805 (1.375-2.402) Age (years) 64.8 (40.0-82.0) Continuous variables given as mean (range) and categorical data as counts
shape parameter γ in the Hill equation was fixed to 1, reducing the exposure-response relationship to a nonsigmoid E max model. The inclusion of delay in the effects of drug exposure on α 2 -integrin expression did not significantly improve model fit, nor did the inclusion of a term describing resistance to the inhibition of α 2 -integrin expression by E7820. Figure 1 shows a visual predictive check of the model describing α 2 -integrin expression, which indicated adequate model fit.
Although all evaluated tumor growth models were able to describe unperturbed growth, only when an exponential model was used could the effects of integrin inhibition be captured adequately over the entire dose range studied in the preclinical experiments. In the exponential model, it was however necessary to include an adjustment for the observed initial slow tumor growth, since almost no tumor growth or even tumor shrinkage in some mice (receiving only vehicle) was observed in the first week. This was captured by including a mono-exponential term describing an initially complete and gradually decreasing resistance to tumor growth, on top of the basic exponential tumor growth differential equation:
with T describing the tumor size in mm, α describing the tumor growth rate, β being the rate of decline of initial resistance to maximal unperturbed tumor growth, and t being the time in days after start of treatment. The inhibitory effect of decreased α 2 -integrin expression on tumor growth was incorporated as a sigmoidal E max function, implemented relative to current tumor size:
E max,T , II 50 and γ are the maximal effect of inhibition of α 2 -integrin expression, the plasma concentration at half maximal effect, and the Hill coefficient, respectively. The Hill coefficient could not be estimated, and was therefore fixed at 5 to allow sigmoidicity in the relationship. This provided significantly better fit than using the non-sigmoidal form. The inclusion of a delay between the effect of α 2 -integrin inhibition on tumor growth did not result in a significantly better model fit. All parameters could be estimated with reasonable precision, except the parameters describing unperturbed growth. The high uncertainty (>100%) in estimation of these parameters could be explained by the fact that these estimates were obtained by fitting only the tumor growth data from the five mice that received vehicle and no drug. A visual predictive check, shown in Fig. 2, however showed that the model described the observed data accurately for these mice, and also for the other mice, when the effects of decreased α 2 -integrin expression were incorporated in the model.
A schematic representation of the structural model describing the tumor growth inhibition and α 2 -integrin expression is shown in Fig. 3, and parameter estimates for the PK-PD model are presented in Tables II, III, and IV. The tumor growth model was able to describe the tumor growth in the mice with adequate precision for all dose levels, as can be seen in the plots in Fig. 2 which show both the observed tumor growth curves and the 80% of model predictions (i.e.
between-mice variation in tumor size, and inhibition of tumor growth rates). The plots reveal that tumor growth was considerably inhibited from doses of 50 mg/kg upward, but that tumor stasis (relative to baseline) at t=21 was only achieved at doses of 100 and 200 mg/kg. In Fig. 4, the relationships between dose level, inhibition of α 2 -integrin expression, and tumor growth inhibition are plotted, which were calculated from simulated tumor growth experiments. The inhibition of α 2 -integrin expression at steady state that is required for tumor stasis in 50% and 90% of mice (I inh,50 and I inh,90 ) were calculated to be 14.7% (RSE 7%)and 17.9% (RSE 8%), respectively.
In total, 462 α 2 -integrin level measurements at 209 unique timepoints were available from 29 patients with solid tumors or lymphoma. The same turnover model that was used in the preclinical experiments was used to describe the α 2 -integrin profiles observed in the clinical trial. It was observed that the estimate for the turnover rate parameter (k out ) in the clinical model (0.099) was fairly similar to the estimate for the preclinical model (0.143), suggesting that rates of integrin inhibition are comparable, and only slightly faster in the preclinical setting. The estimate for IC 50 value obtained for the clinical model was about fourfold higher than obtained for the preclinical model. Although BSV in baseline levels and drug sensitivity were high, significant effects of plasma drug exposure on the inhibition of α 2 -integrin expression were observed, and an E max model best described this relationship. The data allowed the estimation of BSV in baseline α 2 -integrin expression levels and drug sensitivity (IC 50 ). A delay in drug effect did not improve the model, nor did a model that incorporated development of tolerance to E7820. Figure 5 shows a visual predictive check of the model and the data, in which dose levels of less than or equal to 50 mg qd, and dose levels of greater than or equal to 70 mg were grouped. This figure shows that the observed profiles correspond to the median model predicted profiles, and that most of the observed profiles are contained within the 90% confidence interval of the model predictions.
In Fig. 6, the results of simulations for the four highest dose levels that were studied in the phase I trial are plotted, along with the steady state level of α 2 -integrin expression inhibition corresponding to tumor stasis in 50% and 90% of mice. In Table V, the average inhibition of α 2 -integrin expression (I int,av ) after one cycle of E7820 of 28 days in patients is shown. These results demonstrate that a regimen of 50 mg qd is expected to result in less than 50% of patients achieving the defined target of both I inh,50 and I inh,90 . At continuous dosing at 70 mg qd, inhibition of α 2 -integrin expression is expected to achieve the I inh,50 target in >50% of patients, but the median expected integrin level is slightly lower than the I inh,90 . At doses of 100 mg qd, median inhibition is higher than both targets, although the 95% confidence interval was not entirely higher than the I inh,90 target. Only at the highest evaluated but toxic dose of 200 mg qd are 95% of patients expected to achieve inhibition of expression higher than both targets. This dose level has however been shown to induce non-tolerable hematological toxicity (15).
We have presented here a first population PK-PD analysis of preclinical and clinical data of the α 2 -integrin inhibitor E7820, and evaluated relationships between exposure to E7820, α 2 -integrin expression and tumor growth. Integrin inhibitors have been introduced recently as potential new treatment modalities in anti-cancer treatment that inhibit tumor growth by inhibiting angiogenesis. Integrins are heterodimer transmembrane receptors that are essential factors in angiogenesis as they mediate endothelial cell migration, proliferation and survival (16). Preclinical studies have shown that the inhibition of integrins, either by
Table II. PK Model Parameters Based on the Preclinical and Clinical Studies Parameter Preclinical (mice) Clinical PK Estimates (IB) Estimates(5) RSE CL a Clearance 0.471 L h -1 kg -1 6.24 L h -1 5% V a Volume of distribution 0.684 L kg -1 66.0 L 8% F Oral bioavailability 69.3 % ka Absorption rate 3 b h -1 0.889 h -1 7% ω CL BSV in clearance (CV) -51 % 12% ω V BSV in volume of distribution (CV) -48 % 17% ρ CL∼V Correlation CL∼V (CV) -62 % 12% a Clinical estimates for clearance and volume of distribution are given as CL/F and V/F b Absorption rate unknown, but assumed rapid (3 h -1 ) in preclinical experiments (t max was observed at ∼1 h) BSV between subject variation, RSE Relative standard error (%), CV coefficient of variation (%) Table III. PD Model Parameters for Integrin Model Based on the Preclinical and Clinical Studies Parameter Preclinical (mice) Clinical PD: α 2 -integrin model Estimates RSE Estimates RSE I base Baseline integrin level 22.5 % (1%) 8,350 MESF 1% K in Input rate turnover model 3.21 %day -1 (19%) 825.6 MESF day -1 14% K out b Output rate turnover model 0.143 day -1 0.099 day -1 E max,C Maximal effect of E7820 exposure on integrin inhibition 1 a 1 a IC 50 Concentration at 50% of maximal effect 656 ng ml -1 (18%) 2,840 ng ml -1 32% ω I,base BSV in integrin baseline 31 % (40%) 40.0 % 19% ω eff BSV in sensitivity to E7820 --110 % 32% σ exp Exponential residual error (CV) 6.7 % (8%) 11.5 % 23% MESF molecules of equivalent soluble fluorochrome a Fixed b Not estimated, but calculated as I base /K in monoclonal antibodies, peptide drugs, or small molecules can result in tumor remission (16,17). Currently, integrin inhibitors are investigational drugs and are in clinical and preclinical testing as single agents or in combination with other anti-cancer agents.
From experiments in mice, we were able to model with good accuracy the effects of drug exposure on levels of α 2integrin inhibition and tumor growth suppression in mice. The inhibition of α 2 -integrin expression levels was described by an indirect-effect model, which is the most widely used pharmacodynamic model for describing physiological processes. The model was able to describe the decrease in α 2integrin expression over time in response to treatment with E7820 and account for the between-mice variation in drug response. A simple exponential tumor growth model was then used to describe tumor growth profiles in mice; with the predicted α 2 -integrin expression levels driving growth inhibition A Gompertz model (13) provided worse fit than the empirical exponential model. Although tumor growth models have been presented in literature that is somewhat more physiologically based than the exponential model (14,18), this model could also not describe our tumor growth better than the exponential model. More importantly, as the aim of this study was not to scale or extrapolate the tumor growth model to other species or other drugs, but primarily in accurately describing the relationship between α 2 -integrin expression and tumor growth, the exponential model was deemed most adequate. It should be stressed that the term (1 À e ÀbÁt ) included in the model does not describe drug resistance but rather initial slow tumor to growth, and that the rate parameter β was estimated solely based on the unperturbed growth curves. The initial slow growth rate may be due to the transplanted tumor not being fully embedded yet in the tumor environment, and thus not fully able to receive nutrients and oxygen needed for growth. Although administration of E7820 started 7 days after transplantation of the tumor, this period may thus not be long enough to achieve the full potential tumor growth rate. It was noticed that t 1/2 was much lower in mice than in humans (1.01 h versus 7.33 h). Additionally, relatively higher doses were administered to
Table IV. PD Model Parameters for the Tumor Growth Model Based on the Preclinical and Clinical Studies Parameter Preclinical (mice) PD: Tumor growth model Estimates RSE T base Baseline tumor size 5.14 mm 1% α Maximal exponential growth rate 0.0903 day -1 103% β Rate of initial resistance to tumor growth 0.0391 day -1 130% E max,I Maximal effect of integrin inhibition on tumor growth 0.0472 day -1 5% II 50 Relative inhibition of integrin expression at 50% of maximal effect 11.4 % 6% ω T,base BSV in baseline tumor size 6 % 13% ω II50 BSV in sensitivity to integrin expression inhibition 20 % 50% σ exp Exponential residual error (CV) 6 % 6% mice than humans (maximum doses 200 mg/kg in mice vs 200 mg in humans). Therefore, much higher C max concentrations were reached in mice, although E7820 was cleared faster. This may be important: apparently higher concentrations were required in mice compared to humans to elicit a response in integrin expression. However, inhibition of integrin expression in patients was already observed at the MTD of 100 mg, and therefore higher doses were not required to achieve pharmacological activity.
Using simulations from the combined preclinical model for E7820 exposure, α 2 -integrin and tumor growth, targets levels of α 2 -integrin inhibition that are expected to lead to tumor stasis in 50% and 90% of mice, I inh,50 and I inh,90 , were estimated. The PK-PD model describing the relationship between E7820 exposure and α 2 -integrin expression in patients was then used to evaluate if these targets could be met with the regimens that were investigated in the clinical trial. This analysis showed that only at the highest toxic dose 0.0 1.0 2.0 3.0 -50 -40 -30 -20 -10 0 50 mg qd Time (days) Time (days) Time (days) Time (days) Integrin expression (%) 0.0 1.0 2.0 3.0 70 mg qd 0.0 1.0 2.0 3.0 100 mg qd 0.0 1.0 2.0 3.0 200 mg qd Fig. 6. Expected α 2 -integrin expression profiles (relative to baseline, without residual variability) during four continuous dosing regimens (qd). Solid lines indicate median expected α 2 -integrin expression levels, gray areas indicate the 90% model prediction intervals. Dotted and dashed lines indicate the I inh,50 , and I inh,90 (the average daily α 2 -integrin expression inhibition needed for 50% and 90% of mice to reach tumor stasis, respectively)
of 200 mg qd, both targets were met in more than 95% of patients. However, a continuous daily dose of 100 mg would achieve the lower target (I inh,50 ) in more than 95% of patients, and more than 50% of patients would achieve the defined high target. These results suggest that, for tumors that are susceptible to reduced α 2 -integrin expression, the proposed dose of 100 mg qd is likely to achieve sufficient inhibition of α 2 -integrin expression to achieve tumor growth inhibition. Based on this analysis, doses of 50 mg qd or lower are however unlikely to achieve efficacy. This analysis provides evidence to support the notion that relative α 2 -integrin expression as measured on platelets may serve as a clinical biomarker for treatment efficacy. Considerable variation in baseline α 2 -integrin expression was observed between patients, thereby diminishing the predictive performance of absolute expression levels. However, limited experiments in healthy volunteers have shown that the expression is constant over time (1) and analysis of the data from the clinical study shows that expression remained constant in patients receiving low doses of E7820. A significant effect of drug exposure on the inhibition of α 2 -integrin expression was found. This suggests that the level of α 2 -integrin expression as measured on platelets relative to the baseline level could serve as a biomarker to assess target modulation in response to treatment with E7820. Since the α 2 -integrin inhibition targets obtained from the preclinical experiments were low (<20%), this suggest that even moderate inhibition of α 2 -integrin expression on platelets is correlated with tumor growth inhibition.
In the development of new drugs, results from preclinical efficacy and toxicity experiments are generally only used to serve as input for the early clinical stage, i.e., to provide a safe starting dose. Once in the clinical stage, results from preclinical efficacy experiments, i.e., tumor growth experiments, are often not taken into account. In this analysis, in contrast, we present an integrated model-based analysis of all data currently available on E7820: data obtained from preclinical experiments were projected onto data obtained from a clinical trial. As data from the preclinical stage may hold valuable information on the in vivo drug response, especially when combined with data on a suitable pharmacodynamic biomarker, such an analysis may allow an early assessment of efficacy.
It is currently unknown how different tumor types will respond to E7820 exposure, and to the inhibition of integrin expression. It is also not clear yet which tumor types respond best to this compound. Likely, not all tumor types will be sensitive to the inhibition of integrin, and therefore the value of integrin expression as a biomarker must be assessed in more detail in subsequent studies. The preclinical PD model was built using data obtained using a pancreatic cell line. If more preclinical data would become available, the model could be updated with other cell-lines as well, to provide a more robust model, or a model tailored to specific tumor types. Another important consideration with this analysis is the role that integrins play in mouse compared to humans. In the preclinical experiments, integrins on platelets were of mouse origin, while those on tumor were of human origin. How this impacts the results is unclear, and calls for caution in interpretation of the results. Another consideration is that tumor growth rates and integrin turnover rates may differ between mice and humans. Additionally, as in patients, larger, necrotic, and more variably sized tumors are encountered and blood supply to tumors is likely to differ between the two species. Further clinical investigation is therefore needed to confirm the results of this analysis. Additional biomarker and tumor growth collected from a homogeneous group of patients will allow refinement of the current model, and will establish validity of α 2 -integrin as a biomarker for tumor growth inhibition. However, in our opinion, the analysis presented here is the most informative analysis that is possible at the current stage of development of this drug, as it not only takes into account the clinical results but also results from the preclinical phase.
C almodulin (CaM) and neurogranin (Ng) are two abundant proteins in the brain whose interactions are implicated in the enhancement of synaptic responses and cognition. 1À4 CaM is a highly conserved protein in all eukaryotic cells and has been shown to participate in a broad range of cellular functions. 5,6 Ng, on the other hand, is expressed predominantly in the mammalian forebrain, where its level is highest in the hippocampus, neocortex, and amygdala. 7,8 This small molecular weight protein (7.5 kDa) has no known enzymatic activity, and CaM is believed to be the sole Ng binding partner in vivo based on the yeast two-hybrid screening. 9 In vitro, Ng binds CaM and their interactions are weakened by increasing Ca 2þ concentration and by covalent modifications of Ng by phosphorylation and oxidation. 10À14 In the acute hippocampal slices, Ng enhances the high frequency stimulation-induced longterm potentiation (LTP) in the CA1 region through elevation of the neurotransmitter-mediated Ca 2þ transients. 4 Interactions of CaM and Ng have been shown to be sensitive to Ca 2þ concentrations in vitro; 10,14 however, the effect of changing intracellular Ca 2þ , [Ca 2þ ] i , on their interactions in neurons has not been investigated. Since a reduction of Ca 2þ is expected to affect the interactions of CaM and Ng, we sought the "low-calcium" model of epilepsy in hippocampal slices 15À17 to investigate the trafficking of these two proteins in CA1 pyramidal neurons. This experimental model retains many characteristics of focal hippocampal seizures in vivo. 18 In the brain, extracellular Ca 2þ , [Ca 2þ ] o , undergoes dynamic changes that depend on neural activity; for example, [Ca 2þ ] o decreases significantly during and after intense neuronal activity associated with seizures. 19À21 Reduction of the [Ca 2þ ] o blocks synaptic transmission and gradually developed epileptiform activity. 16,17,22À24 Infusion of Ca 2þ chelator, EGTA, into the hippocampus of rat generated epileptic activities, and these responses were reversed by Ca 2þcontaining fluid. 25 In this study, we employed immunochemical staining and confocal microscopy to investigate cellular and subcellular localizations of CaM and Ng in hippocampal slices in response to changing [Ca 2þ ] o . These tissue slices were bathed alternately in the Ca 2þ -containing and Ca 2þ -free buffers for monitoring their effects on the synaptic responses, [Ca 2þ ] i , and trafficking of CaM and Ng. Our results revealed that CaM and Ng were more concentrated in the soma and less abundant in the dendrites of hippocampal CA1 pyramidal neurons under basal conditions. In particular, we found that a high level of CaM was sequestered in the nucleus. Perfusion of the tissue slices with EGTA-containing buffer caused a suppression of synaptic response, induction of epileptic activity, and a slow reduction of intracellular Ca 2þ . This treatment also caused translocation of CaM and Ng from soma to dendrites. The Ca 2þ -sensitive mobilization of CaM was most prominent for CA1 pyramidal neurons and was less pronounced among those in the neighboring CA2 and CA3 areas.
ABSTRACT: Calmodulin (CaM) and neurogranin (Ng) are two abundant neuronal proteins whose interactions are implicated in the regulation of synaptic responses and plasticity. We employed the "low-calcium" model of epilepsy in hippocampal slices to investigate the mobilization of these two proteins in CA1 pyramidal neurons. Perfusion of mouse hippocampal slices with Ca 2þ -free artificial CSF (ACSF) caused a suppression of synaptic transmission and generation of epileptic activity; these responses could be reversed by normal Ca 2þ -containing ACSF. Fluorescence immunochemical staining of control hippocampal slices bathed in normal ACSF revealed that CaM and Ng were more concentrated in soma than in dendrites; especially for CaM, it was concentrated in the nucleus. Perfusion of hippocampal slices with Ca 2þ -free ACSF caused translocation of these two proteins from soma to dendrites, and this trafficking was also reversed by Ca 2þ -containing buffer. A reduction of ∼15 and 40 nM intracellular Ca 2þ , [Ca 2þ ] i , caused half-maximum translocation of Ng and CaM, respectively. Hippocampal CA1 pyramidal neurons were the most responsive to this Ca 2þ -sensitive translocation as compared to those from other areas of the hippocampus. These results illustrated the unique feature of hippocampal CA1 pyramidal neurons in sequestering high concentrations of CaM and Ng in soma and releasing them to distal dendrites at reducing level of [Ca 2þ ] i . KEYWORDS: Calmodulin, neurogranin, calcium, translocation, hippocampus, epileptic activity ' RESULTS AND DISCUSSION Ca 2þ -Sensitive Mobilization of CaM and Ng between Soma and Dendrites. Hippocampal slices bathed in normal ACSF have been used routinely for monitoring the synaptic responses. First, we investigated the cellular and subcellular localization of CaM and Ng of these tissue slices by fluorescence immunochemical staining. The immunoreactivity (IR) of these two proteins appeared to colocalize in CA1 pyramidal neurons (Py) (Figure 1). For CaM, it was largely concentrated in the nucleus, whereas its staining intensity in dendrites within stratum radiatum (Sr) and stratum oriens (So) was very weak (Figure 1A and E). For Ng, its localization in the nucleus, cytoplasm, and apical dendrite was evident but the staining pattern in basal dendrite within So was patchy (Figure 1B and F). Both proteins exhibited strong colocalization in nucleus as seen in the merged images with DAPI staining (Figure 1C and D).
We then tested whether bathing hippocampal slices in Ca 2þfree ACSF had any effect on neuronal activity by monitoring the evoked field excitatory postsynaptic potential (fEPSP) in the CA1 region. Perfusion of EGTA-containing ACSF (EGTA/ ACSF) caused a suppression of the slope of fEPSP, and this response induced by either a short-term (5 min) or long-term (20 min) exposure could be completely reversed by switching back to Ca 2þ -containing ACSF (Figure 2). These results indicate that a reduction of extracellular Ca 2þ by perfusion of EGTA/ACSF up to 20 min does not cause any obvious ill effect to the neurons.
Immunochemical staining patterns of the hippocampal slices exposed to EGTA/ACSF for different lengths of time showed that there was a progressive increase in the IRs of CaM and Ng in dendrites with a concomitant reduction of those in the cell layers (Figure 3). This response was most prominent for the CA1 pyramidal neurons and less so among those in other regions of the hippocampus. For CaM, the increase in dendritic IR was most evident for those slices exposed to EGTA/ACSF for 15 and 20 min (Figure 3D and E). It was most striking that after 20 min of exposure to EGTA/ACSF the CaM IR of CA1 pyramidal cell layer appeared even less intense than that of the dendrites (Figure 3E). For Ng, the increase in dendritic IR was discernible after 5 min of exposure (Figure 3B) and also became more prominent after 15 and 20 min (Figure 3D and E). Re-perfusion of the EGTA/ACSF-treated slices with normal Ca 2þ -containing ACSF restored the original localization patterns of these two proteins (Figure 3F). These findings indicate that CaM and Ng are uniquely responsive to the Ca 2þ -mediated mobilization in CA1 pyramidal neurons.
Distinct Differences among the CA1 and Neighboring CA2 and CA3 Neurons. In hippocampus, Ng and CaM were colocalized in all principle neurons at the various subfields. Under the basal conditions, these proteins exhibited higher concentrations in soma than in dendrites with the exception of pyramidal neurons in subiculum complex, which exhibited a uniform distribution of both Ng and CaM in all cellular compartments (data not shown). Treatment of the slices with EGTA/ACSF caused exit of CaM and Ng from soma to dendrites of CA1 pyramidal neurons; however, this trafficking was less apparent for those neurons residing in CA2 and CA3 subfields (Figure 4). For these latter two neuronal populations, a reduction of [Ca 2þ ] o caused an exit of CaM from nucleus to cytoplasm and proximal dendrites, but much less to distal dendrites (Figure 4E and F). Thus, a prominent somatic colocalization of CaM and Ng was observed in the merged images (Figure 4F). In these EGTA/ACSF-treated tissues, the neuronal populations between CA1 and CA2, and also those of CA3, could be clearly distinguished.
Region. The fluorescence intensity of doubly stained sections was measured extending from soma to distal dendrites (Figure 5). In the control (panel A), the peak staining intensities of CaM (red) and Ng (green) resided in the cell layer and their average intensities in dendrites were approximately 10 and 30% of the value of soma, respectively. After 20 min of exposure to EGTA/ACSF (panel B), the somatic CaM and Ng IRs were greatly reduced compared to the control. Translocation of Ng from soma to dendrites could be detected 5 min after exposure to EGTA/ACSF, and there was a delay for the movement of CaM, which was apparent 10 min after the treatment. After 15 min, the CaM IR spread evenly in the entire dendritic field. Re-perfusion with normal ACSF restored the original cellular distribution patterns of these two proteins (panel C).
EGTA-Induced Changes in the Intracellular Calcium. Hippocampal slices loaded with Fluo-4AM Ca 2þ indicator were imaged by two-photon excitation generated by a titanium:sapphire (Ti:Sa) pulsed infrared laser, which allowed for penetration into tissue slices. Bathing the tissue slices with ACSF resulted in a slow decay of neuronal Fluo-4 fluorescence intensity in spite of the measures to reduce laser transmission and shorten the exposure time during image acquisition. Thus, a control decay curve was established before treating the slices with EGTA/ ACSF (Figure 6). In these time-series experiments, images of the CA1 neurons were collected for 5 min in the presence of normal ACSF, then switched to EGTA/ACSF for 15 min, and finally perfused with EGTA/ACSF þ 10 μM 4-Br-A23187 (Figure 6B). Five minutes after exposure to EGTA (at 10 min time point), the [Ca 2þ ] i was reduced by ∼30% of the control, and after 10 and 15 min exposure the reductions were ∼35 and 40%, respectively. Upon exposure to EGTA/ACSF þ Ca 2þ ionophore for 10 min (at 30 min) the [Ca 2þ ] i was ∼10% of the control. Addition of calcium ionophore in the presence of EGTA/ACSF was used for estimation of the approximate maximal reduction of [Ca 2þ ] i in the neurons. Since the [Ca 2þ ] i of hippocampal neurons in normal Ca 2þ -containing buffer is ∼100 nM 26À28 and the K d of Fluo-4/Ca 2þ is 1 μM in situ, the net reduction in the Fluo-4 fluorescence from the basal level could be used as an approximate measure of the reduction in [Ca 2þ ] i . The observed mobilization of Ng reached near maximum after 5 min of exposure to EGTA/ ACSF (see Figure 5C); we estimated that a half-maximum response would occur at a reduction of ∼15% of Fluo-4 fluorescence, namely, ∼15 nM [Ca 2þ ] i . A similar estimate for the half-maximum mobilization of CaM, after 15 min exposure to EGTA/ACSF, was a reduction of ∼40 nM of [Ca 2þ ] i . The EGTA-mediated reduction in [Ca 2þ ] i could be reversed by re-perfusion with Ca 2þ -containing buffer (Figure 6A). Transient Reduction in Intracellular Ca 2þ Induced Epileptic Activity. Bathing the hippocampal slices in Ca 2þ -free buffer caused an early phase of decline in [Ca 2þ ] i followed by slow changes after 5 min. Within this early time frame, the evoked CA1 fEPSP declined rapidly while the amplitude of POPS rose initially and was followed by a decline in a delayed fashion as compared to that of fEPSP (Figure 7). The waveforms of both fEPSP and POPS during the early phase of decline in [Ca 2þ ] i exhibited characteristic epileptic activity, which consisted of multiple bursts following each electrical stimulation (Figure 7A). These epileptic responses lasted for a few minutes, and synaptic responses became silent afterward. The emergence of epileptic activity was an indication of increasing excitability resulting from complex responses to alteration of Ca 2þ homeostasis.
Perfusion of mouse hippocampal slices with low levels of [Ca 2þ ] o (e0.2 mM) or EGTA-containing solution blocks synaptic transmission of pyramidal neurons and results in the development of seizurelike activity in the CA1 region. 15À17 Reduction of [Ca 2þ ] o induces neuronal hyperexcitability caused by several potential mechanisms, including reduction of the surface-charge screening, Ca 2þ -activated K þ current, synaptic GABAergic inhibition, as well as increasing field effects and gap junctions. 23,29 However, the extent of reduction in [Ca 2þ ] i in this model system has not been investigated. To our surprise, perfusion of the tissue slices with EGTA/ACSF only caused a minor reduction of [Ca 2þ ] i , and this may explain why the perfusion with this solution for 20 min did not cause any obvious deleterious effect of the neurons. Under the basal conditions, [Ca 2þ ] i is maintained at ∼100 nM while the extracellular Ca 2þ concentration is significantly higher at 1À2 mM. Intracellular Ca 2þ buffer, regulators of Ca 2þ dynamics, and efflux pathways are responsible for maintaining such a low [Ca 2þ ] i inside the cell. 30 Fine tune regulation of each of these components prevents a large increase in [Ca 2þ ] i that could cause cell death as in the cases of ischemia-or epilepsy-induced neurotoxicity. These regulatory components are also likely to play a role in preventing excessive loss of [Ca 2þ ] i at low [Ca 2þ ] o . The reduction in [Ca 2þ ] i together with the increase in neuronal excitability may trigger the redistribution of Ng and CaM from soma to distal dendrites in CA1 pyramidal neurons. Recently, we have also observed that high frequency stimulation-mediated induction of LTP in the CA1 region also triggered translocation of both Ng and CaM from soma to dendrites.
Studies with hippocampal neurons in cultures showed that stimulation-induced Ca 2þ entry through L-type Ca 2þ channels and NMDA receptors caused translocation of CaM from cytoplasm to nucleus. 31 Another study indicated that unless pretreatment of hippocampal neuron cultures with Ca 2þ -free buffer containing 2 mM EGTA, elevation of [Ca 2þ ] i by glutamate or NMDA was not effective to cause nuclear translocation of CaM. 32 In these EGTAtreated neurons, EGFP-CaM, visualized by confocal imaging, appeared to localize more in the cytoplasm and proximal dendrites than in the nucleus. This observation agreed with the results shown here for the EGTA/ACSF-treated acute hippocampal slices, in which CaM was redistributed to cytoplasm and dendrites of CA1 pyramidal neurons. Re-perfusion with normal ACSF induced Ca 2þdependent translocation of CaM from cytoplasm and dendrites to the nucleus. Thus, the direction of CaM trafficking may be dependent on its basal levels at different cellular compartments. In the acute hippocampal slices, CaM is already concentrated in the nucleus of CA1 pyramidal neurons, and a perturbation of its resting state causes its exit from nucleus to cytoplasm and dendrites. In this "lowcalcium" model of epilepsy in hippocampal slices, translocation of Ng and CaM from soma to dendrites may contribute to the development of epileptiform activity of CA1 pyramidal neurons.
The net movement of somatic Ng to dendrites was not as pronounced as that of nuclear CaM, which, eventually, became equal to or even slightly less than its concentration in the dendrites. Mobilization of CaM from nucleus to distal dendrites could be caused by the dissociation of CaM from its nuclear binding components, changes in the permeability of the nuclear pore complex for free diffusion, and/or transported by a carrier. It is also possible that an initial exit of Ng from the nucleus facilitates the translocation of CaM. Pyramidal neurons in the hippocampal CA1 region are the most sensitive to the low-calcium induced epileptiform activity. 22 This neuronal population also exhibited a higher sensitivity to the low-calcium mediated translocation of Ng and CaM from soma to dendrites as compared to other neurons in the hippocampus, such as those in the neighboring CA2 and CA3. Morphologically, CA2 neurons are similar to those of the CA3 in size but they are not innervated by the mossy fibers from the dentate gyrus. Molecularly, CA2 neurons express some distinct genes from the neighboring field 33À35 and they, CA3 neurons too, are uniquely spared in Alzheimer's disease. 36 CA2 neurons are also resistant to temporal lobe epilepsy 37 and are resistant to plastic changes by conventional protocols that induce LTP and LTD in the CA1 region. 38 Much of these characteristics of the CA2 neurons are likely related to their differences from the CA1 neurons in handling Ca 2þ -mediated responses (Figure 4). The boundary between CA1 and CA2 is not easily defined because mixed populations of cells are present. However, the current study clearly distinguished the CA1 from the CA2 neurons following treatment of the tissue slices with EGTA/ACSF.
All procedures for using animals were approved by the National Institute of Child Health and Human Development Animal Care and Use Committee. Mice (C57BL\6) were housed in standard cages on a 12-h light/dark cycle and provided a normal chow and water ad libitum. Antibody against Ng (Ab#2641) was raised in rabbit against the C-terminal sequence (GARGGAGGGPSGD) of the protein. The following materials were obtained from the indicated sources: mouse anti-CaM from Zymed Laboratories (South San Francisco, CA); ImmPRESS peroxidase-conjugated anti-mouse and anti-rabbit IgG, horse serum, and Vectashield from Vector Laboratories (Burlingame, CA); Fluo-4 a.m., pluronic acid, Alexa Fluor 594 carboxylic acid succinimidyl ester, and 5-(and-6)-carboxyfluorescein (FITC) succinimidyl ester from Invitrogen (Carlsbad, CA); and tyramine hydrochloride, sodium borate, and hydrogen peroxide-urea adduct tablet from Sigma-Aldrich (St. Louis, MO).
Preparation and Treatment of Hippocampal Slices. Transverse hippocampal slices (400 μm) were kept for recovery after slicing for ∼2 h in oxygenated (95% O 2 /5% CO 2 ) ACSF containing the following (in mM): 124 NaCl, 4.9 KCl, 1.3 MgSO 4 , 2.5 CaCl 2 , 1.2 KH 2 PO 4 , 25.6 NaHCO 3 , and 10 D-glucose, pH 7.4. The slices were submerged in a chamber superfused with oxygenated ACSF at a flow rate of ∼2 mL/min. Glass electrodes (1À4 MΩ) filled with ACSF were used both for stimulation of Schaffer collateral/commissural fibers and for recording of fEPSP from stratum radiatum and the amplitude of POPS from the cell layer of the CA1 area. The slope of fEPSP was calculated for indexing the synaptic response. After maintaining a stabled baseline at a current that gave ∼1/3 of the maximal fEPSP response for at least 20 min, the slice was perfused with Ca 2þ -free ACSF that contained 2.5 mM EGTA for a timed period as stated in the figure legend. Afterward, the perfusate was switched back to Ca 2þ -containing ACSF. At the end of incubation, the tissue slice was quickly removed from the recording chamber and fixed in 4% paraformaldehyde in PBS. Synaptic responses were amplified with an AxoClamp 2B or Multiclamp 700B apparatus, digitized by CED power 1401, and analyzed by Signal 4 software (Cambridge Electronic Design).
Each fixed tissue slice was further sectioned with a vibratome into 50 μm thickness and placed in a 96-well plate for free-float staining. Tissue sections were treated sequentially with PBS containing 0.5% NP-40 and 3% H 2 O 2 each for a minimum of 15 min and blocked with 2.5% horse serum. Tissues were incubated with a mixture of primary antibodies (rabbit anti-Ng Ab#2641, 1/1250; mouse anti-CaM, 1/500) in PBS containing 0.025% horse serum and 0.1% thimerosal overnight at room temperature. CaM was first revealed by incubation with ImmPRESS peroxidase-conjugated anti-mouse IgG for 4 h and Alexa 594-tyramine/H 2 O 2 solution (one hydrogen peroxideurea adduct tablet and stock Alexa 594-tyramine in 10 mL of 50 mM Tris-Cl buffer, pH 8.0) at room temperature for 10 min; the reaction was terminated with 10 mM HCl for 30 min. Then, Ng was revealed by incubation with ImmPRESS anti-rabbit IgG for 4 h and FITC-tryamine/ H 2 O 2 for 10 min, and then terminated with 10 mM HCl. In between incubations, tissues were washed 10 min each with TTBS (20 mM Tris-Cl, pH 7.5, containing 0.5 M NaCl and 0.05% Tween 20) once and PBS twice. Tissues were placed on glass slides, spread with Vectashield containing DAPI, and covered with glass slip. Alexa 594-tyramine was synthesized by incubation of 1 mg of carboxylic acid succinimidyl ester of the dye in 12 μL of DMSO with 4 μL of 0.3 M tyramine hydrochloride and 40 μL of 0.1 M sodium borate buffer, pH 8.6, at room temperature in a dark vessel overnight. FITC-tyramine was similarly prepared by incubation of 10 mg of carboxylic acid succinimidyl ester of the dye in 70 μL of DMSO with 70 μL of 0.3 M tyramine and 15 μLof 0.2 M sodium borate buffer, pH 8.6.
Stained tissue sections were examined in a Zeiss LSM 510 inverted microscope with a pinhole setting for red channel (543 nm) at 1 air unit, and the pin holes of other channels (488 and 405 nm) were optimized to achieve the same optical slice depth with each objective before scanning. Images were captured in 8-bit mode (256 scales) at a high resolution (1024 Â 1024 pixels). For quantification of fluorescent intensity, the acquired multichannel images were analyzed off-line using LSM 510 Live analysis software that measured the profiles along a straight line stretching from distal dendrites to the cell layer. The averaged fluorescence intensity of the dendrites was compared to that of the soma. The ratio of dendritic fluorescent intensity versus that of the soma in the same 10Â image was used as a measure for the extent of translocation.
Hippocampal slices from 14 ( 3 day old mice were loaded with Ca 2þ -indicator, Fluo-4AM. (10 μM in 0.02% pluronic acid), by incubation in oxygenated ACSF at 30°f or 90 min. Images were acquired by using a Zeiss LSM 510 two-photon system fitted with a WPAPO 20Â 1.0 objective. Tissue slices were perfused with oxygenated ACSF (∼2 mL/min), changes in Fluo-4 signal were monitored by excitation at 805 nm, and emitting fluorescence was detected through a BP 500À550 filter. Images were collected at 1 min intervals with the laser transmission kept at a minimum while the detector gain was kept at 70À80% of the maximal sensitivity. Changes in the fluorescent intensity of individual cells were quantified using the LSM 510 Live analysis software based on ROI in the time-series experiments. The intensity of the first image was set as 100%, and the net changes at each time point following treatment were estimated against the control incubated with ACSF alone. The results were expressed as mean ( SE.
dx.doi.org/10.1021/cn200003f |ACS Chem. Neurosci. 2011, 2, 223-230
We thank
CaM, calmodulin; Ng, neurogranin; ACSF, artificial cerebral spinal fluid; [Ca 2þ ] o , extracellular calcium; [Ca 2þ ] i , intracellular calcium; IR, immunoreactivity; DAPI, 4 0 ,6 0 -diamidino-2-phenylindole; LTP, long-term potentiation; fEPSP, field excitatory postsynaptic potential; POPS, population spike.
This work was supported by the
K.-P.H. designed and performed experiments, analyzed data, and wrote the paper. F.L.H. designed and performed experiments, analyzed data, and wrote the paper.
We explored patterns of sexual risk behavior among esquineros, heterosexually-identified, sociallymarginalized Peruvian men using latent class analysis. We used data from the Peru site of the National Institute of Mental Health (NIMH) Collaborative HIV/STD Prevention Trial which included n = 2,109 heterosexually-identified men. The latent class analysis used seven risk behaviors to group esquineros into risk classes. We identified four latent classes, of which two classes had lower probabilities and two classes had higher probabilities of these risk behaviors. Comparing the two lower risk classes to the two higher risk classes yielded significantly more unprotected sex acts (Chi square P value \ 0.001). The risk behaviors in two of the latent classes identified were primarily related to alcohol and drug use. Future HIV/STI prevention interventions may benefit from this information by tailoring messages to fit the observed risk patterns and should focus on drug and alcohol use.
Human Immunodeficiency Virus (HIV) and sexually transmitted infections (STIs) are an important source of morbidity, and can be associated with HIV mortality among young people in the developing world [1]. In resource limited settings, establishing interventions with persons at high-risk for HIV/STI infection can be difficult. In Peru, most of the resources are rightly focused on men who have sex with men, the population most affected by both HIV and STIs [2] and on female sex workers. Other populations, including low-income, socially marginalized, young men have been overlooked and deserve greater attention as they have higher risk for HIV/STI infection than the general population [3].
As behavioral theory and HIV/STI prevention interventions become more specified, understanding the patterns of risk behavior within a population may help to effectively tailor interventions. The concept of audience segmentation posits that a population of interest be composed of different sub-groups and that the message for each subset is more effective if properly tailored to their needs and interests [4,5]. In HIV prevention interventions, the message and manor of intervention delivery also need to be tailored to the sub-group of interest based on patterns of risk behavior in that population. Previous studies using this strategy have based their segmentation on HIV risk and perceived risk [5,6].
In this analysis, we look at segmenting a population of high-risk heterosexually identified men, esquineros, based on their risk behaviors for HIV and STI transmission and acquisition. From ethnographic analysis, we understand that the common connection among the esquineros is their marginalization from society. They are not in school, they do not have traditional jobs, and they are often involved in gangs. Although they may have respect from their peers, other esquineros, in their communities they are often seen as troublemakers, 'vagos' or vagrants [7]. They have been described previously as being the victims of structural violence, whose vulnerability stems from their lack of education, opportunities, with limited advantages or ability to take advantage of the limited opportunities that may present themselves [7]. Their higher risk for HIV/STI may also stem from other behaviors such as drug and alcohol use prior to sex, having sex for compensation, and having multiple partners, among other risk behaviors. Additionally, they report more sex with men than other populations [8]. These behaviors increase risk for HIV/STI among the esquineros [3]; however, how these behaviors relate to one another is not known, it is only clear that these behaviors do not occur uniformly in this population. Therefore, understanding the pattern of these behaviors would be important information to have in order to tailor interventions to specific risk groups.
This population was identified ethnographically and their increased vulnerability to STI/HIV has been documented and is higher than that of the general population of young men in Peru [3]; however, HIV/STI prevention strategies for this sub-population have not been devised. The only intervention to date was unsuccessful [9,10], in part this lack of success may have stemmed from an inadequate adaptation of the intervention to this population [9]. From ethnographic and previous epidemiologic data, we know that HIV risk was not a primary concern of this population and high-risk behavior was ubiquitous, but varied [3,7,8]. Given the ethnographic identification and evidence of varied risk behaviors latent class analysis was devised as a method to better understand the patterns of risk behavior within this population in order to improve subsequent prevention interventions with them.
There is renewed interest in behavioral prevention for both HIV and STIs [11]. However, behavioral prevention efforts have primarily focused on the promotion of condom use and have focused their efforts either on very high-risk sub-populations (men who have sex with men (MSM), commercial sex workers (CSW), etc) or on women. In Peru, several strategies have been tried with general population targets including training pharmacists, internet outreach to MSM, and mobile HIV testing campaigns [12][13][14]. However, these campaigns have not been targeted to the behaviors and needs of specific high-risk sub-populations. No interventions to date have focused on young, male, high-risk individuals who are not exclusively MSM.
As the epidemic moves into its third decade in Peru, a more focused understanding of groups peripheral to MSM is needed. Understanding patterns of risk behavior can help to design more tailored interventions to identified at-risk subpopulations. This paper investigates patterns of risk behaviors among a large population of young, low-income men enrolled in the National Institute of Mental Health (NIMH) Collaborative HIV/STD Prevention Trial to determine what behavioral patterns exist in this population. The patterns of risk identified may help to improve targeted and tailored HIV/STI prevention interventions for this population.
The data for this paper are taken from the baseline assessment of esquineros a group of socially marginalized men enrolled in the Peru site of the NIMH HIV/STD Collaborative Prevention Trial. The trial has been described elsewhere in detail [15]. Briefly, populations at highrisk for HIV and STI infection from five countries were recruited in areas of high social interaction to participate in a behavioral HIV prevention intervention based on diffusion of innovations [16] and past successes with this strategy for HIV prevention [17][18][19]. In Peru, the participants included esquineros and women as well as gayidentified men [3]. The study in Peru was conducted in three coastal cities and peri-urban areas including Lima, the capital of Peru, and Trujillo and Chiclayo, mid-sized cities in northern Peru. Low-income urban and peri-urban communities were chosen based on an index of unmet basic needs and information from key informants regarding the presence of the three sub-populations of interest at venues of social interaction. Project recruiters conducted a census of potential participants belonging to these three sub-populations and enrolled those who were willing to participant in the 2-year trial. Subjects were eligible if they were between 18 and 40 years old, spent time at the sites of social interaction at least twice a week, planned to stay in the study area for the duration of the study (2 years), reported sex in the past 6 months, and provided voluntary informed consent to participate in the trial. The study was approved by the Committee of Human Subjects Research of the University of California at Los Angeles, the San Francisco Department of Public Health, the Universidad Peruana Cayetano Heredia, Johns Hopkins Bloomberg School of Public Health, and the U.S. Naval Medical Research Center in Bethesda, MD in compliance with all federal regulations regarding the protection of human subjects.
This analysis is limited to the esquineros who participated in the NIMH trial. In the ethnographic formative research that was conducted prior to the trial, this population was referred to as 'esquineros' (corner men) or 'vagos' (vagrants) [7]. These are young men who are un-or underemployed and congregate in areas of social exchange in their neighborhoods, such as street corners, soccer fields, parks, bars, etc. Many are members of local gangs that conduct the local drug trade. Although these men are primarily heterosexually-identified, they often have sex with gay men or transvestites from their communities, sometimes in exchange for money or goods such as clothing, haircuts, or alcohol and drugs [7,8].
The data in this paper come from the baseline assessment for the NIMH trial, which was conducted between 2003 and 2005 in three coastal Peruvian cities. The baseline assessment included informed consent, a behavioral interview, pre-test counseling, specimen collection and subsequent laboratory testing for HIV (EIA: BioRad and Biomerieux; WB: BioRad), Herpes Simplex Virus Type 2 (HSV-2) (Herpes Select, Focus Technologies, Cypress, CA) using the cutoff optical density of 3.21. Syphilis testing was conducted using (RPR Biomerieux, Boxtel, The Netherlands) with Treponema pallidum Particle Agglutination Testing (TPPA) confirmation (Fujirebio Diagnostics Inc, Toyko, Japan), gonorrhea, and Chlamydia (CT/NG Amplicor PCR, Roche Diagnostics, Branchburg, NJ, USA).
All data collection procedures were conducted in temporary project offices in the communities under study. The behavioral interview was conducted by trained personnel in Spanish, using Computer Assisted Personal Interview (CAPI) technology, where the interviewer read the questions to the participant and entered their responses into a laptop computer. The interview included questions regarding socio-demographics, health care seeking, sexual behavior in the past 6 months, and detailed behavioral questions regarding the last 5 sex partners, as well as the use of alcohol and drugs. Laboratory results were given to the participants at a subsequent visit approximately 2 weeks after the initial assessment, where participants received personalized post-test counseling, treatment for curable STIs and referral for HIV care, as needed.
Seven variables from the behavioral assessment were used in latent class analysis (LCA). Latent class analysis categorizes binary items into like classes based on the underlying assumption of a common latent variable [20]. LCA is similar to cluster analysis, with the goal of grouping like individuals; however, LCA has the assumption of an underlying latent variable. The outcomes of LCA are the number of latent classes, the probability of each item in each latent class, and posterior probability of latent class membership for each individual. The last outcome classifies individuals by their risk behaviors. In this analysis, the underlying latent variable is risk that could lead to unprotected sex and subsequently STI/HIV infection. The seven items included in the LCA were reporting three or more sex partners in the past 6 months (SEXPART3), having concurrent sex partners in the past 6 months, defined as having at least 2 partners with a temporal overlap in the past 6 months (CONCUR), drug use in the past 30 days (marijuana or cocaine paste) (DRUGUSE), drugs use prior to sex with any of the past 5 sex partners (DRUGSEX), using alcohol prior to sex with any of the past 5 sex partners (ALSEX), reporting a male sex partner among any of the past 5 sex partners (MALE), and reporting sex in exchange for goods or services in the past 6 month (MONEY). Throughout this report, these variables will be referred to using the abbreviated form in parentheses. These items were chosen as they have been shown to be related to unprotected sex in past epidemiologic studies with this population [8].
For the latent class analysis, we fit models with increasing numbers of latent classes specified until reaching the lowest Bayesian Information Criteria (BIC) value. BIC has been favored by researchers given its reliance on both the loglikelihood and the adjusted sample size [21]. Once the number of latent classes was established, to validate the classes, the posterior probabilities for the latent class membership for each participant were exported to Stata to determine the relationship between probable class membership and unprotected sex as well as prevalent STI infection. Entropy, which measures how well each individual fits into a specific class, was also taken into account to determine the fit of the model. The validation of the classes was approached comparing unprotected sex with all and with non-primary partners only reported by members of each class as well as STI prevalence in each class. To assess the stability of the LCA results, the LCA was conducted with random samples of half of the data [22]. The latent class analysis was conducted in MPLUS Version 5 (L. Muthe ´n and Muthe ´n, 1998Muthe ´n, -2006)); data management and descriptive statistics were conducted in Stata 9 (College Station, TX).
A total of 2,146 esquineros participated in the baseline assessment for this trial. Of these 2,109 (98.3%) were used for the analysis as they had complete information on the 7 risk variables of interest. The median age of the men was 22 years with an interquartile range of 20-26, the majority of the men are single (66.7%), approximately half have graduated from high school (49.2%), few report having stable work (22.1%), and 35.6% report not having sufficient food at least once a week or once a month (see Table 1). Table 2 shows the prevalence of each risk behavior of interest in this sample. Using alcohol prior to sex was the most common risk behavior (67.0%) and reporting a male partner among one of the last 5 sex partners was the least common (12.6%). The Cronbach's alpha among these variables is 0.58 showing only moderate internal consistency, suggesting that there is variability in individual responses.
The number of possible patterns given the combinations of the seven variables used in the latent class analysis is 2 7 or 128; 100 of these possible patterns appear in the data. Among the participants, 367 (17.4%) had none of the risk behaviors and 10 (0.5%) had all 7 risk behaviors. The majority of the sample lies between these two extremes. While most of the remaining response patterns showed less than 4% of the sample, there were some more frequent patterns. Other than reporting no risk behaviors, the most prominent patterns of risk behavior were only reporting the use of alcohol prior to sex 386 (18.3%), 114 (5.4%) reported alcohol prior to sex and concurrent sex partners, and 97 (4.6%) reported both alcohol prior to sex and drug use in the past 30 days.
The results of the latent class analysis are shown in Table 3. The models were fit with increasing class sizes until reaching the lowest Bayesian Information Criteria (BIC), reached in the 4-class model. Although the entropy for the 5 class model is slightly better than for the 4 class model, they are very comparable (0.838 vs. 0.855, perfect entropy = 1.0) and both show much higher entropy than the models with fewer or greater numbers of latent classes. The 4-class model was selected based on parsimony and as it has the lowest BIC, it will therefore be use for the remainder of the analysis.
The four-class model yields distinct risk profiles based on class; Table 4 shows the probability of each of the seven items by latent class. Classes 3 and 4 each show higher probability of the seven risk behaviors and classes 1 and 2 have lower probability of the seven risk behaviors. The risk behaviors that are much higher in classes 3 and 4 include exchanging sex for money or goods, reporting a male partner in the past 6 months, reporting 3 or more partners in the past 6 months, and reporting concurrent sex partners.
Each class has a different pattern of risk behaviors, the first class primarily has alcohol related risk behavior, this is the most prevalent risk behavior in the sample and is the baseline risk for the population. The second class has this baseline risk and drug related risk behaviors. The third class has the baseline risk plus high-risk sexual behavior (sex with men and sex for money) and the fourth class has the baseline behavior, drug related behaviors and high-risk sexual behaviors (see Table 4). So although there are two higher risk and two lower risk classes, the risks behaviors in each class are distinct. The types of risk behaviors taken by members of each class appear to be related to the types of sex partners they have. Classes 1 and 2 reported a higher number unprotected sex acts in the past 6 months; however, this is related to having sex with a primary partner. The first two classes mainly have sex with primary sex partners, while classes 3 and 4 have many more non-primary sex partners (See Table 5). While most members of each class report having unprotected sex with a non-primary partner, this is much higher among classes 3 and 4 and these classes also report many more of this type of sex act.
The stability of the latent class structure was confirmed in the latent class analysis on random halves of the data. In each random half, the four-class model had the lowest BIC. The class structure of the random halves of the data also remained stable; the probability of each item in the four classes identified in each random half of the data remained similar to the probabilities found in the overall analysis (data not shown).
The differences in the proportion or median of each variable by class membership are shown in Table 5. Classes 1 and 2 showed lower STI prevalence and unprotected sex. For each of the variables related to sex acts and unprotected sex, there are significant differences between the 4 classes. Additionally, if classes 1 and 2 (lower risk) are combined and compared to the combined classes 3 and 4 (higher risk) there are significant differences for the variables related to sex acts and unprotected sex (See Table 5).
The latent class analysis conducted among this population of esquineros from the NIMH Collaborative HIV/STI
Table 4 Probabilities of reporting each item and prevalence of each class based on the 4 class solution of the latent class analysis Item Class 1 Class 2 Class 3 Class 4 Probability Used alcohol prior to sex, past 6 months (ALSEX) 0.534 0.772 0.867 0.880 Used drugs, past 30 days (DRUGUSE) 0.175 0.840 0.000 1.000 Used drugs prior to sex, past 6 months (DRUGSEX) 0.000 1.000 0.106 0.663 Exchanged money for sex, past 6 months (MONEY) 0.055 0.000 0.354 0.371 Reported a male sex partner, last 5 sex partners last 6 months (MALE) 0.042 0.008 0.301 0.257 Reported 3? sex partners, last 6 months (SEXPART3) 0.060 0.000 0.750 0.728 Reported sex partners within the same time frame, last 6 months (CONCUR) 0.134 0.192 0.680 0.477 Prevalence of each class 0.59 0.08 0.20 0.13 Prevention Trial showed different patterns of risk behavior in the population. The most common class among the esquineros (class 1) showed risk behaviors primarily related to the use of alcohol prior to sex. Two classes (classes 2 and 4) showed substantial risk related directly to drug and alcohol use prior to sex and drug use in the past month. This is an important finding as these behaviors lead to risk taking behaviors such as unprotected sex, although they have not been included as an integral part of HIV/STI interventions in the past [9,15]. The fourth class, which showed high probability of risk behaviors related to both substance use as well as sexual risk, was also the second largest class, showing the high level of risk behaviors in this population. The observed patterns of risk behavior show risk within all of the classes of esquineros and that all classes need the attention of HIV/STI interventions, as prevalence was high among all classes. However, the patterns of behavior offer insight into segments of the population. Risk with primary partners is also an important issue to address via intervention. Although all classes reported sex with non-primary partners, the number of acts this was much lower among classes 1 and 2 and most of the unprotected sex reported in these classes appeared to occur with primary partners. This is a key point for future interventions that should address risk within primary partnerships, potentially related to alcohol and drug use prior to sex.
The esquineros in this study display a constellation of risk behaviors, even in the lower risk classes. Multiple problem behaviors have previously been described in adolescent populations in the United States [23], pointing to an underlying factor that explains the correlation between problem behaviors. However, this was based on research in a general population sample; in this study, the esquineros are already a sub-sample of the general population with risky behaviors. Additional research has shown that HIV-related interventions targeted to effect one behavior have an influence on other behaviors as well [24]. However, to achieve this, effective interventions for esquineros are needed.
Given that this population is difficult to recruit, their growing use of cheaply available internet in Peru [25], may indicate that the internet could be an important way to establish contact and to target interventions to this population [26,27]. Internet-based HIV prevention interventions have been successful with similar populations in the past and may be useful to consider in subsequent studies [27].
LCA has been used extensively by market research firms to better design marketing campaigns, this approach can be incorporated into the design of more effective HIV prevention interventions for at-risk populations. Thus far, LCA has primarily been used to tailor alcohol and drug use interventions, however with the growing recognition that for HIV interventions to be effective they should be more focused to the target population [28,29]; this strategy could be timely and highly beneficial.
This analysis has several limitations. The data come from the cross-sectional baseline assessment of a larger trial. Although behaviors tend to be established and stable, given the cross-sectional nature, there is no way to establish if they precede the measure of prevalent STIs. Additionally, the STIs measured include genital herpes, which may have been present years prior the assessment. Although the remaining STIs may also represent long standing infection, this is less likely with bacterial infections that can spontaneously clear or cause symptoms that facilitate medical intervention especially among men [30,31]. Despite these limitations, the study included a large sample of esquineros, the methods used were highly standardized and controlled, and participation in the study was very high.
This analysis serves to better quantify the risk profiles of the esquineros who have been previously identified as at higher risk of STI/HIV than others in their communities. The patterns of risk behavior are important to improve the focus of HIV/STI prevention efforts and to make them more relevant to the lives of these men. With these risk behavior profiles, it becomes clear that focusing on drug and alcohol use especially prior to sex are important risk factors to be focused on in future prevention efforts.
a Interquartile range
Bold values indicate the lowest BIC and entropy, these are both measures of model fit and the lowest value denotes the best model fit
b Excludes viral STIs (HSV-2 and HIV) and includes only syphilis if the RPR C 1:8 c P value from Wilcoxon rank sum
We would like to thank the study participants and study staff of the
After more than a decade of the AIDS epidemic in Thailand, the number of children whose parents are living with HIV or have died from AIDS is increasing significantly and it has been reported that these children are often discriminated against by their peers. In order to better understand the current situation and to explore possible strategies to support HIV-affected children, this study examined children's attitudes towards HIV and AIDS using questionnaires and focus group discussions with children in Grades threeÁsix in five primary schools in a northern province in Thailand. A total of 513 children (274 boys and 239 girls) answered the questionnaire and five focus groups were organised. The findings showed a strong positive correlation between children's belief that HIV could be transmitted through casual contact and their negative attitudes towards their HIV-affected peers. Most children overestimated the risk of HIV transmission through casual contact and this made their attitudes less tolerant and less supportive. After HIV prevention education (which included information on HIV transmission routes) was given in three of the study schools, the same questionnaire and focus groups were repeated and the findings showed that children's attitudes had become more supportive. These findings suggest that HIV prevention education delivered through primary schools in Thailand can be an effective way to help foster a more supportive and inclusive environment and reduce the stigma and discrimination that decrease educational access and attainment for HIV-affected schoolchildren.
It is estimated that in 2007 there were 17.5 million children who had lost one or both parents to AIDS globally (UNICEF, UNAIDS, WHO, & UNFPA, 2009). Thailand is one of the countries hit hardest by the AIDS epidemic in Asia with 610,000 people currently living with HIV. This has resulted in a very large number of children being affected by HIV and AIDS although not actually infected. Such children are affected through living with HIV-positive parents or having parents who have died from AIDS (UNAIDS, 2008). External support has mostly been provided to meet basic needs and provides financial assistance, educational opportunities and counselling services. However, resources are limited, only a small number of children have access to this assistance (UNICEF, UNAIDS, WHO, & UNFPA, 2008), and there is an urgent need to scale up the support for these children who have been made vulnerable by HIV and AIDS.
Children spend most of their day in school and it is one of the most significant communities they belong to apart from their family. It is now recognised that school plays an important role by providing protection and support for children affected by HIV and AIDS (UNICEF et al., 2008). For example, education on literacy, numeracy, vocational skills and other life skills can equip children to better cope with their future lives (Carr-Hill, Katabaro, Katahoire, & Oulai, 2002;International HIV/AIDS Alliance, 2003). Schools are an important entry point for children to receive social welfare services, health services and food. Children can receive informal support from their peers and teachers to recover from their distress at home and to regain the sense of normality (UNICEF et al., 2008). Keeping children in school can provide psychosocial support and help to reduce the risk of HIV infection, exploitation and abuse (Coombe, 2002), but only when the school environment is safe and inclusive (UNICEF et al., 2008).
It is also widely recognised that in many highburden countries HIV-affected children experience discrimination and exclusion in school due to stigma related to HIV and AIDS (Castle, 2004;UNICEF, 2003). Given the importance of positive peer relationships for school-aged children (Hartup, 1992;Jones, 1995), negative peer relationships caused by HIV-related stigma could be a significant barrier to children being able to realise the benefits that schooling can bring. In order to maximise the potential of school to support these children, it is necessary to understand schoolchildren's views about HIV and AIDS, and find out how well HIV-affected children are accepted within their school peer group.
Despite its importance, there has been little research on children's attitudes towards HIV and towards their peers who are affected by HIV and AIDS. The aim of this study was to fill this knowledge gap by examining primary schoolchildren's attitudes towards their HIV-affected peers in Northern Thailand. This study was part of a larger enquiry, which explored the impact of HIV and AIDS on children and the role of the primary school in supporting these children (Ishikawa, 2007).
This study was conducted in the Lamphun province in Northern Thailand, where HIV prevalence reached over 7% for pregnant women at the peak of the epidemic (UNDP, 2004). Three primary schools, which were receiving support from a local NGO for children affected by HIV and AIDS such as outside school activities for these children (i.e., programme school), and two primary schools nearby without the NGO support (i.e., non-programme school) were selected for this study. Children's attitudes towards HIV and AIDS were explored using a questionnaire and focus groups for primary schoolchildren in the third to sixth grade. The questionnaire was developed by selecting relevant questions from existing questionnaires on HIV and AIDS Á including the UNAIDS General Population Survey (2000) and the tool developed by WHO and UNESCO (1994) Á and then revising them to fit the school setting and the situation of HIV-affected children. The questionnaire had 14 questions; 12 of these questions focused on children's attitudes towards HIV-affected children and two questions asked about their attitudes towards people living with HIV in general (see Box 1). The questionnaire included seven knowledge items on casual contact in school settings because studies have shown that misunderstanding or overestimation of the risk of HIV transmission by casual contact contributes to negative attitudes towards people living with HIV (Castle, 2004;Herek, Capitanio, & Widaman, 2002). The total number of correct answers to these seven items provided a knowledge score that was used for further analysis.
The questionnaire was administered at each school to all the children attending the class on the day of the study visit, under the guidance of research assistants. Each question was read-out by an assistant and children wrote down their answers on the questionnaire. In total, 513 children (274 boys and 239 girls; 96.6%) in the third to sixth grade, aged 8Á14 years completed the questionnaire. Among them, 41 children (19 boys and 22 girls; 8.0%) were identified as being affected by HIV and AIDS (i.e., children whose parents were living with HIV or have died of AIDS) by their teachers. Although the HIV status of these children was unknown, most were probably not infected with HIV due to the high coverage of the programme to prevent mother-to-child transmission of HIV in Thailand.
A total of five focus groups were also conducted. Each group had 10 children (five boys and five girls) who were randomly selected after excluding children affected by HIV and AIDS. In these groups, the children discussed about HIV and AIDS, their attitudes towards HIV-affected children and how they could help these children (see Box 2 for the discussion guide).
In order to examine whether or not there were any associations between children's correct knowledge of the risk of HIV transmission through
Box 1. Questions: children's attitudes towards HIV and AIDS. (1) Should people with HIV live far away from other people? (2) If you knew that a food seller had HIV, would you buy food from her/him? (3) If a student is infected with HIV, should she/he be separated from other students? (4) If there is a student with HIV in your class, would you play with her/him? (5) Would you eat lunch with a student with HIV? (6) Would you stay away from a student whose parents are infected with HIV? (7) Would you share a packet of snacks with a student with HIV? (8) Should a student with HIV be allowed to study together with other students? (9) Should a student with HIV use separate plates and glasses from other students? (10) Are you afraid of playing with a student with HIV? (11) If a student with HIV asks for your help, would you be willing to help her/him? (12) Should a student with HIV be allowed to play with other students? (13) Would you drink water from the same glass as a student with HIV? (14) Would you lend your pen to a student with HIV?
casual contact and their attitudes towards their peers affected by HIV and AIDS, information on the risk of HIV transmission in school settings was provided to all children in the programme schools. This information was delivered through a game in which pictures of different ways that children could have casual contact, such as studying together with a student with HIV, were shown and children were asked to say whether there was a high risk, low risk or no risk of HIV transmission. Children who had the highest score for correct knowledge received a small prize. After giving this information through the game, the attitudes of children were examined again by asking them to complete the same questionnaire that had been given to them previously. This procedure was repeated in all five schools.
Quantitative data were analysed using SPSS 17.0 for Windows and transcripts of focus groups were made and organised with the help of the NVivo software. A children's attitude scale was developed and the correlation between their attitudes score and some variables, such as knowledge about HIV and AIDS and past contact with people with HIV, were examined. The findings from these quantitative data were further explored in depth by the analysis of qualitative data from the focus groups.
This study was approved by the ethical committee of the Institute of Education, University of London, and official permission for the study was obtained from the provincial educational authority in Thailand as well as from each participating school. Prior to the questionnaire and the focus groups, children were told that participation in the study was voluntary and they could withdraw from the study at any time.
Children's answers to the attitude items are presented in Table 1. Not all children answered all the questions so that the totals were usually less than 513. When asked if they would be willing to help a student living with HIV, 137 children (64 boys and 73 girls; 29.4%) said they were ''willing to help always'' and 216 children (103 boys and 113 girls; 46.4%) said they ''would help sometimes''.
During the focus groups, there were some positive statements regarding children infected with HIV. For example, a girl in Grade 6 said ''I feel sorry for a student with AIDS'', a boy in Grade 6 said ''I will encourage him'' (a boy in Grade 6) and a girl in Grade 4 said ''I will not bully and tease a student with AIDS, and if other children bully her, I will help her''.
However, 74.7% of the children were against children with HIV studying together with those not infected and boys were more likely to give a negative answer than girls (x 2 023.165; df 02; p B0.001). It was also found that most of the children were afraid of playing with children with HIV. Only 36 children (20 boys and 16 girls; 7.7%) answered that they were not afraid at all of playing with children with HIV, whereas 218 children (102 boys and 116 girls; 46.5%) answered that they were a little afraid, and 215 children (125 boys and 90 girls; 45.8%) answered that they were very afraid.
During the focus groups, children often stated that they would stay away from a student with HIV. Some students, mostly boys, said they would play with students infected with HIV but not go close to her/him:
It is better to stay far from a student with AIDS. I'm afraid to get AIDS. (A boy in Grade 4) I will play with a student with AIDS as usual and try to encourage him. But I will be careful not to get AIDS from him and I will not let him know (that I am trying to be careful). (A boy in Grade 4) This negative attitude also applied to children who were not infected but were affected by HIV and AIDS, although it was less strongly expressed. Only 121 children (74 boys and 47 girls; 25.9%) answered that they would not stay away from children whose parents have HIV, which suggested that more than 70% of children were assuming that if the parents had HIV, their children were infected as well.
More negative attitudes were expressed towards more intimate contact with children infected with HIV. When asked whether they would eat their lunch together with a student with HIV, 38 children (16 boys and 22 girls; 8.1%) answered ''yes''. When asked whether they would use the same glass to drink water with a student with HIV, only 12 children (seven boys and five girls; 2.6%) answered ''yes''. These sentiments were also expressed by children in the focus groups: The results of the focus groups revealed that the children believe that if parents were HIV-positive the child was also infected with HIV. They were sympathetic and compassionate towards these children and no prejudice towards their parents was mentioned. However, at the same time, they were aware of the risk of HIV infection and were trying to protect themselves by avoiding physical contact with them. The results of the questionnaire also showed that the children's negative attitudes were largely influenced by their fear of HIV infection due to their misunderstandings of HIV transmission routes and overestimation of the risk of HIV infection through casual contact. It was also found that more than 80% of the children reported that their parents have forbidden them to play with children affected by HIV and AIDS.
In order to further examine children's attitudes, an attitude scale was developed. The findings of the focus groups with the children indicated that they usually assume that HIV-affected children are HIVpositive. Special attention was therefore paid to their attitudes towards children with HIV. Eleven out of the 14 items (Box 1), which focus on children with HIV, were used to develop a scale as follows:
Step 1: The correlation between each item and the raw sum of all the 11 attitude items were calculated.
Step 2: The 11 items were entered into a principal component analysis and two factors with eigenvalues greater than 1 were extracted; then, correlations between each factor and each attitude item were calculated.
Step 3: The items that had a high correlation both with the raw sum of the 11 items as well as with the two factors were selected, which gave a total of five items and these were used as the components of the children's attitude scale. These five items were:
(1) If a student was infected with HIV, should she/ he be separated from other students? (2) If there is a student with HIV in your class, would you play with her/him? (3) Would you eat lunch with a student with HIV? (4) Should a student with HIV be allowed to study together with other students? (5) Should a student with HIV be allowed to play with other students?
According to the answer, each response was scored 1 (negative attitude) to 3 (positive attitude) and a sum of these five items was defined as the attitude score. A simple sum was used because there was no external rationale for differential weighting. The most tolerant score (i.e., positive attitude) is 15, whereas the least tolerant score (i.e., negative attitude) is 5. Cronbach's a for this scale was 0.876. The findings showed that children's attitude scores ranged from 5 to 15 (median 8). Figure 1 shows the distribution of the scores. This figure shows that the children tend to have negative attitudes and the distribution was positively skewed.
Then, factors that associated with the children's attitude scores were explored using an analysis of variance. Table 2 shows the effect of each variable on the attitude score. It was found that older children had higher attitude scores (F 3, 435 021.32; pB0.001; x 2 00.128), and children who had a higher knowledge score also had more positive attitudes (F 6, 435 06.46; pB0.001; x 2 00.082). Children who answered that their parents always told them not to play with children affected by HIV and AIDS had a more negative attitude than the children whose parents never or sometimes told them so (F 2, 435 041.32; pB0.001; x 2 00.160). It also showed that girls had a higher attitude score than boys (F 1, 435 08.55; p00.004; x 2 00.019). The contact with people with HIV did not affect children's attitude score. In multivariate analysis, children's age, sex, knowledge score and their parents' attitudes (i.e., forbid to play with affected children) remained statistically significant (Table 3). Among these factors, children's grade (t 458 09.112; pB0.001), parents' attitudes (t 458 08.102; pB0.001) and knowledge score (t 458 06.048; pB0.001) showed a large impact on the children's attitude score.
As previously mentioned in Section ''Methods'', information on how HIV can and cannot be transmitted in the school setting was given to children in the three programme schools and the children were asked to complete a post-intervention questionnaire which was the same as the previous one. The focus groups were also repeated. The children in the two non-programme schools were not given the information. In total, 513 children (272 boys and 241 girls; 96.6%) answered the post-intervention questionnaire. Although there was no significant difference in the knowledge score between schools at the time of the pre-intervention questionnaire (Table 4), the knowledge score of children in the programme schools was higher than that of the non-programme schools after they had received the information provided (Mann ÁWhitney test: U09647.0; pB0.001).
Although a slightly higher mean attitude score was observed in the programme schools compared to the non-programme schools from the pre-intervention questionnaire, the attitude score had improved in the course of time for both groups with a more significant improvement in the programme schools (MannÁWhitney test: U012,881.0; pB0.001; Figure 2). After the information on how HIV cannot be transmitted was provided, the children's attitude score improved from 9.0 (median Á pre-questionnaire) to 13.0 (median Á post-questionnaire) in the programme schools. The attitude score also improved in the non-programme schools from 8.0 (median) to 10.0 (median).
Changes of children's attitude were further examined by calculating the reduction in attitude gap (i.e., 100*(15Áattitude post-score)/(15Áattitude pre-score)). Multivariate analysis was run in order to examine the factors affecting changes in children's attitudes. It was shown that the provision of the information on HIV transmission (t 394 05.450; pB0.001), the attitude score at the pre-questionnaire (t 394 04.829; pB0.001) and the knowledge score at the post-questionnaire (t 394 0Á3.004; p00.003) had a significant association with the reduction in attitude gap (Table 5). Children's grade and the attitudes of parents, which had a large impact on the attitude score of the children at the pre-questionnaire, did not have a significant impact on the attitude at the post-questionnaire (grade: t 394 0Á1.219; p00.224; parents' attitudes: t 394 0Á0.402; p00.688).
During the second set of focus groups, more positive statements were heard in the programme schools than in the non-programme schools. AIDS Care 241 I will play with a child with AIDS, because you can't get HIV easily. I'm not afraid. (A boy in Grade 6, programme school) I think we should not tease a student with AIDS. (A girl in Grade 4, programme school)
In contrast, there were still similar statements as the previous focus groups in the non-programme schools. In summary, the children's attitudes changed positively and the changes were more significant in the programme schools, which received information on non-transmission routes of HIV.
This study found that children were feeling sorry for children affected by HIV and AIDS and were willing to help them. Their sympathy and compassion towards affected children mostly originated from the fact that the affected children had parents who are HIV-positive or had lost their parents due to AIDS. This compassion was also due to their belief that these affected children were also infected, through no fault of their own, with HIV. In this study, the prejudice due to their parents' social status, which Clemo (1992) found, was not apparent, and the affected children were regarded as ''innocent victims'' (Busza, 2001;Devine, Plant, & Harrison, 1999). However, it was also found that children assumed that the children ''affected'' by HIV and AIDS were ''infected'' with HIV, thus they were afraid of them. As a result, the affected children were being stigmatised in the school. This corroborates the findings of Herek (1999), who reported that stigma arose from both the fear of HIV infection, as well as from the fatality of the illness.
As past studies suggest (e.g., Boer & Emons, 2004;Castle, 2004), it was found in this present study that misunderstandings about HIV being transmitted through casual contact significantly influenced children's attitudes. Children were aware of the transmissibility of HIV and most of them were overestimating the risk of HIV transmission through casual contact. This misconception seemed to have increased their fear of HIV infection and, as discussed above, led them to believe that if parents are HIV-positive, the children are also infected with HIV, which resulted in their negative attitudes towards these children affected by HIV and AIDS. A significant improvement in their attitude, after receiving correct information on the non-transmissibility of HIV through casual contact in the schools, also supports this finding.
The results have also shown the impact of the parents' attitudes on the children. Most parents were reported to have forbidden their children to play with children affected by HIV and AIDS and this seemed to have led to children having negative attitudes. It was also found that older children and girls had a more positive attitude towards their HIV-affected peers. This may be due to the fact that social skills are developed as children grow and mature, and to the fact that girls are generally more socially competent than boys of the same age. This finding may also explain the fact that boys and younger children affected by AIDS were more likely to be bullied by their peers (Ishikawa, 2007). There were some potential limitations in this study. First, the samples included children affected by HIV and AIDS who may themselves have had more positive attitudes towards children with HIV. However, as the results showed, children's status as affected by HIV and AIDS as well as their past contact with people with HIV did not show significant correlation with their attitudes. Second, the improvement in their attitudes may have been due to the maturation effect as children receive more information and knowledge as they grow, which was shown by the fact that attitudes of children in the non-programme schools have also improved. However, the fact that a more significant change in the attitudes was observed in the programme schools strongly suggests that these changes were not only due to the effect of maturation.
This study has demonstrated that schoolchildren do stigmatise and discriminate against their HIV-affected
60 50 40 30 20 10 0 60 50 40 30 20 10 0 60 50 40 30 20 10 0 60 50 40 30 20 10 0 5 6 7 8 9 10 11 12 13 14 15 n=225 n=213 n=201 n=242 Attitude score (pre) Attitude score (post) No. of children No. of children Programme school Non-programme school No. of children No. of children 5 6 7 8 9 10 11 12 13 14 15 5 6 7 8 9 10 11 12 13 14 15 5 6 7 8 9 10 11 12 13 14 15
The authors would like to acknowledge all the children, caregivers and teachers who took part in and contributed to this study. This study was partly funded by the
Background: To evaluate the efficacy of highly-active antiretroviral therapy (HAART) in individuals taking cytochrome P450 enzyme-inducing antiepileptics (EI-EADs), we evaluated the virologic response to HAART with or without concurrent antiepileptic use.
Methods: Participants in the US Military HIV Natural History Study were included if taking HAART for ≥6 months with concurrent use of EI-AEDs phenytoin, carbamazepine, or phenobarbital for ≥28 days. Virologic outcomes were compared to HAART-treated participants taking AEDs that are not CYP450 enzyme-inducing (NEI-AED group) as well as to a matched group of individuals not taking AEDs (non-AED group). For participants with multiple HAART regimens with AED overlap, the first 3 overlaps were studied.
Results: EI-AED participants (n = 19) had greater virologic failure (62.5%) compared to NEI-AED participants (n = 85; 26.7%) for the first HAART/AED overlap period (OR 4.58 [1.47-14.25]; P = 0.009). Analysis of multiple overlap periods yielded consistent results ]; P = 0.006). Virologic failure was also greater in the EI-AED versus NEI-AED group with multiple HAART/AED overlaps when adjusted for both year of and viral load at HAART initiation ; P = 0.005). Compared to the non-AED group (n = 190), EI-AED participants had greater virologic failure (62.5% vs. 42.5%; P = 0.134), however this result was only significant when adjusted for viral load at HAART initiation (OR 4.30 [1.02-18.07]; P = 0.046).
Conclusions: Consistent with data from pharmacokinetic studies demonstrating that EI-AED use may result in subtherapeutic levels of HAART, EI-AED use is associated with greater risk of virologic failure compared to NEI-AEDs when co-administered with HAART. Concurrent use of EI-AEDs and HAART should be avoided when possible.
Seizure disorders are common in HIV-infected individuals, with an incidence of up to 11% in several cohort studies [1][2][3]. Antiepileptic drugs (AEDs) are frequently prescribed to patients with HIV not only for pre-existing epilepsy, but also in the setting of CNS opportunistic infections and for other neurologic and psychiatric conditions including neuropathy, refractory headaches, depression, and bipolar disorder [4]. Concurrent use of highly-active antiretroviral therapy (HAART) and AEDs has the potential for highly complex and clinically significant drug interactions.
Several first-generation AEDs, including phenytoin, carbamazepine, and phenobarbital are metabolized by the cytochrome P450 (CYP450) enzyme system. This pathway of drug metabolism is also utilized by protease inhibitors (PIs), non-nucleoside reverse transcriptase inhibitors (NNRTIs), and the CCR5 inhibitor maraviroc [4,5]. In addition to competing for enzyme binding sites, some antiretrovirals and AEDs intrinsically induce CYP450 metabolism with the potential to decrease blood levels of both agents. As a consequence, this drug interaction may limit the effectiveness of HAART and predispose patients to adverse HIV treatment outcomes, including virologic failure, accumulation of antiretroviral drug resistance mutations, and HIV disease progression. In turn, lower AED blood levels may also have deleterious effects due to lack of efficacy, including loss of seizure control or inadequate control of neuropathic pain.
In many parts of the world where HIV is highly prevalent, such as sub-Saharan Africa and parts of Asia, only CYP450 enzyme-inducing AEDs (EI-AEDs) are available [6]. Thus, treating both HIV and epilepsy in these areas may lead to suboptimal treatment of both conditions due to the potential for clinically significant drug interactions. The US Department of Health and Human Services (DHHS) guidelines [7] recommend considering alternative anticonvulsants or monitoring drug levels, while WHO guidelines [8] recommend cautious use, drug level monitoring, or avoidance of these combinations. Since there are inadequate clinical data and the virologic efficacy is largely unknown for HAART in the setting of concurrent EI-AED use, we performed a retrospective case control study to determine the virologic outcomes in participants with concomitant HAART/EI-AED use in the US military HIV Natural History Study (NHS).
Participants were identified in the database of over 5,000 patients enrolled in the NHS since 1986, a cohort of consenting military members, retirees, and beneficiaries 18 years or older with HIV [9,10]. Individuals are seen approximately every 6 months at participating United States military treatment facilities. Data are systematically collected, including demographic characteristics, information on medication use, laboratory data, and reports of clinical events with medical record confirmation.
The NHS database was searched for individuals on concurrent HAART and EI-AEDs, which included phenytoin, carbamazepine, and phenobarbital. HAART regimens were PI-or NNRTI-based as previously defined [10]. Regimens consisting only of triple nucleoside reverse transcriptase inhibitors (NRTIs) were excluded due to the lack of drug interactions between EI-AEDs and the NRTI class. Participants included were those taking a PI-or NNRTI-based HAART regimen for ≥6 consecutive months and during this period were taking an EI-AED drug for ≥28 consecutive days. Since participants in the EI-AED group may have taken medications that can affect CYP450 metabolism other than HAART and EI-AEDs, use of rifamycin antibiotics (CYP450 inducers) and azole antifungals (CYP450 inhibitors) were reported.
Two control groups were used for comparisons to the EI-AED group, the first consisted of participants prescribed AEDs that are not CYP450 enzyme-inducing, the NEI-AED group. NEI-AEDs included levetiracetam, lamotrigine, zonisamide, ethosuxamide, topiramate, gabapentin, tiagabine, and pregabalin. Oxcarbamazepine (a weak CYP450 inducer) and valproic acid (a CYP450 inhibitor) were excluded. Since participants in the EI-AED group were prescribed AEDs primarily for seizure treatment and prophylaxis as well as neuropathy, participants in the NEI-AED group were restricted to those taking these drugs for the same indications. Analogous to the EI-AED group, the NEI-AED group was required to have a HAART duration ≥ 6 months and an AED overlap during the HAART period of ≥28 days.
The second control group was a subgroup of all participants on HAART without any AED use in the NHS (non-AED group), matched (10:1) to each EI-AED patient according to year of HAART initiation and number of previous HAART regimens. Participants in this control group must have been on their HAART regimen for ≥ 6 months to be included. Since all cases in the EI-AED group had a documented date of HIV infection prior to the year 2000, all potential controls were limited to those with dates of HIV infection prior to 2000.
Virologic failure was defined as having all plasma viral loads (VLs) in the first 6 months of HAART (minimum of 2 values) ≥400 copies/mL and/or the participant having 2 consecutive VLs ≥400 copies/mL after 6 months of HAART. Other virologic outcomes assessed included the percentage of participants with VL <400 copies/mL at 6 and 12 months of HAART, and the average of VL (log 10 ) for each individual within the HAART period. For virologic suppression at 6 and 12 months, the VL closest to 6 and 12 months after HAART initiation, respectively, were used.
Logistic regression was used to compare the proportion with virologic failure and viral suppression after starting HAART in the EI-AED and each of the comparison groups. For each outcome, the number and percent of participants with the outcome are reported with the relative odds and 95% CI, cases versus controls. Analysis of variance and covariance was used to compare cases and controls for average VLs during HAART. Mean case-control differences are reported ± standard error (SE). Because some individuals in the EI-AED and NEI-AED groups had multiple HAART episodes with concurrent AED use, GEE analyses was used to compare virologic outcomes using multiple HAART/AED overlap episodes (up to the first 3 episodes). For binary outcomes, a logic link was used in the GEE analyses; for continuous outcomes a normal link was used.
Both univariate and multivariate analyses were performed. For the analyses comparing the EI-AED and NEI-AED groups, a model adjusting for year of and VL prior to HAART initiation was performed. For comparison of the EI-AED and non-AED groups, adjustment was made only for prior VL, since cases were matched to controls for year starting HAART. SAS, version 9.2 was used for all analyses; PROC GENMOD was used for the GEE analyses.
Based on inclusion criteria, 19 participants were treated concurrently with EI-AEDs and HAART, with 12, 6, and 1 taking phenytoin, carbamazepine, and phenobarbital for the first HAART/EI-AED overlap period, respectively. EI-AEDs were used for seizure disorder in 17 of 19 participants; further characterization of seizure disorders included CNS toxoplasmosis (2 with seizures, 1 for prophylaxis), herpes simplex meningitis/encephalitis (n = 2), progressive multifocal leukoencephalopathy (n = 1), seizures following a motor vehicle accident (n = 1), and the remainder with seizure disorders without further characterization (n = 10). Two individuals used EI-AEDs for neuropathic pain. For the NEI-AED group, 85 participants met inclusion criteria; 82 received gabapentin, 2 pregabalin, and 1 levetiracetam at the first overlap. The majority of participants were prescribed NEI-AEDs for the indication of neuropathic pain (81/85, 95%), with the remainder for seizure disorder (4/85, 5%). Two participants were prescribed rifampin and 1 participant had a history of itraconazole use, however none of these medications were prescribed during the HAART periods investigated in this study.
Demographic factors including age at HIV diagnosis and race were similar between the EI-AED, NEI-AED and non-AED groups. However, participants in the EI-AED group were younger at first HAART/AED overlap compared to the NEI-AED group (40.1 vs. 45.1 years, P = 0.027; Table 1), reflective of the EI-AED group starting HAART during an earlier calendar year (median 1998 versus 2003). Mean CD4 cell count and log 10 VL at HAART initiation were similar between the 3 groups, although VL levels tended to be higher in the EI-AED compared to the NEI-AED group (3.8 versus 3.1 log 10 copies/mL; P = 0.075). The EI-AED group had a higher proportion of individuals with AIDS-defining events prior to the HAART period analyzed, however this was only significant in comparison to the non-AED group (57.9% vs. 21.1%; P < 0.001). The majority of participants in all groups were HAART treatment experienced, however the EI-AED group had a higher percentage of HAART-naive individuals (21.1%) compared to the NEI-AED group (7.1%; P = 0.01).
In evaluating HAART/AED overlap periods, 7 (36.8%) EI-AED participants had a single period of overlap, while 5 (26.3%) and 7 individuals (36.8%) had 2 or ≥3 overlap periods, respectively. The number of overlaps was similar for the NEI-AED group. The duration of first HAART/AED overlap was no different between the groups, with 7.0 months (range 1.0-96.4) overlap in the EI-AED group compared to 9.1 months (1.3-65.4) for the NEI-AED group (P = 0.231). The groups were also similar when all eligible HAART/AED periods were considered, with 21.3 (1.0-155.4) and 22.1 (1.6-120.3) months of overlap for the EI-AED and NEI-AED groups, respectively (P = 0.798).
In comparing outcomes for the first HAART/AED overlap, virologic failure was significantly greater in the EI-AED group (62.5%) compared to the NEI-AED group (26.7%; P = 0.009; Table 2). The average log 10 VL during the overlap period was also higher in the EI-AED group (3.3 ± 1.3 vs. 2.4 ± 1.2; P = 0.006). The percentage of participants with VL <400 copies/mL was significantly lower in the EI-AED group compared to the NEI-AED group at 6 months (33.3% vs. 71.4%; P = 0.016) and 12 months (36.4% vs. 75%; P = 0.018), respectively.
Results were similar when multiple HAART/AED overlap periods per individual were included in the analyses, which added approximately twice the number of HAART/ AED episodes (Table 2). Virologic failure was more common in HAART episodes for EI-AED (63.3%) compared to NEI-AED individuals (27.9%; P = 0.006) and the average log 10 VL during the overlap period was significantly higher in the EI-AED group (3.3 ± 1.3 vs. 2.5 ± 1.3; P = 0.005). The percentage of participant HAART episodes with VL < 400 copies/mL was also lower in the EI-AED group compared to the NEI-AED group at 6 months (28.6% vs. 69.4%; P = 0.002) and 12 months (39.1% vs. 74%; P = 0.004).
Analysis of virologic outcomes adjusting for year of and VL at HAART initiation yielded similar odds ratios to that in the univariate analyses. This was true for both analyses of the initial overlap period and in using multiple overlaps. However, only for the multiple overlap analyses were the odds ratios significantly different from one. The estimated odds ratio for virologic failure for the EI-AED group compared to the NEI-AED group using multiple episodes was 4.19 (95% CI [1.54-11.44]; P = 0.005).
For the EI-AED group, virologic failure was higher compared to the non-AED group (62.5% vs. 42.5%) in
HIV-infected patients commonly require treatment with AEDs due to neurologic and psychiatric conditions. Drug interactions between EI-AEDs and HAART are highly complex and may result in loss of efficacy for one or both treatments. In examining this interaction retrospectively in a military HIV cohort with free access to healthcare and medications, we found greater virologic failure in individuals taking EI-AEDs compared to NEI-AEDs when used in combination with HAART. Since first line agents for epilepsy in most low and middle income countries are limited to EI-AEDs, the clinical ramifications of HAART/EI-AED drug interactions may be substantial. Despite the widespread use of EI-AEDs and the potential for significant drug interactions with HAART, clinical studies are extremely limited [11]. A randomized, parallel-arm study examined the pharmacokinetic interaction between lopinavir/ritonavir (400 mg/100 mg twice daily) and phenytoin (300 mg daily) in healthy volunteers [12]. In the first arm of 12 participants, the addition of phenytoin reduced the area under the concentration-time curve (AUC) of lopinavir and ritonavir by 33% and 28%, respectively after 12 days of overlap compared to the pre-phenytoin period. Notably, the effect of increased lopinavir clearance secondary to CPY3A4 induction by phenytoin was not offset by the presence of low dose ritonavir used as a "boosting" agent. The second arm of 8 participants showed a 31% reduction in phenytoin AUC after the addition of lopinavir/ritonavir demonstrating a two-way drug interaction between classes. A similar result was shown in a randomized, crossover study of 18 healthy individuals receiving either efavirenz (600 mg daily) or carbamazepine (titrated to 400 mg daily) followed by 14-21 days of overlap with the other drug [13]. Compared to pre-overlap levels, efavirenz AUC and minimum (Cmin) and maximum (Cmax) concentrations were reduced by approximately 17% to 43% while carbamazepine AUC decreased by 27%. Though the majority of drug-drug interactions result in reduced plasma concentrations, carbamazepine toxicity may occur secondary to inhibition of CYP3A4 when used with low dose ritonavir [4,14]. Since most studies were performed in healthy volunteers, extrapolation of these findings to patients with HIV infection and epilepsy is difficult because the clinical implications of these interactions have not been adequately studied.
This is the first study demonstrating clinically meaningful outcomes in participants receiving overlapping treatment with EI-AEDs and HAART. The impact is so robust, that we were able to demonstrate this despite the small number of individuals receiving EI-AEDs. Since it is more difficult to enter military service with pre-existing epilepsy, the overall incidence of epilepsy is low in our cohort. Yet, the close follow-up in this prospective observational cohort makes it uniquely ideal for an assessment of clinical consequences of this interaction. Despite the small number of participants taking EI-AEDs in our study, these agents are still commonly used even in the United States. EI-AEDs are favored by some insurance plans due to their lower cost, so it is likely that a cohort with a higher prevalence of epilepsy would have included more participants on EI-AEDs. It is notable that of the 21 participants diagnosed with a seizure disorder in this study, 17 were taking EI-AEDs.
The comparison of EI-AEDs versus NEI-AEDs combined with HAART in our study showed worse virologic outcomes in the EI-AED group. The inclusion criteria for the NEI-AED group were chosen to best approximate the participants in the EI-AED group, specifically targeting use of NEI-AEDs for the indications of seizure disorder or neuropathic pain. In cases where the specific indication for AED use was known, the majority of individuals were prescribed AEDs in the setting of CNS opportunistic infections. The relatively small number of individuals in the EI-AED group limited the power of the study. This was likely due to the increased availability of newer AEDs that are not CYP450-enzyme inducing over the past decade. Other limitations include the differing proportions of seizure disorders and neuropathic pain in the two groups. As a reflection of this, the drugs in the NEI-AED group are agents commonly used for neuropathic pain. Multivariate analyses were performed in an attempt to minimize some of the differences in HAART period between groups, with results demonstrating worse virologic outcomes in the EI-AED group. Other unmeasured factors included HIV drug resistance, potency of HAART regimens, adherence to both drug classes, absence of ARV and AED blood levels, and the inability to study individual EI-AED and ARV pairings due to small sample size. The EI-AED group also had a higher percentage of AIDS events and higher VL prior to the first AED/HAART overlap compared to NEI-AED group. This suggests that the EI-AED group may have more advanced HIV disease and greater risk of virologic failure. It is important to note, however, that the EI-AED group had less treatment experience than the NEI-AED group, with 42% and 15% having <1 year of HAART experience, respectively. This difference in treatment experience in the EI-AED group may potentially offset the risk of treatment failure posed by having a higher percentage of prior AIDS events.
For comparison with both NEI-AED and non-AED control groups, the EI-AED group had consistently worse virologic outcomes, especially in multivariate analyses. Non-AED individuals fared better than EI-AED participants, however the results were significant only after adjustment for VL at HAART initiation. There are several unmeasured factors that may have contributed to these findings including medication doses and adherence, HIV drug resistance, and other uncharacterized variables unique to patients with seizure disorders or neuropathic pain. Measures of AED efficacy, including seizure control, were not completely captured. Our initial hypothesis was that concurrent HAART/EI-AED use would lead to subtherapeutic blood levels of HAART, elevated VLs, and eventually virologic failure. We chose a minimum HAART/AED overlap period of ≥28 days due to the small number of patients exposed to EI-AEDs for any duration. The median duration for all overlaps was 9 months for the EI-AED group. Since epilepsy and seizure disorders typically require longterm, if not life-long treatment, many patients will be taking EI-AEDs and HAART for extended periods of time. Even though the percentage of participants with virologic failure was high at 63.3%, it is possible that the HAART/EI-AED overlap time was insufficient to develop regimen failure for some individuals and additional failures would occur with continued use of both classes.
The introduction of EI-AEDs in patients with HIV may complicate a regimen that is already subject to other challenges from drug interactions. For example, the burden of tuberculosis in sub-Saharan Africa requires many HIV-infected patients to receive concurrent antituberculous treatment (ATT). The cornerstone of ATT regimens is the rifamycin class of antibiotics. As inducers of CYP450, rifamycins can also enhance the metabolism of AEDs and antiretrovirals, adding further management challenges [15]. Compared to rifampin, rifabutin has less CYP450 induction and is favored for ATT in the setting of HAART. In the current study, no participants were treated for active tuberculosis during the study period.
The World Health Organization's list of essential medicines includes the EI-AEDs carbamazepine, phenytoin, and phenobarbital [16]. In addition to the availability of only EI-AEDs in many areas of the world, the expanding use of AEDs in patients with HIV has made the management of HIV and comorbid conditions challenging in these locations, as well as in more developed countries. For example, distal sensory polyneuropathy (DSP), often treated with AEDs, occurs in up to 57% of patients with HIV and the risk of developing DSP is increased with underlying nutritional deficiencies [17][18][19].
In the setting of concurrent HAART/EI-AED use, the US Department of Health and Human Services (DHHS) guidelines [7] recommend providers consider use of alternative agents and/or monitoring of blood levels of HAART/EI-AEDs. Therapeutic drug monitoring (TDM) of HAART is not routinely recommended for the management of HIV patients. However, DHHS guidelines suggest that TDM may be useful in situations with clinically significant drug-drug interactions that may result in reduced efficacy, including use of certain AEDs. TDM may identify reduced blood levels of HAART as a result of drug-drug interactions prompting the provider to consider increasing the dose of ARVs. However, this may lead to a higher rate of adverse effects and ultimately impact drug tolerance and adherence.
Despite the potential merits of this approach, TDM has several limitations including cost and limited availability. According to a recent Cochrane review [20], TDM trials are generally small and underpowered, have short follow-up time, and poor compliance with TDM recommendations. TDM trials have also been performed in countries with higher income and may not be generalized to resource-limited settings. Since the majority of EI-AED use is in low and middle income countries due to the greater cost of NEI-AEDs, TDM is unlikely to be an option for clinicians in these areas.
In settings where EI-AEDs must be used, it is important to recognize treatment failure early with frequent VL monitoring and clinical assessments for efficacy of HAART and EI-AEDs. HAART regimens composed of the integrase inhibitor raltegravir in combination with 2 NRTIs would enable clinicians to avoid drug-drug interactions, however raltegravir may not be available in many areas. Although there are concerns that a triple NRTI regimen may be less durable compared to other HAART regimens, this may be another reasonable option given the lack of drug interactions between NRTIs and EI-AEDs [21].
For the treatment of epilepsy in the setting of HIV infection, alternative agents to EI-AEDs should be administered if available. Valproic acid is available in many regions, however this drug is a weak inhibitor of CYP450 and may lead to increased HAART levels and toxicity, especially when used with lopinavir/ritonavir [4,22]. Due to its additional property as a non-selective histone deacyltase (HDAC) inhibitor, in vitro studies indicate valproic acid may increase HIV outgrowth from resting CD4+ T cells [23]. However, the in vivo effects and clinical relevance have not been firmly established [24,25]. Despite the potential concerns, valproic acid appears to be a safe alternative to EI-AEDs when used with HAART [26].
EI-AEDs should be avoided in favor of NEI-AEDs in patients requiring concurrent HAART and AED therapy due to the higher potential of virologic failure and reduced efficacy. In areas where EI-AED use cannot be avoided, closer and more frequent monitoring of HIV and seizure control is warranted. Alternative HAART regimens, such as triple NRTIs or integrase-based regimens, and use of TDM may be beneficial in managing or avoiding these complex drug interactions when available. In low to middle income regions such as sub-Saharan Africa and parts of Asia, treatment of HIV and comorbid epilepsy and other neurologic conditions will continue to pose great challenges until additional resources become available, such as NEI-AEDs and a wider repertoire of antiretrovirals.
OR (odds ratio): Odds of virologic event for EI-AED cases versus odds for NEI-AED controls; VL, viral load (copies/mL); * Up to three intervals used per subject
OR (odds ratio): Odds of virologic event for EI-AED cases versus odds for non-AED controls; VL, viral load (copies/mL)
The content of this publication is the sole responsibility of the authors and does not necessarily reflect the views or policies of the
Department of Health and Human Services, the DoD or the Departments of the Army, Navy or Air Force. Mention of trade names, commercial products, or organizations does not imply endorsement by the U.S. Government. This work was presented, in part, at the 18 th Conference on Retroviruses and Opportunistic Infections, Boston, MA, USA. The Infectious Disease Clinical Research Program HIV Working Group includes Mark Kortepeter, Helen Chun, Cathy Decker, Susan Fraser, Joshua Hartzell, Gunther Hsue, Arthur Johnson, Alan Lifson, Grace Macalino, Robert O'Connell, John Powers, Roseanne Ressner, Edmund Tramont, Tyler Warkentian, Paige Waterman, Sheila Peel, Connor Eggleston, Scott Merritt, Susan Banks, Michael Zapor, Brian Agan, Michelle Linfesty, Mary Bavaro, Timothy Whitman, Glenn Wortmann, and Lynn Eberly. Support for this work (IDCRP-000-03) was provided by the Infectious Disease Clinical Research Program (IDCRP), a Department of Defense (DoD) program executed through the Uniformed Services University of the Health Sciences. This project has been funded in whole, or in part, with federal funds from the National Institute of Allergy and Infectious Diseases, National Institutes of Health (NIH), under Inter-Agency Agreement Y1-AI-5072.
Competing interests JFO, GAG, DMS, GLB, AG, ACW, NC, TL, and MLL declare that they have no competing interests. JAF has served on the scientific advisory board of UCB, Johnson & Johnson, Eisai, Novartis, Valeant, Icagen, Intranasal, Sepracor, and Marinus. Dr. French is the president of the Epilepsy Study Consortium that receives funding from multiple pharmaceutical companies. JMG is a consultant for Pfizer.
All authors participated in the design of the study and manuscript preparation. GAG performed the statistical analysis. All authors read and approved the final manuscript.
Infectious Disease Clinical Research Program, Uniformed Services University of the Health Sciences, Bethesda, MD, USA. 2 Infectious Disease Service, Brooke Army Medical Center, San Antonio TX, USA. 3 Division of Biostatistics, University of Minnesota, Minneapolis, MN, USA. 4 NYU Comprehensive Epilepsy Center, New York, NY, USA. 5 Department of Pharmacy Practice and Administration, Philadelphia College of Pharmacy, Philadelphia, PA, USA. 6 Department of Neurology, Mount Sinai School of Medicine, New York, NY, USA. 7 International Neurologic & Psychiatric Epidemiology Program, Michigan State University, East Lansing, MI, USA. 8 Division of Infectious Diseases, National Naval Medical Center, Bethesda, MD, USA. 9 Infectious Disease Service, Walter Reed Army Medical Center, Washington, DC, USA. Infectious Disease Clinic, Naval Medical Center San Diego, San Diego, CA, USA.
In July 2010, WHO published new recommendations on providing antiretroviral therapy to adults and adolescents, including starting ART earlier, usually at a CD4 count of 350 or lower, specific regimens for first-and second-line therapies, and other recommendations. This paper estimates the potential impact and cost of the revised guidelines by first, calculating the number of people that would be in need of antiretroviral therapy (ART) with different eligibility criteria, and second, calculating the costs associated with the potential impact. Results indicate that switching the eligibility criterion from CD4 count < 200 to < 350 increases the need for ART in low-and middle-income countries (country-level) by 50% (range 34% to 70%). The costs of ART programs only to increase coverage to 80% by 2015 would be 44% more (range 29% to 63%) when switching the eligibility criterion to CD4 count < 350. When testing and outreach costs are included, total costs increase by 62%, from US$26.3 billion under the previous eligibility criterion of treating those with CD4 < 200 to US$42.5 billion using the revised eligibility criterion of treating those with CD4 < 350.
In July 2010, the World Health Organization (WHO) published new recommendations on providing antiretroviral therapy (ART) to adults and adolescents in resource-limited settings that revised the guidelines previously published in 2006. The new recommendations encourage starting ART earlier, usually at a CD4 count of 350 or lower, specifies regimens for first and second line therapies, and contains other recommendations regarding laboratory monitoring and other elements [1]. The revised guidelines were developed based on systematic reviews of the evidence, consultation with key stakeholders, and consideration of the impact and cost of potential changes. This paper describes the model and analysis prepared to examine the potential impact and cost of the revised guidelines.
The analysis consists of two parts: first, we construct a model to calculate the number of people that would be in need of ART with different eligibility criteria, in order to calculate the potential impact of the new guidelines. Second, we calculate the costs associated with the potential impact in order to evaluate the financial implications of the new guidelines.
The model tracks the HIV+ population by CD4 count using an approach similar to one used in South Africa recently to estimate the need for treatment (see Figure 1) [2]. The values and sources for all of the parameters described below can be seen in Supplementary Material available online at doi:10.1155/2011/738271 Annex A.
We assume that all newly infected people start with CD4 counts above 500, and that their CD4 counts decline over time. The transition probabilities λ1, λ2, λ3, and λ4 represent the probability of progressing from one CD4 category to the next; the derivation of these probabilities is discussed in detail below. In each category there is some probability of death from HIV-related causes, designated as μ1, μ2, μ3, μ4, and μ5 as well as a chance of death from non-AIDS causes, μ0 (not shown in the figure). The probability of HIV-related death increases as CD4 counts decrease.
The number of people in the different CD4 count categories represents the HIV-infected population that is not on ART. The number of people eligible for treatment is
λ1 λ2 λ3 λ4 New infections CD4 350-500 CD4 250-350 CD4 200-250 μ1 μ2 μ3 μ4 μ5 c1 c2 c3 c 4 c5 FL ART 350-500 FL ART 250-350 FL ART 200-250 s1 s2 α2 s3 α3 s 4 α4 s 5α5 HIV-related death r0 r1 r2 r3 r4 r5 β1 β2 β3 β4 β5 SL ART 350-500 SL ART 250-350 SL ART 200-250 r 1 CD4 >500 CD4 <200 FL ART >500 FL ART <200 SL ART >500 SL ART <200 r r r r 0 10 20 30 40 50 60 70 80 90 100 National Township Health workers Educators Karonga Kenya South Africa South Africa South Africa South Africa Malawi 350-499 200-349 250-349 200-249 (%) Cape town >500 >350 <200 Figure 2: Distribution of HIV+ Population not on ART by CD4 Count.
the number in each CD4 count category that is below the recommended level for initiating ART.
Depending on the eligibility criterion and the level of first-line ART coverage a percentage of those eligible for treatment will start first-line ART (c1, c2, c3, c4, c5). Those on ART are categorized by their CD4 count at the initiation of treatment. The model does not track the temporal decline of CD4 counts of those on treatment. Those on first-line ART have a probability of failure depending on their CD4 count at initiation, α1, α2, α3, α4, and α5.
The number starting ART each year is determined by the assumed coverage and the number of people eligible for treatment. We assume that those starting on ART will be distributed among the eligible CD4 categories such that an equal percentage of people in each eligible CD4 category initiate treatment.
Those failing on first line ART will either start on second line ART (according to second line coverage s1, s2, s3, s4 and s5) or die from HIV-related causes. Those on second line have some probability of dying from HIV-related causes each year (β1, β2, β3, β4, β5).
The number of HIV-related deaths each year is the sum of HIV-related deaths from those not on ART and those on ART.
The historical annual number of new infections is exogenous to the model and is based on a Spectrum projection using historical surveillance and survey data to determine HIV prevalence and incidence trends [3]. The future number of new infections is also based on the Spectrum projection but can be modified by expanding treatment. For those not on ART infectiousness varies by CD4 count (as a result of variations in viral load) as indicated by r1, r2, r3, r4, and r5. Infectiousness is high during primary infection, r0, low during the asymptomatic period (r1, r2, r3, and r4) and high during the symptomatic period, r5. Those on ART have reduced infectiousness, r . As a result the future number of new infections can be influenced by the dynamics of CD4 decline and the coverage of ART.
We have estimated the transition probabilities by fitting the model to data on the distribution of the HIV-infected population by CD4 count and the pattern of progression from HIV infection to AIDS death. Data on the distribution of the HIV-infected populations by CD4 count are available from studies in a township near Johannesburg, South Africa (community-based survey of 1000 men and women aged 15-49 [4]), health care workers in Gauteng, South Africa (all 2032 professional and support staff at two hospitals [5]), educators in South Africa (national survey of 21,669 public school educators from all provinces of South Africa [6]), Cape Town, South Africa (observational cohort from two public sector clinics consisting of 2086 patients [7]), Karonga, Malawi (demographic surveillance site studying all adults aged 18-59 and including about 150 HIV-positive individuals [8]), and Kenya (nationally representative sample of adults 15-64 [9]). The distribution of these populations by CD4 count category is shown in Figure 2. Data are also available from several cohort studies on the overall progression from HIV infection to HIV-related death. The Analysing Longitudinal Population-based HIV/AIDS data on Africa (ALPHA) network has conducted a pooled analysis using data from several cohorts to estimate the proportion surviving by the number of years since infection [10]. Only the Kenya data set is a nationally representative sample, and it is the only one that provides information on all CD4 categories of interest. Thus we have estimated the parameter values using only the Kenya data set, along with the age-adjusted, net survival curve based on the East and Southern Africa cohorts from the ALPHA network, but checked the results against the other data sets.
We fit the model to both sets of data simultaneously. One version of the model was set up for Kenya and used the Spectrum estimates of the number of new infections from 1980 to 2007 and the reported number of people on ART from 2000 to 2007. We compared the data on distribution by CD4 count from the Kenya AIDS Indicator Survey (KAIS) with the model projection for 2007. Another version of the model followed a cohort of 1000 new HIV infections as they progress through the various CD4 categories and to death. The resulting proportions surviving were compared with the ALPHA network survival curve for East and Southern Africa. We searched for the single set of transition probabilities that provided the best fit in both cases. The model used a time step of one-tenth of a year in order to accommodate the short duration in the 200-250 category that could be less than one year. The fits are shown in Figures 3(a) and 3(b). The resulting parameters are shown in Supplementary Annex Table A1.
The fit of the model to Karonga (Malawi) and Orange Farm (South Africa) data sets using the parameter values derived from the fit to the Kenya data and the annual number of new infections in Malawi and South Africa is shown in Figures 4(a) and 4(b).
Four categories of cost are considered: antiretroviral (ARV) drugs, laboratory costs, service delivery costs, and identification (outreach and testing). The cost of ARV drugs is determined from the number of people on first and second line, the distribution of patients by regimen and the costs of each regimen. Following previous work, we examine two sets of alternative regimens: one that contains a fast phase-out of d4T, and another that contains a slower phase-out of d4T [11]. Drug costs may be different for patients in low and middle income countries. Current costs are based on WHO and Clinton Foundation reports (Table 1).
Laboratory costs are calculated separately for new and continuing patients and can vary by regimen. Currently, laboratory costs are calculated as the annual median cost for lab tests across recent literature. Recent studies in various countries (Cote d'Ivoire, Ethiopia, Mexico, Nigeria, South Africa, Thailand, Uganda, Zambia) are used as the basis [12][13][14][15][16][17][18][19][20]. The median cost is $250 per year for new patients and $190 per patient per year for continuing patients.
Service delivery costs are based on a standard number of inpatient days and outpatient visits per patient per year and country specific costs for inpatient days and outpatient visits. For this analysis we used the same studies referenced above for laboratory costs (with the exception of Cote d'Ivoire and the addition of another South Africa study [21]) to calculate the median number of outpatient visits per year as 9.5. Only three of these studies also had data on the number of inpatient days for ART patients [12,14,21]; we used these to calculate the median number of inpatient days for ART patients per year as 1.56. The country-specific costs per inpatient day are the costs of one bed day at a primarylevel hospital as reported in the WHO-CHOICE database of service delivery costs [22]. The cost of an outpatient visit is for a 20-minute outpatient visit at a health centre, from the same WHO database. Representative regional costs are shown in Table 2.
Outreach and testing costs vary primarily by the type of population reached. The model considers 10 population categories for testing: The unit cost of VCT services average about $16 per client. We have used this cost also for provider-initiated testing and counseling. No additional testing and counseling costs are included for pregnant women since the costs of testing and counseling are already covered in the Prevention of Mother-To-Child Transmission (PMTCT) programs. Similarly we assume that outreach and counseling for sex workers, IDU, and MSM are already covered in prevention programs for those populations, and add only $1 for the costs of the test itself. For general population testing we have doubled the personnel costs associated with VCT to allow for additional outreach programs in addition to the testing and counseling costs. The resulting cost is $23 per person tested. The number of tests for each population group will depend on the eligibility criterion and the coverage. We assume that patients with symptoms who are found to be HIV+ will be in the lowest CD4 count category. We assume that those who are found to be HIV+ in the other population groups will be distributed by CD4 count according to the distribution of all HIV+ people excluding those <200.
The model has been applied to all low-and middle-income countries (LMIC) and to seven countries individually: Burkina Faso, Mexico, Nigeria, Russia, Tanzania, Ukraine, and Vietnam. The number of new infections each year and the number of people on ART through 2008 were taken from the Spectrum projections for each country. Estimates of the population sizes are based on national estimates prepared as part of the effort to estimate global resource needs [23].
Table 3 displays the results for LMIC for the additional cost and impact of scaling up ART coverage to reach 80% by 2015, assuming that the criterion for eligibility to treatment switches from a CD4 count 200 to 350 in 2010. In addition, the financial implications of two ARV regimens are presented; first with a slower phase-out of d4T, and second with a fast phase-out of d4T. In order to compare across countries and across scenarios, all cost and impact figures are discounted to 2010 using an annual discount rate of 3 percent.
Our estimates suggest that switching the eligibility criterion from CD4 count <200 to <350 increases the number of person-years of ART from 40.7 million to over 61 million, a 50% increase. There is a concomitant reduction in the number of AIDS deaths, with the number decreasing by 21% when the eligibility criterion switches to CD4 count <350. The number of new HIV infections is also reduced, due to the lower infectivity that occurs when people receive ART; new HIV infections are reduced by 11% when the eligibility criterion changes.
The financial costs of providing ART to meet the new need from increasing the eligibility criterion also increase in a similar way to the increase displayed in the number of person-years of ART. The costs of providing ART would be 44% higher with a switch to providing ART to those with CD4 count <350. Although the overall costs are higher with the fast phase-out of d4T relative to the slower phase-out of d4T, the difference is quite small.
Note that the slightly lower percentage increase in costs versus the number of person-years on ART reflects the relatively greater numbers of people on first-line therapy with the increase in eligibility criterion. When the additional testing costs incurred in order to identify the new patients are included, however, total costs increase relatively more than the number of person-years of ART. Total costs increase from US$26.3 billion (US$27.0 billion) to US$42.5 billion (US$43.5 billion) if the eligibility criterion is CD4 count <350 and there is a fast (slower) phase-out of d4T, an increase of 62% (61%). Combining the results for incremental costs and deaths averted suggests that the cost per AIDS death averted is approximately US$9,700 if the eligibility criterion switches to CD4 count <350.
In order to perform a sensitivity analysis, we vary the costs of laboratory testing and service delivery costs using the interquartile distribution of laboratory testing costs from the studies cited above. Using the first quartile function result, laboratory and service delivery costs are reduced by 31%, while using the third quartile function result increases laboratory and service delivery costs by 64%. Overall, this translates to a range in total costs (not presented here) of US$42.5 billion to US$55.5 billion for the scenario with a slower phase-out of d4T, and a range in total costs of US$37.2 billion to US$56.5 billion for a fast phase-out of d4T. In order to compare results for different epidemic types and different regions, we performed the analysis for seven countries: Burkina Faso, Mexico, Nigeria, Russia, Tanzania, Ukraine, and Vietnam. Results indicate that there is not a great deal of variation across countries (see Figure 5). While the average percentage increase in the number of person-years on ART for LMIC was 50% when the eligibility criterion switched from CD4 count <200 to <350, this varies across countries from a low increase of 34% in Burkina Faso to a high increase of 70% in Vietnam. A similar pattern can be observed for AIDS deaths; the country level results range from a reduction of 16% in Burkina Faso to a reduction of 23% in Vietnam when the eligibility criterion switches to CD4 count <350. Finally, the changes in the country-level additional ART costs associated with changing the eligibility criterion mirror the changes in the results for LMIC; for LMIC, the additional ART costs increase by 44% when the eligibility criterion switches to CD4 count <350, while the increases at the country level vary from 30% (Burkina Faso) to 63% (Vietnam).
In this paper, we model both the impact and cost of the new 2010 WHO recommendations for providing antiretroviral therapy to adults and adolescents in resource-limited settings. We examine the impact of changing the eligibility criterion for antiretroviral therapy from CD4 count <200 to CD4 count <350 on the number of person-years on ART, the number of AIDS deaths averted, and the costs of the change including the costs of additional tests and recruitment costs. We also examine the financial impact of switching away from d4T towards other recommended regimens.
We find that, although the total costs for providing ART increase, the percentage increase is slightly less than the increase in number of person-years on ART. The number of person-years on ART increases for LMIC by 50%, varying between 34% and 70% at the country level, while the cost of providing ART increases by 44% for LMIC, varying between 30% and 63% at the country level when the eligibility criterion changes to CD4 count <350. There is minimal impact on the incremental cost when phasing out d4T either fast or more slowly when the eligibility criterion varies. When testing and outreach costs are included, total costs increase by 62%, from US$26.3 billion under the previous eligibility criterion of treating those with CD4 <200 to US$42.5 billion using the revised eligibility criterion of treating those with CD4 <350.
In addition, the number of AIDS deaths decreases at the global level by 21% when the eligibility criterion switches to CD4 count <350, with country-level results varying between decreases of 16% and 23%. Combining the data results in a cost per AIDS deaths averted varying between approximately US$7,100 and US$9,700 (US$4,800 and US$14,000 along with US$6,400 and $16,300) depending on the change in eligibility criterion.
Source: Authors' calculations.
Teresa Swift, University of Bristol Dunn and colleagues (2011) provide an interesting analysis of the ethical issues involved in deep brain stimulation (DBS) for people with treatment-resistant depression (TRD). I agree with the conclusions the authors draw regarding TRD patients' capacity to consent but would like to take up the authors' call for greater examination of the way in which desperation may affect patients' decision making about research participation. I propose that desperation affects not capacity to consent but voluntariness and that any attempts to explore desperation should reflect this.
In the context of informed consent, voluntariness is legally defined in relation to external constraints imposed by others, in forms such as coercion, undue influence, force, or fraud (see, e.g., Jackson 2009). Other accounts of voluntariness, however, acknowledge the role of circumstances. Roberts (2003) has previously argued that illness-related factors and psychological issues, among other things, may affect the voluntariness of decisions. Likewise, Nelson and Merz (2002) note that threats to voluntariness can arise from potential participants' vulnerabilities, while Hewlett (1996) specifically criticizes the informed consent model for failing to address the influence of circumstances on consent to clinical research participation. Patient circumstances may of course include the nature of their illness and any feelings of desperation that their condition creates. These accounts are supported by a definition of voluntariness provided by Olsaretti (1998). Olsaretti makes a distinction between free choices and voluntary choices, arguing that they are not necessarily related but are often confused with one another.
Choices should be termed free if they are not subject to the influences of other people. Choices are nonvoluntary, however, if they are made because no other option is acceptable to a person in terms of that person's well-being other than the option ultimately chosen. Choices are voluntary when there is an acceptable alternative or when, even if there is no acceptable alternative, the only option available is chosen because the agent likes it and would choose it even if another acceptable option also existed. Olsaretti gives an example of nonvoluntariness, as follows:
Daisy lives in a city surrounded by desert. She desires to leave, but knows she would not survive the journey through the sands, and therefore chooses to stay. Daisy is free to leavenobody prevents her-but she acts non-voluntarily, since she stays only because all other possibilities would be fatal. (Olsaretti 2004, 138-139) If we apply this notion of voluntariness to the health care context, it may be the case that a patient can only make a voluntary choice to participate in research if more than one acceptable option is available to her. Otherwise, as Hewlett notes, her choice is merely theoretical. Thus, if the choice a desperate patient faces is between undergoing an experimental treatment, or declining and accepting the lack of effective standard treatment for her, the decision to enrol in a trial may seem the only acceptable option. The decision may be a competent one, fully informed and even free, but it may not be voluntary. Dunn and colleagues consider whether desperate patients lack genuine autonomy in some way, even if they cannot be presumed to lack decision-making capacity; Olsaretti's notion of voluntariness appears to provide an explanation of the specific way in which the desperate patient's autonomy may be impaired without implicating capacity.
In order to address concerns about desperation, however, Dunn and colleagues propose the use of instruments assessing patients' attitudes to research and its risks and benefits in order to strengthen the process of informed consent. If such an instrument is to tap into the concept of desperation, the crucial question to ask potential research participants, based on the account of voluntariness just given, is whether they feel they have any acceptable alternative other than to say yes to entering a trial when it is offered, and whether they would have declined the trial offer if such an alternative had existed.
If Olsaretti's definition of voluntariness is accepted, however, the ability of the desperate patient to give consent, as it is presently defined, may not technically be affected, even though the voluntariness of the person's decision making might be. There is also the question of what action recruiting researchers should take in light of the information gathered from such an instrument. Dunn and colleagues propose a more thorough consent process, which is always to be commended. However, while detailed attention to the information and comprehension aspects of informed consent may indeed help to correct misperceptions or misplaced expectations, this process is unlikely to be able to assuage the desperation that may drive a patient to undergo an experimental intervention even when fully and accurately informed about significant risks and uncertain benefits and in possession of the cognitive capacity to make such a decision. Agrawal (2003) states that it is important to characterize ethical concerns correctly in order to apply the appropriate safeguards. For desperate patients, as with other vulnerable individuals, if a particular aspect of autonomy cannot be enhanced at the consent level (and is not even incorporated into the consent concept in the case of Olsaretti's definition of voluntariness), the key safeguard may lie at a different stage of the research process. Agrawal believes that vulnerability is more useful as a term if one defines what a person is vulnerable to. In the case of a research trial, certain patients may be vulnerable to exploitation, i.e., vulnerable to accepting an unfair distribution of the risks and benefits of the research. Potentially exploitative offers should of course be addressed by the ethical researcher at the trial design stage, but the appropriate safeguard against exploitation, Agrawal argues, is ethics committee review to ensure that any patient invited into a trial is offered a fair balance of risks and potential benefits if the person participates. This is not simply a matter of equipoise but also of providing a "fair deal" within each trial group (since two trial groups may be equally "unfair" and still satisfy equipoise). Even the desperate patient who finds a trial offer irresistible should not, therefore, be faced with a poor risk/potential benefit ratio that exploits that person's desperation. As Resnik (2002) argues, "There is nothing inherently wrong with conducting research on subjects that suffer from . . . misfortunes or vulnerabilities, provided, of course, that one does not take unfair advantage of those subjects" (2002,29). Resnik's conclusion accommodates Dunn and colleagues' legitimate concerns that to exclude desperate patients from research is to ignore the ethical principle of justice.
In summary, I argue that while the relationship between TRD and consent may be explicated in terms of capacity, desperation may have its effect on voluntariness (as defined by Olsaretti) rather than on capacity and that its effect may be difficult to ameliorate through informed consent measures. Since my commentary is also a theoretical exploration of the issue, I support Dunn and colleagues' call for more empirical research into the relationship between desperation, autonomy, and research participation.
January-March, Volume 2, Number 1, 2011 ajob Neuroscience 45
January-March, Volume 2, Number 1, 2011
U.S.
/entries/decisioncapacity Dunn, L. B., P. E. Holtzheimer, J. G. Hoop, H. S. Mayberg, L. Roberts, and P. S. Appelbaum. 2011. Ethical issues in deep brain stimulation research for treatment-resistant depression: Focus on risk and consent.
We present the design and implementation of VISAGE (VISual AGgregator and Explorer), a query interface for clinical research. We follow a user-centered development approach and incorporate visual, ontological, searchable and explorative features in three interrelated components: Query Builder, Query Manager and Query Explorer. The Query Explorer provides novel on-line data mining capabilities for purposes such as hypothesis generation or cohort identification. The VISAGE query interface has been implemented as a significant component of Physio-MIMI, an NCRR-funded, multi-CTSA-site pilot project. Preliminary evaluation results show that VISAGE is more efficient for query construction than the i2b2 web-client.
VISAGE is the query interface being developed for Physio-MIMI, an NCRR-funded, multi-CTSA-site project [7] to improve informatics support for researchers conducting clinical studies.
The Physio-MIMI data integration environment has two salient features. First, it is a federated system linking data across institutions without requiring a common data model or uniform data source systems. This would greatly reduce data warehousing activities such ETL, often a significant overhead for data integration. Second, Physio-MIMI is tightly focused on serving the needs of clinical research investigators. VISAGE must therefore provide robust data mining capabilities and must support federated queries, while still being user friendly.
The goal is for VISAGE to be directly used by clinical researchers, for activities such as data exploration seeking to formulate, clarify, and determine the availability of support for potential hypotheses as well as for cohort identification for clinical trials. Such an interface would enable an evolution of the data access paradigm: the current paradigm (left of Fig. 1) is one in which clinical investigators communicate a data request to an Analyst or Database Manager (1) who in turn translates the request into a database query and interrogates the database (2) to obtain requested data, finally returning results (3). The time span between 1 and 3 can be weeks if not months, and steps 1-3 often need to be repeated as the query criteria are refined. VISAGE seeks to change this to a paradigm which empowers clinical investigators with data access and exploration tools directly (right of Fig. 1). In this case clinical investigators (1) and data analysts (2) access data directly, and then perform collaborative data exploration (3). This paper reports the design, implementation and preliminary evaluation results of VISAGE. We take a user-centered approach, proven essential for successful user interface development for websites [4]. This approach requires the engagement of the end-user in all steps of the developmental process, such as needs analysis, user and task analysis, functional analysis and requirement analysis. To improve usability, VISAGE incorporates visual, ontological, searchable and explorative features in three main components: (1) Query Builder, with ontology-driven terminology support and visual controls such as slider bar and radio button; (2) Query Manager, which stores and labels queries for reuse and sharing; and (3) Query Explorer, for comparative analysis of one or multiple sets of query results for purposes such as screening, case-control comparison and longitudinal studies. Together, these compo-nents help efficient query construction, query sharing and reuse, and data exploration, which are important objectives of the Physio-MIMI project.
The query interface is increasingly recognized as a bottleneck for the rate of return for investments and innovations in clinical research [1,3,5,8]. Improving query interfaces to clinical databases can only result from an approach that centers around the work requirements and cognitive characteristics of the end user [4], not the structure of the data. To date, few interfaces are usable directly by clinical investigators, with the i2b2 web client [3,5] a possible exception. Aspects of query interface design that facilitate its use by investigators include query-by-example, tree-based construction, being database structure agnostic, obtaining counts in real time before the query is finished and executed, and saving queries for reuse.
The goal of Phyiso-MIMI is to develop informatics tools to be used directly by researchers to facilitate data access in a federated model for the purposes of hypothesis testing, cohort identification, data mining, and clinical research training. In order to accomplish this goal a new approach to the query interface was necessary.
To make VISAGE usable directly by clinical researchers, we adopt the agile development methodology [2]. A key requirement of this methodology is the close interaction between the developers and the users. In designing VISAGE, we follow user-centered design [4] principles, which involve use cases, user and task analysis and functional analysis, described in the rest of this section. In developing VISAGE, clinical researchers have been integrated in the same Physio-MIMI team, with weekly and monthly meetings focusing on design refinements based on feedback from live demonstration and user testing of working components.
The Physio-MIMI team consists of informaticians and clinical researchers from three CTSA Institutions listed in the author area. The needs of the intended user community were evaluated during a face-to-face meeting of the clinical researchers and the design team, and were refined through monthly telephone meetings between the developer team and the end-user team.
During these meetings the clinical researchers helped identify use cases to highlight the power of VISAGE to query physiological and clinical data from one or more repositories residing at one or more institutions. The first use case involves the identification of a research cohort that meets specific demographic, physiological, and clinical criteria and the subsequent identification of a second similar cohort to serve as control subjects. It is necessary for the user to be able to specify the selection criteria and quickly obtain results of the number of records available in the selected data repository(-ies), to save the query, and then to repeat the query modifying one or more criteria in order to identify a second cohort.
The needs analysis also revealed features of VISAGE that would be desired for it to be useful to clinical researchers. First, it was clear that the users wanted to be able to identify clinical criteria for use in the query based on clinical and logical terminology, not technical or database schema terminology. Similarly, the interface needed to allow for searching for available terms based on a number of synonyms (for example, searching for BMI or Body Mass Index). Second, users wanted to receive immediate feedback on the counts returned by a query rather than having to submit and wait each time criteria are adjusted. This allows the users to see the impact of adding or modifying criteria and more quickly construct the query that meets the current need. Third, the ability to direct a single query to one or more underlying sources of data without explicit knowledge of each of the different database structures. Fourth, the ability to save and reuse queries to avoid having to repeat the process of specifying very detailed criteria in order to change a single aspect of the query.
The needs analysis made it clear that the overarching use case for VISAGE, of which there are several more specific thematic variations, is a clinical researcher exploring available data with the intent of discovering the nature, scope, and provenance of the data as it may apply to the researcher's interests and intended uses. Among the variations thus far envisioned are the following: (1) searching for hitherto unnoticed patterns of association and correlation among the available data that suggest or reinforce nascent research hypotheses; (2) deriving and assembling clinical, demographic, behavioral, and assay data sets for use in statistical analyses that can be used in the justification of funding proposals for research studies; and (3) profiling patient populations to determine the availability of cohorts who could be recruited as subjects in proposed research studies.
Such tasks are commonly referred to as data mining, typical down-stream steps that require in-depth analysis, by statisticians or computer scientists, of queried data sets for the discovery of patterns and associations. VISAGE's Query Explorer interface serves to incorpo-rate those activities that are typically carried out in such down-stream data mining analysis, in order to support discovery-driven query exploration by clinical investigators directly. Understandably, what can be achieved by online analysis based on an extended query interface will neither be as powerful nor as comprehensive as dedicated off-line study which may take weeks or months to complete. VISAGE is not designed to replace the role of data mining; rather, it complements data mining by incorporating steps that may be routinely performed before a more in-depth, off-line analysis.
In order to support hypothesis generation and testing and cohort identification, the key challenge is an interface that greatly accelerates access to relevant data sets: past queries should be quickly recallable; new queries should be easily constructible; existing queries should be readily modifiable.
The sense of exploration would quickly diminish if it takes too much effort or too much time for a set of queries to return meaningful results. To help achieve a speedy response of the system during the highly explorative phase of the user, VISAGE provides the user a choice of three tiered query results: counts only; counts with attribute vectors; attribute vectors with associated files (physiological signal data, genetic data, or other large binary files such as images). Typically, results are limited to counts and aggregate statistics until the user achieved a sense of which direction to pursue further.
The Query Explorer and some of the design features are aimed at reducing the user's effort in formulating new queries and revising existing ones. The visual slider bars have the added advantage of error reduction for constraint specification.
Ontological Support. Due to the complexity of the clinical and physiological data to be available through Physio-MIMI, a federated model was preferred. Rather than forcing each data source to conform to a standard database schema, Physio-MIMI is based on the mapping of individual databases to a common Sleep Domain Ontology (SDO), also being developed as part of Physio-MIMI. The SDO consists of a set of concepts (terms) in the sleep medicine domain and the relationships between the concepts. The concepts are organized in hierarchical (SubClass, IS-A) relationships, as well as others such as "partOf", "findingSite", "associated-Morphology", etc. The Query Builder, backed by the domain ontology, provides a searchable list of terms as the starting point. And for each term, it provides the user with context-specific navigation to explore its relationships -allowing the user to traverse up or down the parent-child hierarchical relationships as well as along the other axis relevant to the term in order to further refine the query. By employing the SDO, a standard set of terminology can be employed while allowing individual data contributors to maintain data according to their desired schema. The ability of VISAGE to query across disparate databases acros institutions is therefore dependent on this ontological mapping. The Query Builder provides the user interface to formulate the necessary patterns -allowing the construction of a logical query. The logical query is translated into a local database query based on the mapping between the ontology model and the database specific data model.
This section focuses on the resulting design and implementation of Query Builder (Fig. 2) and Query Explorer (Fig. 3). The Query Manager saves queries (optionally their results) for reuse, which can be searched by keywords in title, description, or the query itself (e.g. for finding queries about a specific symptom or disorder). We omit the description of Query Manager since the functionalities of this component is similar to that of an email management application.
The query builder interface includes functional areas 1-12 (Fig. 2).
A main technical objective of Physio-MIMI is to demonstrate the feasibility, using clinical sleep research as an exemplar, to integrate clinical research databases across institutions and support a federated query interface without requiring dedicated data warehouses to be constructed at local institutions. We leverage two pieces of prior technology: MIMI 1 (Case Western) and Honest Broker 2 (Michigan). Strategic decisions were made at project inception to address challenges involved in:
• Federation avoiding heavy movement of large data sets • Incremental integration without requiring standard/uniform source database systems or formats (bottom-up than top-down)
• Ontology-driven integration and query interface without prior data standardization • Emphasis on "up-stream" informatics support for ongoing projects, rather than typical "down-stream" data integration for sharing after project completion
The query interface is increasingly recognized as a bottleneck for the rate of return for investments and innovations in clinical research. In order for vast amounts of collected data to be of significant use, there must be an effective method to query and reuse the data. VISAGE is the query interface for Physio-MIMI aimed to be usable directly by clinical research investigators, rather than traditional database analysts. In accordance with user-centered design principles, the design and implementation of VISAGE offer three separate yet linked components to deliver the desired functionality:
• The Query Builder offers ontological support to allow the user to select concepts on which to query. When concepts are added, controls specific to the concept allow the user to specify the requirements • The Query Manager stores the user's saved queries and allows the sharing of queries with other users • The Query Explorer gives the user a powerful interface to rapidly explore one or more queries at a time, comparing them on the basis of several concepts at once using charts and statistics
The (1) Database Selector lets the user select which database in the system to run the query against. The (2) Search Bar allows the user to search the hierarchy of terms, displaying those that match in the (3) Term Selection Area below. When terms are clicked, they are added to the (4) Term Display Area. Terms can be selected with the checkboxes, and selected terms can be grouped together or separated by clicking ( 5 1 2 3 4 5 6 7 8 9 10 11 12 Domain Expert Query Builder Honest Broker Core Honest Broker Adapter Institutional Firewall VISAGE allows informaticians to quickly make data sources available for querying by supplying tools for secure database connectivity and online tools for mapping database elements to SDO concepts. Once mapped, the database can be available to query and will appear to the user in the Database Selector. Using the Query Builder researchers can quickly generate a query across multiple databases or compare results of the same criteria against different databases. The (2) Search Bar allows the user to search the hierarchy of terms, displaying those that match in the (3) Term Selection Area below. When terms are clicked, they are added to the (4) Term Display Area. As mentioned above, the user can search for any synonyms of concepts in the ontology and be presented with the appropriate ontological concept. The searchable list of terms is backed by the SDO and provides the user the ability to navigate using ontological relations to further refine the query, in a similar manner to i2b2 [3]. To use the VISAGE interface, the clinical researcher needs only to understand the clinical model (domain ontology), and the Query Builder provides the interface for formulating the necessary patterns for the construction of a logical query. The logical query is then translated into a database-specific query based on the mapping between the ontology model and the database schema.
By default, the query's logic is in Conjunctive Normal Form, which means records need only satisfy one condition in each group to be included in the query result set. To change to Disjunctive Normal Form, the (6) Flip action is made available. The grouping logic is denoted by the color of the box. Elements in a green box are logically connected by AND, while elements in a light blue box are joined by OR. Terms can be selected with the checkboxes and grouped together or separated by clicking (5) Group or Ungroup, allowing for different parenthetical groupings of terms for the conjunctive or disjunctive relationships. Additional term manipulation functionality includes (7) Rearrangement, which lets a user drag and drop the terms to arrange them how he wishes, and (8) Deletion, which allows removal of terms that the user may have mistakenly added to the query. To specify inclusion conditions, each term added to a query comes with term-specific controls.
For categorical data, (9) Checkboxes display the possible values for categorical variables. The values for categorical variables are also derived from the SDO, and map to specific values in the underlying database schema(s). The user need only know the conceptual categories not the underlying structure, and due to the VIS-AGE database mapping individual databases need not code categorical variables in the same manner. For continuous variables, (10) Sliders allow easy and expressive creation of intervals, with ranges of inclusion specified by light blue shading as well as numeric display. The Sliders have the additional advantage of allowing for the creation of multiple disjoint intervals, something that is often not possible in interfaces that provide manual specification of continuous ranges.
When the user is finished adding terms and modifying inclusion conditions, the number of records that satisfy the conditions is displayed in the (11) Result Count Area. Finally, the user can (12) Describe/Save/Update the query to the Query Manager for future use in the Query Explorer or re-use in the Query Builder.
The Query Explorer allows the records returned by one or more queries to be further investigated. Not only can the user view distributions of the terms that were used as criteria in the specification of the query, but any other available term can be selected for exploration within that result set. The Query Explorer provides numeric distributional information including frequency and percent for each level of categorical variables, and mean, standard deviation, and range for continuous variable. The Query Explorer also provides graphical displays of distributions including pie charts and histograms for categorical and continuous variables, respectively. Discovery-driven query exploration may start with one, two or multiple queries in a query group, arranged in a specific order by the end user, not unlike a workflow. The queries in a query group are "aligned" to allow the user to zero in on selected attributes to gain a sense of value distribution of the selected attribute among the patients represented in the query results.
By exploring the value distribution of a certain variable within a set of query results, a user may discover how some of the baseline query criteria influence the value distribution of specific attributes (e.g., as pie-chart in Fig. 3), without issuing another query with an additional attribute specified. For example, Fig. 3 illustrates an explorative step for the query used in Fig. 2, where no gender criteria is included. The Query Explorer interface allows one to search and select variables that may or may not be present in the original query. The piechart on left shows the gender distribution in the result for the selected query (in Fig. 2). The histogram of age distribution is displayed on the right. One can imagine that by selecting two or more queries, one can explore potential patterns for with a case population and a control population (one query for each), or for Longitudinal Studies (same query with varying time points).
A preliminary evaluation was performed on the efficiency of VISAGE for query construction. Three common queries with increasing levels of logical complexity on patient demographics were selected. Two expert users created the queries in both VISAGE and the i2b2 web client, respectively. The number of clicks and time needed for creating the queries were recorded and tabulated in the next
table. Query VISAGE i2b2 Web Client # of clicks time (sec.) clicks time (sec.) 1 5 13 14 59 2 6 16 25 119 3 20 52 37 160
As can be seen from the table, VISAGE reduced time and effort (in terms of the number of clicks) to a half or nearly a third. However, we caution that this evaluation is very preliminary and it only looks one specific aspect of the query interface. For example, a larger number of query samples and users with varying computer experiences should be included for a more comprehensive usability evaluation, followed by a rigorous statistical analysis.
The development of VISAGE has focused on a powerful interface that is intuitive, usable and simple. Agile [2] and user-centered methodologies [4] are used for the query interface development. It entails that a clear separation between design and implementation is neither feasible, nor necessary. Design versions are usually at a conceptual or functional level, and the details are relegated to the prototyping phase, which drives the design revision. This is the reason that the complete interface design of VISAGE is embodied in its implementation in Section 4. Rapid prototyping of VISAGE is achieved through the use of various Open Source Web development tools and frameworks including Ruby on Rails, Prototype, and script.aculo.us JavaScript libraries. All of these are web-based (Web 2.0) and work across platforms.
We like to thank other members of the
To understand the hemodynamics of hepatocellular carcinoma (HCC) is important for the precise imaging diagnosis and treatment, because there is an intense correlation between their hemodynamics and pathophysiology. Angiogenesis such as sinusoidal capillarization and unpaired arteries shows gradual increase during multi-step hepatocarcinogenesis from high-grade dysplastic nodule to classic hypervascular HCC. In accordance with this angiogenesis, the intranodular portal supply is decreased, whereas the intranodular arterial supply is first decreased during the early stage of hepatocarcinogenesis and then increased in parallel with increasing grade of malignancy of the nodules. On the other hand, the main drainage vessels of hepatocellular nodules change from hepatic veins to hepatic sinusoids and then to portal veins during multi-step hepatocarcinogenesis, mainly due to disappearance of the hepatic veins from the nodules. Therefore, in early HCC, no perinodular corona enhancement is seen on portal to equilibrium phase CT, but it is definite in hypervascular classical HCC. Corona enhancement is thicker in encapsulated HCC and thin in HCC without pseudocapsule. To understand these hemodynamic changes during multi-step hepatocarcinogenesis is important, especially for early diagnosis and treatment of HCCs.
Hepatocellular carcinoma (HCC) is the most common primary liver cancer worldwide. Approximately 80% of Japanese HCC cases are derived from HCV-associated liver cirrhosis and chronic hepatitis, and the remaining less than 20% of the patients are HBV positive. The patients with hepatitis B or C cirrhosis are especially classified as a very high-risk group. Ultrasonography is performed every 3-4 months for the very high-risk group. Because of the introduction of this surveillance system, the size of HCCs firstly detected during 2002-2003 (n = 33731) was less than 2 cm in 32.5% of all cases, 2.1-5.0 cm 47.0%, respectively [1]. However, various types of hepatocellular nodules such as dysplastic nodule (DN) are also detected during screening procedures. Ultrasound and CT features of DNs and early HCCs are similar, and a precise differential diagnosis is impossible. Pathologically, human HCC develops in a multistep fashion from DN to classic hypervascular HCC. Therefore, for the early diagnosis of HCC, understanding of the concept of multi-step hepatocarcinogenesis and the sequential changes of imaging findings in accordance with multi-step hepatocarcinogenesis is important.
To understand the hemodynamics of HCC is important for the precise imaging diagnosis and treatment, because there is an intense correlation between its hemodynamic and pathophysiology. For this purpose, dynamic MDCT is most valuable because of its high spatial and contrast resolution. However, because of the dual blood supply of the liver and intravenous injection of the contrast medium, the precise analysis of hemodynamics by conventional MDCT is often difficult. By the introduction of dynamic CT during selective arteriography, including CT during arterial portography (CTAP) [2,3] and CT during hepatic arteriography (CTHA) [4], it has become possible to visualize the distribution of the intra-hepatic portal and arterial blood flow separately with extremely high contrast resolution, and as a result, to analyze precisely the correlation between blood supply and pathophysiology. In this article, blood flow imaging features of HCC will be discussed based on the CTAP and CTHA imaging and pathophysiologic correlations with special reference to multistep hepatocarcinogenesis.
The concept of multi-step hepatocarcinogenesis and related small hepatocellular nodules in the patients with chronic liver diseases, particularly those with cirrhosis or chronic hepatitis caused by hepatitis B or C viruses, was developed mainly in Japan. However, it had not been widely accepted throughout the world and the diagnostic criteria of these nodules different even among the world specialists. However, in 2009, the International Consensus Group for Hepatocellular Neoplasia organized by the world's leading liver pathologists finally reached agreement [5].
According to this report, these nodules are divided into large regenerative nodule, low grade DN (L-DN), high-grade DN (H-DN), and HCC. In addition, small HCC (less than 2 cm) is divided into early HCC and progressed HCC. Early HCC has a vaguely nodular appearance and is well differentiated. Progressed HCC has a distinctly nodular pattern and is mostly moderately differentiated, often with evidence of microvascular invasion. L-DNs are vaguely or distinct nodular with mild increase in cell density and no cytologic atypia. H-DNs are more likely to show a vaguely nodular pattern with architectural and/or cytologic atypia, but the atypia is insufficient for a diagnosis of HCC. They show increased cell density, sometimes more than two times higher than the surrounding nontumoral liver, often with an irregular trabecular pattern. Unpaired arteries are found in most lesions, but usually not in great numbers. A nodule with largely H-DN features containing a subnodule of well-differentiated HCC can be seen. Early HCCs are vaguely nodular and are characterized by various combinations of the following major histologic features; (1) increased cell density more than two times that of the surrounding tissue, with an increased nuclear/ cytoplasm ratio and irregular thin-trabecular pattern; (2) varying numbers of portal tracts within the nodule (intratumoral portal tracts); (3) pseudoglandular pattern; (4) diffuse fatty change; and (5) varying numbers of unpaired arteries. Any of the features listed above may be diffused throughout the lesion or may be restricted to an expansile subnodule (nodule-in-nodule). Most importantly, because all of these features may also be found in H-DNs, it is important to note that stromal invasion remains most helpful in differentiating early HCC from H-DNs. However, the application of these criteria is challenging because most histologic criteria are arrayed on a gradual spectrum and cannot be easily summarized as present or absent.
Because of these reasons as described above, it should be realized that there must be various degree of overlaps among imaging features of these nodules and they may show gradual changes during multi-step hepatocarcinogenesis.
Vascular endothelial growth factor (VEGF) is known to play a critical role in the neovascularization in the development and progression of malignant neoplasms [6,7]. VEGF is produced by tumor cells, and its binding with VEGF receptors such as Flt-1 and Flk-1, which are expressed on vascular endothelial cells, leads to the proliferation and migration of endothelial cells. In addition, VEGF receptors expressed on tumor cells are involved in tumor proliferation in an autocrine loop via interaction with VEGF produced by the tumor cells themselves [7]. Park et al. [8] reported that the expression of VEGF was correlated with angiogenesis and cell proliferation in hepatocarcinogenesis. On the other hand, tumors often encounter hypoxic conditions during their growth. Under such conditions, hypoxia inducible factor-1a (HIF-1a) promotes the transcriptional activity of angiogenesis-related molecules such as VEGF and erythropoietin by affecting the hypoxia response element and HIF-1a located in nuclei [7].
We analyzed these changes of the angiogenesis during multi-step hepatocarcinogenesis by immunohistochemical and molecular studies [7]. According to our analysis, it was found that hepatocellular areas around the portal tracts in DNs, including those with sinusoidal capillarization and unpaired arteries, were strongly positive for HIF-1a, whereas this molecule was faintly expressed in the surrounding livers. Cytoplasmic overexpression and intranuclear expression of HIF-1a, a more increased expression pattern, were also observed in HCC, suggesting that cytoplasmic HIF-1a might have been moved into the nuclei in activated HCC cells. HIF-1a is involved in the upregulation of genes harboring the hypoxia response element such as VEGF, suggesting that increased expression of HIF-1a in the areas around the portal tracts of DNs may be responsible for increased expression of VEGF and its receptor followed by sinusoidal capillarization and increased numbers of unpaired arteries in DNs and also in the angiogenesis in HCC. These expressions gradually spread into the entire nodule in accordance with the elevation of the grade of malignancy of the nodules (Fig. 1).
Figure 2 shows an early HCC consisting of H-DN with a small part of highly differentiated HCC with stromal invasion. On CTHA, a well-differentiated focus demonstrates a faint enhancement and this portion reveals more expression of sinusoidal capillarization and unpaired arteries than that in the surrounding H-DN. dance with the elevation of the grade of malignancy of the nodules during hepatocarcinogenesis [Double immunohistochemical staining for CD 34 (blue) and a-smooth muscle actin (SMA) (brown)].
We previously described that the intranodular blood supply evaluated by CTAP and CTHA changed in accordance with hepatocarcinogenesis from DN to overt HCC [4,9]. On CTAP, the intranodular portal supply could be divided into four types relative to the sur-rounding cirrhotic liver [4] (Figs. 3, 4, 5, 6, 7), namely, isodense nodule relative to the surrounding liver indicating almost the same intranodular portal supply (type A), slightly hypodense nodule indicating decreased but not absent intranodular portal blood flow (type B), a part of the nodule showing a definitely hypodense area indicating a partially absent intranodular portal blood supply (type C) and definitely hypodense indicating an absent intranodular portal supply (type D). The correlation between the histologic types of nodules and these CTAP findings revealed the significant correlation or B Serial specimens from 1 to 3 with double immunohistochemical staining for CD 34 and a-SMA show the communication between intratumoral blood sinusoids (arrowhead) and hepatic venules (arrows) in the tumor (9100). cating almost the same intranodular arterial supply relative to the surrounding liver (arrow).
strong tendency between DN and type A, early HCC and type B, well-differentiated HCC and type C, and moderately or poorly differentiated HCC and type D (Figs. 3, 4, 5, 6, 7). On CTHA, the intranodular arterial supply could be also categorized into four types relative to the surrounding cirrhotic liver (Figs. 3, 4, 5, 6, 7), namely, isodense nodule indicating almost the same intranodular arterial blood supply relative to the surrounding liver (type I), hypodense indicating decreased arterial blood supply (type II), a part of the nodule demonstrating hyperdensity indicating a partially increased arterial supply (type III) and entirely hyperdense indicating entirely increased arterial supply (type IV). The correlation between the histologic types of nodules and these CTHA hypodense focus indicating partial portal perfusion defect (arrow). On CTHA (right), it is visualized as an entirely isodense nodule with an internal definitely hypervascular focus (arrow). B Double immunohistochemical staining of the boundary between the tumor and liver parenchyma shows abundant communications between intratumoral blood sinusoids and hepatic sinusoids (9100).
findings revealed the significant correlation or strong tendency between type I and L-DN and early HCC, type II and H-DN and early HCC, type III and well-differentiated HCC and type IV and moderately or poorly differentiated HCC (Figs. 3, 4, 5, 6, 7). These results suggested that the intranodular portal supply relative to the surrounding liver parenchyma is decreased, whereas the intranodular arterial supply is first decreased during the early stage of hepatocarcinogenesis and then increased in parallel with increasing grade of malignancy of the nodules as shown in Fig. 8. However, the differences of histological findings among these nodules are sequential, and the exact diagnosis of the entire nodule is occasionally difficult because of internal histological heterogeneity. Therefore, as shown in the original article [7], there was a fairly wide range of overlap in blood supply patterns among the various types of hepatocellular nodules.
To verify the histological background of the findings obtained by CTAP and CTHA, we analyzed morphometrically the vascular supply of DN and HCCs [10,11], and suggested that the portal tracts including portal vein and hepatic artery were decreased in accordance with increasing grade of malignancy and virtually absent in HCCs. In contrast, abnormal arteries due to tumor angiogenesis developed in H-DN during the course of hepatocarcinogenesis, and were markedly increased in number in moderately differentiated HCCs. from 1 to 4 with double immunohistochemical staining for CD34 and aSMA show the communication between intracapsular portal venules (a-d) and intratumoral blood sinusoids (arrowheads) (9100).
By single level dynamic thin-section CT during the bolus injection of a small amount of contrast medium, we revealed in vivo hemodynamics in hypervascular classical HCC, namely, the arterial blood flow into the tumor drains into surrounding hepatic sinusoids (corona enhancement) (Fig. 7) [12]. This drainage was well visualized in the late phase of CTHA which was taken after the stoppage of the infusion of the contrast medium into the hepatic artery. Histological examination revealed continuity between a tumor sinusoid and a portal venule in the pseudocapsule (encapsulated HCC) Fig. 8. Multi-step hepatocarcinogenesis and changes of intranodular blood supply. Intranodular portal supply gradually decreases in accordance with the elevation of the grade of malignancy of the nodules and finally disappears in moderately differentiated HCCs. On the other hand, arterial supply first decreases at the early stage of hepatocarcinogenesis and then acutely increases, and finally the entire nodule is fed only by artery in moderately differentiated HCCs.
Fig. 9. Multi-step changes of drainage vessels and peritumoral enhancement during hepatocarcinogenesis. In DNs or early HCCs, the main drainage route from the tumor is intranodular or perinodular hepatic vein. However, because hepatic veins disappear from the tumor during very early stage of hepatocarcinogenesis, drainage vessels change to hepatic sinusoids. In moderately differentiated HCC with pseudocapsule formation, the communication between tumor sinusoids and the surrounding hepatic sinusoids are also blocked, and then, the portal venules in the pseudocapsule finally become the main drainage vessel from the tumor. In accordance with the changes of the drainage vessels, thin to thick corona enhancement appears surrounding the tumor.
or surrounding hepatic sinusoids (HCC without pseudocapsule) [11,12]. According to our recent histological study correlated with CTAP and CTHA, the main drainage vessels of hepatocellular nodules change from hepatic veins to hepatic sinusoids and then to portal veins during multi-step hepatocarcinogenesis, mainly due to disappearance of the hepatic veins from the nodules [11]. Therefore, in early HCC, no perinodular corona enhancement is seen on portal to equilibrium phase CT, but it is definite in hypervascular classical HCC (Figs. 3, 6, 7, 9). Corona enhancement is thicker in encapsulated HCC and thin in HCC without pseudocapsule (Figs. 6, 7, 9). The drainage flow from hypervascular HCC variously modified the imaging findings, a feature useful for differential diagnosis. Drainage from the tumor makes the tumor appear larger than it really is on various kinds of blood flow imaging findings. It is the first site of the intrahepatic metastasis of HCC, and daughter nodules are commonly seen in the drainage area. Iodized oil flowed into the surrounding liver through this drainage route and enhanced the effect of transcatheter arterial chemoembolization [13]. The drainage area should be included in RFA area to prevent local recurrence.
Metastatic liver cancers show thin corona enhancement or early peritumoral enhancement on single-level dynamic CTHA [14]. We named the former as ''drainage pattern'' and the latter as ''arterio-portal (AP) shunt pattern''. In cases with drainage pattern, the tumor shows hypervascularity in early phase and thin peritumoral enhancement in late phase similar to hypervascular HCC without pseudocapsule, and the drainage route may be the connection between tumor sinusoids and hepatic sinusoids surrounding the tumor [14,15]. In cases with AP shunt pattern, the tumor shows no definite enhancement except faint staining on the peripheral margin of the tumor, but early peritumoral enhancement with occasional wedge-shaped expansion is seen. The mechanism of this early enhancement is unknown, but peritumoral multicentric AP shunts due to the obstruction of the portal or hepatic venules can be one of the possible causes. Because abundant fibrous tissue is often contained in this kind of tumor, internal delayed enhancement is commonly associated. Mass-forming type of cholangiocarcinoma usually demonstrates AP shunt pattern. Among malignant primary liver cancer, cholangiolocellular carcinoma (bile ductular carcinoma) shows unique hemodynamics [16]. It typically shows tumor hypervascularity with surrounding enhancement resembling AP shunt pattern in early phase and delayed internal enhancement on late phase, probably due to abundant cancer cells and fibrous tissues in the tumor with multiple entrapped portal tracts in the tumor (replacing infiltration type growth with the portal tracts incorporated into the tumor). Benign hypervascular hepatic masses such as cavernous hemanigoma, focal nodular hyperplasia (FNH) [17], angiomyolipoma and peliosis hepatis usually do not show corona enhancement, probably due to main drainage to hepatic vein. Two exceptions are hepatic adenoma and hypervascular hyperplastic nodule associated with alcoholic cirrhotic livers which commonly demonstrate corona enhancement [18]. Understanding these hemodynamic differences among various kinds of hepatic mass lesions are important for differential diagnosis and treatment.
In conclusion, it is very important to know the hemodynamics of HCCs and related hepatocellular nodules for the understanding of pathophysiology and precise imaging diagnosis and treatment of HCCs. For this purpose, angiography-assisted CT is most valuable and accurate, but because of its invasiveness, blood flow imaging with contrast ultrasound, dynamic CT and MR imaging is necessary.
Open Access. This article is distributed under the terms of the Creative Commons Attribution Noncommercial License which permits any noncommercial use, distribution, and reproduction in any medium, provided the original author(s) and source are credited.
Objectives: To compare the effect of nebulized racemic epinephrine to nebulized racemic albuterol on successful discharge from the emergency department (ED).
Methods: Children up to their 18th month of life presenting to two teaching hospital EDs with a clinical diagnosis of bronchiolitis who were ill enough to warrant treatment but did not need immediate intubation were eligible for this double-blind randomized controlled trial (RCT). Patients received either three doses of racemic albuterol or one dose of racemic epinephrine plus two saline nebulizers. Disposition was decided 2 hours after the first nebulizer. Successful discharge was defined as not requiring additional bronchodilators in the ED after study drug administration and not subsequently admitted within 72 hours. Adjusted relative risks (aRR) were estimated using the modified Poisson regression with successful discharge as the dependent variable and study drug and severity of illness as exposures. Secondary analysis was performed for patients aged less than 12 months and first presentation.
Results: The authors analyzed 703 patients; 352 patients were given albuterol and 351 epinephrine. A total of 173 in the albuterol group and 160 in the epinephrine group were successfully discharged (crude RR = 1.08, 95% confidence interval [CI] = 0.92 to 1.26). When adjusted for severity of illness, patients who received albuterol were significantly more likely than patients receiving epinephrine to be successfully discharged (aRR = 1.18, 95% CI = 1.02 to 1.36). This was also true among those with first presentation and in those less than 12 months of age.
Conclusions: In children up to the 18th month of life, ED treatment of bronchiolitis with nebulized racemic albuterol led to more successful discharges than nebulized epinephrine.
ronchiolitis is a common childhood disease. It is a leading cause of hospitalization worldwide and accounts for substantial morbidity and a mortality of less than 1%. 1,2 Between 1980 and1996, annual infant hospitalization rates for bronchiolitis doubled to 34 per 1,000, while rates for other lower respiratory tract diseases remained stable. 3 In 2002, bronchiolitis resulted in 149,000 hospital admissions and annual costs of $543 million in the United States. 4 Bronchodilator therapy in bronchiolitis is controversial. Although commonly used, 5 its efficacy is not universally accepted. Small studies [6][7][8][9] and systematic reviews have shown small, short-term improvement in clinical scores in outpatient use of albuterol. 10 There is, however, no convincing evidence that albuterol decreases the admission rate in outpatients. 10 Some, [11][12][13] but not all, [14][15][16] small studies suggest that epinephrine may decrease admissions in outpatients. A Cochrane review of these studies was inconclusive, however, because of the small numbers overall. 17 For the largest study included in that review (comparing racemic albuterol and epinephrine in outpatients), the power to detect a change in admissions of 15% from 50% was 23%. Even if all the cases in that review had been combined in a single study, the power would have been 58%.
The inconclusive results obtained to date suggest that if there is a difference in disposition between racemic albuterol-and epinephrine-treated patients, it is modest. It seems likely that the most important factor predicting the need for admission will be the severity of the bronchiolitis. Without a severity assessment tool validated to predict the need for admission in bronchiolitis, unmeasured but uneven distribution of factors that influence disposition can confound a modest, yet important treatment effect.
The most commonly used tool for severity assessment is the respiratory distress assessment instrument (RDAI). It was originally described in an influential early study evaluating the effect of subcutaneous epinephrine in bronchiolitis. 12 The RDAI produces precise scores and is reliable among different users. Although subsequently validated in asthma, it has not been validated for predicting admission in bronchiolitis. 18 Despite being widely used, because the RDAI limits its assessment to the respiratory system, 12 applying it to infants with bronchiolitis is problematic. By failing to incorporate age it assigns a 3-week-old infant with wheezing a low score, but assigns a high score to an 11-month ''happy wheezer'' with audible wheeze and chest wall retractions. 12 It also does not address other clinical findings such as dehydration, which influence disposition in infants and toddlers.
Prior to undertaking this randomized controlled trial (RCT), we derived a severity-of-illness tool in bronchiolitis patients defining ''admission not needed,'' ''length of hospital stay up to and including the median,'' and ''length of hospital stay greater than the median,'' as ordinal outcome measures. During the derivation of this ordinal regression model, we found that age, dehydration, retraction severity, and tachycardia predicted the outcome. 19 We subsequently validated this model at a different children's hospital and measured its interrater reliability at a third site. 20 This ordinal regression model reflects the systemic consequences of bronchiolitis in addition to the respiratory component and measures a meaningful outcome. It is presented in Figure 1.
The primary objective of this study was to compare the effect of nebulized racemic albuterol to nebulized racemic epinephrine on discharge rates among children presenting to the emergency department (ED) with bronchiolitis while adjusting for severity of illness. Our secondary objectives were to determine the effect of these bronchodilators in the subgroups of infants less than 12 months of age and those presenting for the first time.
We conducted a two-site double-blind RCT comparing nebulized racemic albuterol to nebulized racemic epi-nephrine. Both sites had institutional review board (IRB) approval for the study.
The primary site was a county hospital ED with 53,000 attendances annually (23% children), serving a mixed urban, rural, and suburban population. The secondary site was a community teaching ED with an annual census of 80,000 (20% children). Both hospitals have university affiliations and emergency medicine residencies. The secondary site also has a pediatric residency program. Recruitment occurred from November 1, 2003, to May 1, 2006, and from November 1, 2004, to May 1, 2006, at the primary and secondary sites, respectively.
We defined bronchiolitis operationally as clinical evidence of lower airway obstruction (physical findings of wheezing and chest wall retractions) following an upper respiratory tract infection 21 in children up to the 18th month of life. There was no lower age limit. The adopted upper age cutoff represented a midpoint between the ranges of accepted age limit (12 or 24 months) for defining bronchiolitis. 4,18,[21][22][23] We excluded children with bronchiolitis who required no treatment, those with illness so severe as to require immediate intubation, and those who received bronchodilators in the ED prior to screening. In our attempt to have broad inclusion criteria, we did not exclude patients with a history of prior episodes of wheezing, lung disease, or other comorbidity. Diagnosis was made by an attending physician or midlevel provider.
All patients with a diagnosis of bronchiolitis judged by the attending physician or physician assistant to require treatment were eligible to be enrolled. The treating physician screened and enrolled patients as they presented to the ED. Signed informed consent was obtained from a parent of all children enrolled.
Research assistants (RAs) assisted study investigators at the primary site. All clinicians and RAs received regular training reinforcement in study procedures throughout each bronchiolitis season. At the secondary site, patients were screened and enrolled on a convenience basis when a study attending physician or research nurse was available.
Patients were randomized in blocks of 50 to receive either three consecutive doses of nebulized racemic albuterol or a single dose of nebulized racemic epinephrine followed by two nebulized saline treatments. The large block size reduced the risk of allocation bias and physicians being able to guess the identity of the drug. Randomization was performed using a computer-generated random number series and was stratified per site. The pharmacist had sole access to this list and the codes were secured until the study was complete.
The dose of racemic albuterol given was 0.625 mg for infants weighing less than 5 kg and 1.25 mg for those weighing 5 kg or more. The dose of racemic epinephrine was 11.25 mg (0.5 mL of 2.25% solution) regardless of weight. The hospital pharmacist prepared all of the drug packets. Each were identical and presented in sequentially numbered light-protective polythene packets containing three syringes with 0.5 mL of a clear liquid, three 2.5-mL saline vials, a study label, and a package insert with instructions to the user.
All patients received 2.5 mL of nebulized saline during the consent process prior to randomization. Both saline mist and study drugs were delivered using an oxygen-driven VixOne small-volume nebulizer (Westmed, Inc., Tucson, AZ) at 20-minute intervals. A flow rate of 7 L ⁄ min was used with wall-mounted oxygen switches and between 6 and 8 L ⁄ min with cylinder oxygen. Patients received either three nebulized racemic albuterol doses or one nebulized racemic epinephrine dose, followed by two saline nebulizers.
Data were collected prospectively using standardized forms to document history and physical exam. Data collected included age, history of prematurity, medical history including previous episodes of wheezing, use of albuterol or prednisone before arrival, and any family history of asthma. A nurse recorded the vital signs, including oxygen saturation, at triage, and a physician recorded physical examination findings. When available, RAs scribed the data form and assisted with the consent process.
Two hours following the administration of the first dose of the study drug, the treating physician could discharge, admit, or give additional bronchodilators. Patients in this last group were classified as having a prolonged ED stay. Admission criteria are given in Figure 2, although physicians could override these criteria using clinical judgment. 24 Viral antigen testing, steroid use, and discharge medications were ordered according to the treating physician's preference.
The primary outcome was ''successful discharge,'' defined as discharge following study drug administration, not requiring additional bronchodilators in the ED and not resulting in admission within 72 hours of discharge. Patients who were classified as having a prolonged ED stay were counted as admissions. Reattendance within 72 hours (at a clinic, doctor's office, or ED) that did not result in a hospital admission was considered a successful discharge. This outcome was chosen to ensure that the discharge was appropriate and not just temporarily deferred. We verified any further medical contact by reviewing hospital records, telephone follow-up, and checking county coroner's records to verify vital status where follow-up was unsuccessful.
Severity-of-illness assignment 20 was performed electronically after data collection was complete but before the randomization code was broken. The tool is complex (Figure 1) and at present is designed for research rather than clinical use.
We calculated a sample size required to provide 90% power (with two-sided alpha level of 0.05) based on admission rates of 33% for patients treated with racemic epinephrine and 50% for racemic albuterol. 11 We then rounded this simple sample estimate (n = 374) up to 500 as recommended by Long, 25 anticipating logistic regression as the method of analysis. An entire bronchiolitis season was used as the minimum unit of recruitment to avoid potential bias from recruiting for part of the season. There may be differences between children who present early or late in the season.
Analyses were performed on an ''intent-to-treat'' basis. We analyzed only the first enrollment of patients who were enrolled more than once during the study.
We modeled severity as categorical variables rather than as an ordinal variable. We created a fourth severity category, containing patients missing one or more of the variables needed to determine illness severity. Study site of enrollment was modeled as a categorical variable. Owing to the relatively common outcome, we estimated adjusted relative risks, (aRR), using modified Poisson regression as described by Zou, 26 rather than odds ratios. We present aRR with 95% confidence intervals (CIs) using successful discharge as the dependent variable and study drug, severity of illness, and study site as exposures. We also constructed other models to account for the patients with missing observations on the severity scale by using multiple imputations and by excluding these cases from the model. We also performed a Mantel-Haenszel simple stratified analysis, stratifying only for disease severity. Cases that met admission criteria at enrollment were included in the primary analysis. It was felt that if a significant proportion of these patients improved sufficiently to be discharged posttreatment, this would be an important finding. We excluded these in a sensitivity analysis. We also excluded the prolonged ED treatment group in a sensitivity analysis. We analyzed patients with recurrent wheezing and patients over 12 months of age both by excluding them and by constructing a model with categorical variables for these and other variables of potential interest. These sensitivity analyses used a Poisson model with stepwise backward selection to examine the effect of all the primary and secondary endpoints by including the variables: patients with recurrent wheezing, prematurity, age less than 2 months, steroid use, hypoxia, and a family history of asthma.
We performed double data entry using a customized Filemaker-Pro (Filemaker-Pro, Version 6, Santa Clara, CA) database. Statistical analysis was performed using Stata 9.2 (Release 10, Statacorp, College Station, TX).
A total of 721 patients were enrolled, 655 at one site and 66 at the other. Following the exclusion of those with eligibility violations or multiple enrollments, 703 were analyzed (Figure 3). The mean age of all patients enrolled was 6 months (median 4.8 months), 407 (58%) were male, and 610 (87%) were less than 12 months of age. Respiratory syncytial virus antigen was present in 313 of the 553 (57%) patients tested in the ED. The baseline characteristics are broadly similar and shown in Table 1.
There were 352 patients in the racemic albuterol group and 351 patients in the racemic epinephrine group. Figure 4 shows the age distribution of patients by treatment group. Patient flow to their ultimate disposition is shown in Figure 5. Bronchiolitis was classified as mild, moderate, and severe in 175, 419, and 47 children, respectively. The racemic albuterol group had significantly more (p = 0.007) moderately ill, but fewer mildly ill, patients compared to the racemic epinephrine group. Table 2 depicts the drug assignment and disposition of each severity-of-illness category. Patients were successfully discharged in 333 of 703 (47.4%) cases. Successful discharge decreased directly with severity of illness: moderate (aRR = 0.49, 95% CI = 0.42 to 0.57) and severe (aRR = 0.20, 95% CI = 0.10 to 0.39), compared with mild disease. The proportion of those successfully discharged in each treatment arm for each severity-of-illness category is shown in Figure 6.
Admission rates at the primary site were 52% (331 ⁄ 640) and at the secondary site were 62% (39 ⁄ 63). There were marginally significantly fewer discharges in the secondary site (aRR = 0.76, 95% CI = 0.56 to 1.03). While discharge rates and outpatient treatment failures were higher at the primary site, the use of clinical judgment rather than other admission criteria, other patient characteristics, and steroid use were similar at both sites. Thirty-seven had a second ED visit within 3 days. Six were readmitted within 3 days of initial ED discharge in the racemic albuterol group and 10 patients in the racemic epinephrine group, all from the primary site. Forty-nine of 177 (27.7%) patients who met admission criteria at entry from the primary site were discharged, while all of those who met admission criteria at entry were admitted from the secondary site. Our primary outcome results include adjustment for site.
Crude analysis (not adjusting for severity of illness) showed no difference between groups (RR 1.08, 95% CI = 0.92 to 1.26). After stratification by severity of illness, patients who received racemic albuterol were significantly more likely than patients receiving racemic epinephrine to be successfully discharged (aRR = 1.18, 95% CI = 1.02 to 1.36). This result was the same using Poisson regression and Mantel-Haenszel stratified analysis. Excluding those in the ''prolonged ED treatment'' group did not change the treatment results (aRR = 1.16, 95% CI = 1.01 to 1.33) in Poisson regression. Racemic albuterol's advantage did not change in the sensitivity analysis controlling for potential confounders age less than 12 months, prematurity, hypoxia, and history of prior wheezing (aRR 1.18, 95% CI = 1.02 to 1.37) (Table 3).
A total of 187 patients met admission criteria at entry. Some of these were successfully discharged (see Table 2). Excluding those who met admission criteria on entry had little effect on the results (aRR 1.22, 95% CI = 1.04 to 1.43). Interrater reliability of the severity assessment tool (assessed only at the primary site) was substantial (j = 0.68). 27
The data required to calculate the severity of illness were incomplete in 62 of 703 (8.8%) cases. Risk of discharge for this group was indistinguishable from the mild severity group (aRR = 0.89, 95% CI = 0.73 to 1.09). Neither excluding these cases nor using multiple imputations changed the results (analysis not shown).
Including only infants less than 12 months, using age less than 12 months as a categorical variable or excluding subjects aged less than 2 months did not change the significant advantage of racemic albuterol over racemic epinephrine. Patients with recurrent wheezing were slightly less likely to be admitted (Table 3).
Tachycardia (defined by heart rate above the 97th percentile for age) was noted before treatment in 88 of 352 (25%) in the racemic albuterol group and 74 of 351 (21%) in the racemic epinephrine group. Posttreatment tachycardia was only noted in 52 (15%) in the former and in 42 (12%) in the latter (p = 0.32). Two patients
Table 1 Baseline Patient Variables Variable Epinephrine (n = 351) Albuterol (n = 352) p-Value (Univariate Testing) Median age in months (IQR) 4.8 (6.2) 4.9 (6.1) 1.00 Male, N = 407 (57.9%) 213 (60.7) 194 (55.1) 0.15 Ex-premature (%) 61 (17.4) 69 (19.6) 0.50 Asthma or RAD in any member of immediate family 70 (19.9) 69 (19.6) 0.93 On albuterol before arrival in the ED, n (%) 60 (17.1) 64 (18.2) 0.77 On prednisone before arrival in the ED, n (%) 15 (4.3) 19 (5.4) 0.66 Temperature ‡ 38°C (%) 129 (36.8) 140 (39.8) 0.62 Age £ 2 months (%) 76 (21.7) 75 (21.3) 1.00 Age > 12 months (%) 52 (14.8) 41 (11.7) 0.22 First-time wheezing (%) 198 ⁄ 280 (70.7) 205 ⁄ 287 (71.4) 0.85 RSV antigen-positive (%) 152 ⁄ 276 (55.1) 161 ⁄ 277 (58.1) 0.49 SaO 2 £ 92% (%) 35 ⁄ 344 (10.2) 25 ⁄ 342 (7.3) 0.22 SaO 2 (±SD) 97 (4) 97 (3) 0.63 Mean temperature, °F (±SD) 100.0 (1.7) 99.9 (1.7) 0.55 Mean HR (±SD) 156 (24) 158 (23) 0.35 Mean RR (±SD) 47 (13) 46 (12) 0.50 Tachycardic (%) 74 ⁄ 350 (21.1) 88 ⁄ 352 (25.0) 0.25 Hydration status (%) n = 342 n = 346 Normal 303 (88.6) 303 (87.6) 5% dehydration 24 (7.0) 38 (11.0) 0.02 10% dehydration 14 (4.1) 5 (1.5) 15% dehydration 1 (0.3) 0 (0) Increased work of breathing retraction severity (%) n = 313 n = 331 None 20 (6.4) 27 (8.2) Mild 162 (51.8) 161 (48.6) 0.62 Moderate 115 (36.7) 130 (39.3) Severe 16 (5.1) 13 (3.9) ED = emergency department; HR = heart rate; IQR = interquartile range; RAD = reactive airway disease; RR = respiratory rate; RSV = respiratory syncytial virus; SaO 2 = oxygen saturation. developed perioral cold-induced urticaria. This resolved rapidly, and one received promethazine. Both were in the racemic epinephrine group and had used cylinder oxygen. One death in the racemic epinephrine group was reported within 30 days.
Our crude results did not show a difference between agents. When we adjusted for severity, we found a lower risk of admission with the use of nebulized racemic albuterol when compared to racemic epinephrine in children presenting to the ED with bronchiolitis. This was an unexpected result and contrasted with some, 11,15,28 although not all, [7][8][9]16 existing studies. However, to the best of our knowledge, our study is the largest to date examining this question. Our results held true in all subgroups: hypoxia, age less than 12 months, prematurity, and a family history of asthma do not alter the advantage of racemic albuterol over racemic epinephrine in reducing admissions. Broad inclusion criteria helped ensure adequate sample size. Unlike many smaller studies, our sample represents the breadth of bronchiolitis seen in the ED each season. One of the ironies of bronchiolitis research is that while bronchiolitis is agreed to be common, most studies of it are small. This is in part due to difficulties in defining the condition and distinguishing it from asthma. The term asthma is itself so broad and illdefined that at least one editorialist has called for the term to be classified as a symptom rather than a diagnosis. 29 Furthermore, it is doubtful if asthma should ever be diagnosed in the age group that we studied. Our definition of bronchiolitis is consistent with the often heterogeneous 28 definitions used by others. 21 We have also addressed elements seen in other definitions by both subgroup and multivariate analysis and consistently obtained results favoring racemic albuterol.
Comparing pretreatment groups by the probability of admission sets a substantially higher bar than simply comparing the distribution of individual variables between treatment arms. The latter method can lead to the appearance of balance when in fact the preintervention probability of admission differs between groups.
The former requires a validated tool to predict admission and an adjusted analysis in the event that groups are not balanced or much larger sample sizes to increase the probability of treatment group balance or subject pair matching at enrollment. From a practical standpoint, implementation of subject-pair matching would require an automated system deployed across several aggressively recruiting sites that allows rapid input of patient data, classifies severity of disease, and finds a matching case. Such a solution is clearly resource-intensive. We opted for simple enrollment and subsequent adjustment for disease severity using a validated tool, which was feasible and costeffective and minimized barriers to recruitment by treating clinicians.
Had we ignored severity of illness, or assumed equal pretreatment probability of discharge in each treatment arm, we would have erroneously concluded that there was no difference between the drugs. Our design addressed the limitation that two relatively large groups may have a similar distribution of potentially confounding variables overall in each group, but have different combinations of these confounders at the level of individual subjects.
Failure to address this can result in erroneous interpretation of RCTs, as has emerged when large trials were reanalyzed. 30 We had anticipated this possibility in the design stage and decided a priori to apply a severity-of-illness tool in our analysis to adjust for the effect of severity of illness on the probability of discharge within this sample. Thereafter, we evaluated the additional effect of the drugs on discharge risk. We consider this methodology to be a design strength compared with previous RCTs. 30 The choice of severity-of-illness tool is important. We used the National Children's Hospital (NCH) severityof-illness tool because it has been validated with respect to need for admission. The relative complexity of this tool precluded stratification prior to enrollment. The more widely used RDAI has been reported as having almost perfect agreement (j = 0.9) for the wheezing component and substantial agreement (j = 0.64) for the presence of retractions component in its original description. An overall kappa was not reported. 12 We and others have found worse agreement for auscultatory findings in infants, 21,31 but that overall the reliability of our severity-of-illness tool (j = 0.68) was sufficient to permit its use. Previously employed severity-of-illness or clinical scores have not been validated to predict disposition in bronchiolitis, whereas the model we used has. 19 Our data show racemic albuterol leads to lower discharge rates in each severity-of-illness stratum (Figure 6), but that when these strata are combined there is no difference. The Yule-Simpson ''paradox'' describes just this situation, where the success of several subgroups may be reduced or even reversed when these subgroups are combined. 32 The conditions required for the ''paradox,'' namely, the combination of an imbalance in the proportion of each subgroup receiving each intervention and a different event rate in each subgroup, are present in our data. 33 This combination requires inclusion of this confounding variable in a multivariate regression analysis. 34 Almost identical results were obtained using Mantel-Haenszel stratified analyses.
Another difference between our study and others is that determination of disposition was at 2 hours following the initial nebulized treatment. This reflects the widespread clinical practice of observing patients following epinephrine, rather than immediately discharging them. Our observation period ensured that disposition was not prematurely decided based on initial improvement. Studies have shown a benefit in clinical scores in the initial 15-60 minutes for infants receiving epinephrine, 13,35 but overall this benefit is short-lived. 17 This study, however, does not exclude a role for epinephrine. Further research is needed to address which subsequent agent should be used following inadequate response to albuterol or whether there is a synergistic effect to using both.
The age cutoff for inclusion in this study was a compromise between the widely accepted definitions of less than 1 or 2 years of age. Reducing the upper age limit could reduce the enrollment of patients with recurrent wheezing and perhaps some of those who may subsequently develop asthma. We reanalyzed the sample population introducing age less than 12 months and first presentation as model variables. We also reanalyzed the sample after excluding patients over 12 months. In these alternative models, we continued to find results favoring racemic albuterol. Similarly, a history of recurrent wheezing or a family history of asthma did not change the treatment effects. Consequently, despite our broad inclusion criteria, our results likely hold across even the subsets of children presenting to the ED with bronchiolitis, subsets represented by some prior, more narrowly defined, but inconclusive studies.
The number of eligible patients not enrolled in the study was unknown, as the primary site IRB did not permit the collection of any information about patients either who were not enrolled due to refusal of consent or who had not been approached for consent. Anecdotal experience suggests that this number was small at the primary site. We were not able to determine the specific reason for admission, as we did not require physicians to specifically document this.
The relatively low sample size in the secondary site could be attributed to physician difference in diagnosing bronchiolitis or the availability of researchers. While it likely represents a convenience sample, the results of the secondary site paralleled those at the primary site. The numbers at the primary site, where the recruitment period was longer, were a reflection of the efforts of the investigators (who actively recruited patients from the waiting room), availability of RAs up to 16 hours a day, and widespread knowledge of the study throughout the community and the emergency medical services. At the secondary site, a research nurse was available 40 hours a week. We addressed this difference in site recruitment by including site as a variable in the model, where it approached statistical significance in its own right but did not alter albuterol's advantage.
Ideally, we would have performed stratified randomization, thereby avoiding the need for stratified or adjusted analysis. However, the severity-of-illness tool we used is complex and applying it at the randomization stage would have posed a significant barrier to recruitment. It could also introduce bias if an investigator was enrolling the patient and knew the severity score.
We accounted for missing data by creating a fourth severity-of-illness category, by excluding these cases, and by estimating their missing values using multiple imputation. 35,36 In all cases, the treatment effects were essentially the same regardless of statistical strategy. Nonetheless, this is much less desirable than having completed data sheets. This is particularly the case as the number of incomplete data sheets was higher in the racemic epinephrine than the racemic albuterol group, potentially biasing the results. The nonsignificant aRR of the incomplete data group means that it cannot be distinguished from the mild group, potentially indicating more mild-type cases in the epinephrine group. Neither dropping these cases nor multiple imputations for these missing values changed our results. Some patients could have received more or less racemic albuterol than if we had dosed on a strictly milligram per kilogram fashion, and some may regard our doses as low. This could bias our results against albuterol.
The last dose of active medication in the albuterol group was administered closer to the time that the disposition decision was made than in the racemic epinephrine group. This would bias the results against epinephrine assuming that racemic albuterol has an advantage over saline.
Using a single dose of racemic epinephrine, rather than three doses, reflects clinical practice. The dose of racemic epinephrine (11.25 mg regardless of weight; others have used 0.15 to 0.9 mg ⁄ kg) 15,28 is comparable to that used for treating croup. This relatively high dosage minimized the potential disadvantage of using a single dose.
All patients received saline mist during the consent process prior to randomization. If saline mist has a benefit, this would bias the study to the null. We used saline nebulizers following the racemic epinephrine treatment to maintain blinding. Again, any beneficial effect of saline would have biased the study to the null.
Racemic albuterol was prescribed equally among those discharged from each treatment arm. This would be expected to narrow the differences in unscheduled admissions between treatment groups following discharge and increase the likelihood of a Type 2 error.
We would have liked to control for discharge medications, but to achieve this would have required unanimous agreement from all the EPs, pediatricians, and family physicians serving the study site and would likely have prevented the study from proceeding. We found discharge medications to be similar between groups.
In children up to the 18th month of life presenting to the ED with a clinical diagnosis of bronchiolitis, racemic albuterol, rather than racemic epinephrine, should be the initial agent chosen, as doing so modestly increases the rate of successful discharge.
*Reference category is mild disease. n for this model is 686. The same results for severity of illness and drug were obtained using Mantel-Haenszel stratified analysis.
ACAD EMERG MED • April 2008, Vol. 15, No. 4 • www.aemj.org
The authors acknowledge the assistance of the following:
B iological electrophiles result from oxidative metabolism of exog- enous compounds or endogenous cellular constituents, and they contribute to pathophysiologies such as toxicity and carcinogenicity. The chemical toxicology of electrophiles is dominated by covalent addition to intracellular nucleophiles. Reaction with DNA leads to the production of adducts that block replication or induce mutations. The chemistry and biology of electrophile-DNA reactions have been extensively studied, providing in many cases a detailed understanding of the relation between adduct structure and mutational consequences. By contrast, the linkage between protein modification and cellular response is poorly understood.
In this Account, we describe our efforts to define the chemistry of protein modification and its biological consequences using lipid-derived R,β-unsaturated aldehydes as model electrophiles. In our global approach, two large data sets are analyzed: one represents the identity of proteins modified over a wide range of electrophile concentrations, and the second comprises changes in gene expression observed under similar conditions. Informatics tools show theoretical connections based primarily on transcription factors hypothetically shared between the two data sets, downstream of adducted proteins and upstream of affected genes. This method highlights potential electrophile-sensitive signaling pathways and transcriptional processes for further evaluation.
Peroxidation of cellular phospholipids generates a complex mixture of both membrane-bound and diffusible electrophiles. The latter include reactive species such as malondialdehyde, 4-oxononenal, and 4-hydroxynonenal (HNE). Enriching HNE-adducted proteins for proteomic analysis was a technical challenge, solved with click chemistry that generated biotin-tagged protein adducts. For this purpose, HNE analogues bearing terminal azide or alkyne functionalities were synthesized. Cellular lysates were first exposed to a single type of HNE analogue (azido-or alkynyl-HNE), and then click reactions were performed against the cognate alkynyl-and azido-biotin derivative. The resulting biotin-labeled proteins were captured and enriched over a streptavidin matrix for subsequent mass spectrometric analysis. We thereby identified a multitude of HNE targets. Simultaneous microarray analysis of changes in gene expression triggered by HNE also produced an abundance of data. Functional analysis of both data sets generated the hypothesis that an important pathway of cellular response derives from electrophile modification of protein chaperones, resulting in the release of transcription factors that are their clients. Informatic analysis of the protein modification and microarray data sets identified several transcription factors as potential mediators of the cellular response to HNE-adducted proteins. Among these, heat shock factor 1 (HSF1) was confirmed as a sensitive and robust effector of HNE-induced changes in gene expression. Activation of HSF1 appears, in part, to be mediated by the electrophilic adduction of Hsp70 and Hsp90, which normally maintain HSF1 in an inactive cytosolic complex.
The identification of HSF1 as a mediator of biological effects downstream of HSF1 has provided new opportunities for research, illustrating the potential of our systems-based approach. Accordingly, we characterized HSF1-mediated gene expression in protecting against electrophile-induced toxicity. Among the genes induced by HSF1, Bcl-2-associated athanogene 3 (BAG3) is notable for its actions in promoting cell survival through stabilization of antiapoptotic Bcl-2 proteins, appearing to have a critical role in mediating cellular protection against electrophile-induced death.
Cells respond to a diverse array of environmental and endogenous stimuli through cell surface and intracellular receptors. Once engaged, these receptors trigger signaling cascades that culminate in a cellular response. These communication systems represent networks developed through evolution that enable cellular adaptation to the surrounding environment. In addition to these receptor-mediated pathways, cells can also respond to nonspecific challenges presented by reactive oxidants and electrophiles. The structures and comparative reactivities of these agents are highly diverse, so as a general rule specific receptors for reactive intermediates have not evolved. Rather, generic mechanisms exist that mediate cellular responses to chemical stress. The signaling pathways activated by these reactive species either enhance cytoprotective processes or trigger cell death in the case of overwhelming stress.
Electrophiles represent a significant threat to cellular law and order because they react with a multitude of intracellular nucleophiles including DNA, RNA, phospholipids, and proteins. Electrophiles are generated during the enzyme-mediated oxidation of foreign compounds (toxicants, foods, pharmaceuticals), as well as by the oxidation of endogenous biomolecules (lipids, amino acids, carbohydrates). Much of our knowledge regarding the biochemistry of electrophiles comes from extensive research into their role in carcinogenesis. Decades of work have addressed various aspects of electrophile stress within this context and have helped to define the mechanisms for electrophile generation from various sources; the covalent modification of DNA bases; the biochemical mechanisms of mutagenesis and repair; and the correlation between electrophilic DNA adducts and cellular transformation (Figure 1). These investigations have been aided by a powerful approach in which single DNA adducts are incorporated into a viral genome or shuttle vector, and the fates of the adducts are subsequently evaluated in intact cells. 1 This method has provided direct correlations between the structures of DNA adducts, their metabolism in vivo, and their mutagenic potential.
Less effort has been devoted to the study of macromolecular targets other than DNA. However, it is increasingly apparent that they represent important sites for electrophilic adduction with respect to both the degree and the biological significance of their modification. 2 In addition, there is a growing appreciation that cells are not passive targets for electrophilic damage but can adapt and protect themselves from subsequent stress. Much of this heightened interest derives from improved technologies that have enabled researchers to better address the role of macromolecular modification (especially protein modification) in sensing and mediating biological responses to reactive species. For example, the reactions between electrophiles and amino acid side chains are relatively well-defined, and techniques exist for detecting electrophilic modifications to individual peptides or proteins. 3 However, establishing a definitive link between the adduction of a specific protein and a biological response has proven to be a more challenging task. Some progress has been made by examining proteins that participate in well-defined signaling pathways and addressing the effect of their modification on the function of the signaling network. For example, modification of IκB kinase by the reactive electrophiles ∆ 12,14 -15deoxy-PGJ 2 , 4-hydroxynonenal (HNE), or parthenolide blocks activation of the NF-κB signaling pathway 4-6 (Figure 2). Modification of IκB kinase interferes with the release of NF-κB subunits from IκB and attenuates NF-κB-mediated gene expression. 4 Valuable information has been gleaned from such focused investigations, but the pace of discovery has been slow. Moreover, the relative importance of individual signaling pathways within the context of the overall cellular response is difficult to assess in this manner. Recent advances that enable large sets of modified proteins to be identified, combined with global analyses of gene expression, have provided an opportunity for a systems-based approach to study electrophile stress at the cellular level. We will describe our initial attempts to construct such an approach within the context of our ongoing investigations, systematically examining the cellular responses to electrophiles generated from oxidized lipids.
Research in our laboratory has focused on electrophiles generated as a result of glycerophospholipid peroxidation, the spontaneous oxidation of the unsaturated fatty acyl side chains esterified to the glycerol backbone (Figure 3). Oxidants generated by a variety of pathways remove bis-allylic hydrogen atoms to generate pentadienyl radicals that are scavenged by O 2 to form lipid peroxyl radicals. The lipid peroxyl radicals remove a hydrogen atom from a neighboring poly-unsaturated fatty acid residue to propagate the radical chain and produce a fatty acyl hydroperoxide bound to the phospholipid. 7 The sn-2 position of all membrane glycerophospholipids consists of mono-or polyunsaturated fatty acids, so the potential for lipid peroxidation is enormous. In fact, by monitoring certain products of lipid peroxidation as their urinary metabolites, it has been unequivocally established that phospholipid peroxidation occurs continuously in humans and can be increased by oxidant challenges such as cigarette smoking or xenobiotic exposure. 8 Reduction of the initially formed
fatty acyl hydroperoxides by one-electron reductants generates alkoxyl radicals that decompose to a plethora of products, some of which contain reactive functional groups such as epoxides and aldehydes. Depending on the chemistry of hydroperoxide decomposition, fragmentation products can be produced in which the electrophile is released and diffuses throughout the cell or in which the electrophile remains bound at the sn-2 position of the phospholipid.
We are particularly interested in the chemical biology of R,β-unsaturated aldehydes, such as malondialdehyde and HNE. Research from several groups, including our own, has established that these electrophiles react with DNA to generate mutagenic adducts that are detectable even in the genomes of healthy individuals. 9 In addition to DNA, malondialdehyde and HNE react with proteins. Lipid-derived adducts to protein are detectable in mammalian tissue and can alter the properties of the target proteins. 2 The reactivity of HNE within the cell is affected by such diverse factors as pH, glutathione content, and protein concentration, which together influence the complement of adducted proteins and degree of modification. Until recently, our understanding of HNE biology was mainly derived from the study of individual target proteins on a case-by-case basis. For example, the contribution of HNE to pathological events such as neurodegeneration, pain, inflammation, and cellular aging has largely been defined through the study of individual protein targets. 5,[10][11][12]
We have initiated a program to relate protein modification by electrophiles to cellular responses in a global fashion. The approach is outlined in Figure 4. Cells are treated with an electrophile (in this case HNE), and two large data sets are generated. The first is an inventory of the proteins modified by HNE, and the second is a compilation of the gene expression changes determined by microarray analysis. Transcriptional activation or inhibition is not the only response of a cell to stress, but it integrates many changes in signaling networks that culminate in an ultimate cellular outcome. These protein modification and gene expression data sets are then linked by informatic analysis of the signaling networks engaged. For example, by monitoring changes in gene expression, one can infer which transcription factors are either activated or inhibited during the cellular response to electrophile treatment. Extrapolating data in this manner can help to identify which signaling networks are affected by the electrophile. The next step is to examine protein modification data, to ask if one can plausibly connect a HNE-modified protein to the affected signal transduction pathways. This enables one to formulate hypotheses linking protein modification to transcriptional response that can be tested using molecular biological and biochemical approaches (e.g., reporter-based assays of gene expression, siRNA knockdown, etc.). Major challenges presented by this approach include developing the capture chemistry (the method by which the modified proteins are enriched for subsequent identification) and bioinformatic analysis (the tools used to link protein modification to transcriptional changes).
Reaction of HNE with proteins occurs primarily by Michael addition to histidine, cysteine, and lysine residues (Figure 5). 13 The initial Michael adducts cyclize to hemiacetals that have the capacity to cross-link lysine residues. HNE also reacts with lysine to form a Schiff base, but this is quantitatively less significant than Michael addition. The Schiff base can cyclize and dehydrate to form a stable pyrrole adduct. 14 In a highly oxidizing cellular environment, protein cross-linking reactions involving histidine or lysine residues has also been demon-
strated to arise from the bifunctional generation of both Michael and Schiff base adducts. Until recently, the identification of HNE-modified proteins relied primarily on immunochemical methods for adduct capture or detection. These methods are powerful, but they have significant limitations. Many antibodies are raised against specific HNE-amino acid adducts, so they may not recognize a broad range of adducts, whereas some antibodies may cross-react with proteins or adducts generated by oxidants or electrophiles other than HNE.
In order to supersede the use of antibody-based methods, we sought an approach that would be specific for HNE-modified proteins and would capture HNE adducts regardless of their chemical structure. Biotin hydrazide has been used to derivatize proteins modified by carbonyl-containing lipid oxidation products for enrichment with avidin-based matrices. [15][16][17] Although biotin hydrazide reacts with free protein aldehydes, it does not react with HNE adducts derived from the lysine Schiff base and it can react with adducts derived from other aldehydes and ketones. 18 We utilized click chemistry to biotin-label HNE-modified proteins. 19 Click chemistry describes the 1,3-dipolar cycloaddition reaction between azide (or alkyne) labeled probes to conjugate alkyne (or azide) labeled reporter tags. This approach has proven extremely useful for interrogating biochemical targets in the complex environment of the cell. 20 For our purposes, the probe is an analogue of HNE, modified at the terminal carbon with an azide (azido-HNE) or substituted at the ω and ω-1 positions with an alkyne (alkynyl-HNE) 19 (Figure 6). The attachment of either an azide or alkyne tag to HNE is a subtle change so that HNE, azido-HNE, and alkynyl-HNE exhibit comparable cytotoxicity and abilities to stimulate gene expression (Figure 6). We used click chemistry, streptavidin-based enrichment, and mass spectrometry to compile large data sets of protein targets of azido-HNE or alkynyl-HNE in the human colon cancer cell line, RKO. 19 Individual proteins were validated as HNE targets based on the concentration-response to HNE analogues, as well as on the specificity of their modification. Protein modification data were then used in conjunction with results from gene expression studies to generate hypotheses on the signaling processes affected by HNE and the resulting cellular consequences.
We examined the effects of HNE on gene expression in the RKO cell line under conditions that were similar to those used to evaluate protein modification, varying both concentration and time of compound exposure. 21 Changes in global transcript levels were monitored using microarray techniques. Clustering genes into groups regulated by a common signaling pathway proved useful in characterizing the cellular response to HNE. For example, mapping transcriptional changes based on upstream regulatory pathways revealed activation of the DNA damage (HDM2; TP53INP1), antioxidant (HMOX1; SCL3A2; GCLM; NQO1), ER stress (ASNS; CTH; 1). Activation of these pathways was individually confirmed using real-time PCR, Western blot analysis, and luciferase reporter assays.
We used a global approach to relate microarray and protein adduction data in a more directed manner. This was achieved using various software tools for sorting and mining both protein modification and gene expression data sets. For example, WebGestalt, which stands for "WEB-based Gene SeT AnaLysis Toolkit", is a program that integrates information from various public resources, uncovering genetic and biochemical relationships in complex data sets. 22 We have also employed GenMAPP, which facilitates the analysis of microarray data within the context of biochemical processes and human disease. 23 The results of these analyses are summarized in Figure summarized in Figure 7 illustrate the complexity of the chemistry and cellular responses to electrophile stress. It also identifies candidate transcription factors and pathways for study.
To prioritize our studies of the relation between protein modification and cellular responses, we constructed expression vectors containing individual response elements upstream of the luciferase gene, which we transfected into RKO cells prior to treatment with HNE. This enabled a semiquantitative comparison of the magnitude of the transcriptional response induced from different signaling pathways following HNE treatment. The most robust response was observed with expression vectors under the control of the heat-shock response element. This provided a biological rationale for efforts to define the mechanism of activation of the heat-shock signaling pathway.
The inventory of HNE-modified proteins contained several heat shock proteins, including Hsp60, GRP78, Hsp70 (Hsp72), and Hsp90. 19 Likewise, microarray data revealed a dramatic increase in heat shock-regulated transcripts in HNE-treated RKO cells. 21 The inducible expression of heat shock genes caused by heat or chemical stress is principally controlled by the latent transcription factor, heat shock factor 1 (HSF1). A fundamental step in the activation of HSF1 is nuclear translocation. 24,25 In the absence of stress, HSF1 is retained in the cytoplasm by inhibitory associations with Hsp70, Hsp90, and various cochaperones. 26 We demonstrated that HNE promotes the nuclear translocation of HSF1 using Western blot analysis, and showed the enhanced transcription of a luciferase reporter gene under the regulation of a conserved heat shock element. 27 We hypothesized that the process by which HNE enhances heat shock gene expression involves the modification of Hsp70, 90, and other chaperones, causing the release and nuclear translocation of HSF1. HNE treatment has been shown in vitro to adduct specific amino acid residues on both Hsp70 and Hsp90, which correlates with their reduced ability to bind and properly fold client proteins. 28,29 We performed coimmunoprecipitation experiments with myc-tagged Hsp70, demonstrating that its association with HSF1 is disrupted by HNE treatment. This occurs at HNE concentrations that affect Hsp70 modification, HSF1 nuclear translocation, luciferase expression from heterologous expression vectors, and heat shock gene expression. We are presently determining the sites of Hsp90 and Hsp70 modification and the functional consequences of HNE modification both in terms of HSF1 binding as well as its affiliation with other specific client proteins.
To assess the importance of the heat shock response in mediating the cellular response to electrophile stress, siRNA was used to silence HSF1, thereby attenuating heat shock gene expression in HNE-treated cells. 27 Control and HSF1-deficient cells were then exposed to HNE, and various responses measured including viability and gene expression changes. siRNA knockdown of HSF1 was nearly complete, and residual expression of the transcription factor was vanishingly small. Cells lacking HSF1 were profoundly more sensitive to the toxic effects of HNE than were cells that retained HSF1 (Figure 8). This suggests that HSF1-induced gene expression is an important protective response that is mounted by cells following electrophile treatment. In fact, in our studies, the cytoprotective role of HSF1 exceeds that of Nrf2 (the transcription factor responsible for the antioxidant response) based on a comparison of HNE toxicity and apoptotic markers in HSF1-
silenced and Nrf2-silenced cells. 27 This observation implies that heat shock, at least in our cellular model, has greater significance than the antioxidant response in abating cell death caused by exposure to reactive electrophiles. It also illustrates the value of simultaneously monitoring multiple signaling pathways rather than a single one.
The mechanisms underlying the reduced viability of HSF1deficient cells are varied, but the ultimate result is an increased sensitivity to HNE-induced apoptosis. HSF1-silenced cells showed dramatically elevated levels of JNK1 phosphorylation, which is a trigger for apoptosis, as well as decreased levels of Bcl-xL, which is an inhibitor of apoptosis. The reduction in Bcl-xL is due to diminished protein levels rather than reduced mRNA expression, which suggests that the inability of HSF1-silenced cells to mount a heat shock response leads to an increased turnover of Bcl-xL protein.
The molecular basis by which HSF1 attenuates cell death was examined in greater detail by microarray analysis, by comparing gene expression profiles between control and HSF1-silenced cells, following the addition of either vehicle (0.5% DMSO) or HNE. 30 Gene ontology (GO) analysis was performed in order to categorize HSF1-regulated genes by specific attributes (GO terms) defined by the GO Consortium. 31 GO terms fall under three broad categories, including "Cellular Component", "Biological Process", and "Molecular Function." For example, under the category of "Biological Process", we closely examined HSF1-dependent genes with either known or hypothetical antiapoptotic functions (Table 1). Although the expression of over 1000 transcripts showed some degree of dependence on HSF1, relatively few antiapoptotic transcripts (GO term: Negative Regulation of Apoptosis) were represented in the data; examples include CLU (clusterin); CRYAB (R,β-crystallin); HSPB1 (Hsp27), and BAG3 (Bcl-2-associated athano-gene 3). 30 At the concentrations of HNE evaluated, cell death occurs through an apoptotic pathway dependent on protein synthesis and involving cytochrome c release, caspase activation, and PARP cleavage. 32 Bcl-2 inhibits HNE-induced apoptosis by preventing mitochondrial pore formation and the resulting efflux of cytochrome c and other proapoptotic factors. Since BAG3 was identified by others as a Bcl-2-interacting protein, 33 we hypothesized that BAG3 induction by HSF1 plays a critical role in mitigating cell death, possibly by enhancing the expression or prolonging the half-lives of Bcl-2, Bcl-xL, and related antiapoptotic proteins. (Figure 9).
BAG3 belongs to a family of protein cochaperones (BAG1-6) that regulate diverse cellular processes, including proliferation, migration, and apoptosis. 34 To evaluate the proposed role of BAG3 in facilitating cell survival, siRNA was used to silence its induction in HNE-treated cells. Knockdown of BAG3 enhanced HNE-induced cell death to the same extent as silencing HSF1, confirming the importance of BAG3 in the heat-shock-mediated defense against reactive electrophiles. We also observed that silencing BAG3 was associated with a dramatic loss in proteins belonging to the antiapoptotic Bcl-2 family, including Bcl-2, Bcl-xL, and Mcl-1. There was no reduction in the levels of mRNAs for any of the genes, suggesting that the reduction in their protein levels is due to enhanced protein turnover. Mcl-1 is of particular interest because it is one of the 10 most upregulated genes across all cancers and accounts for the resistance of many neoplasms to chemotherapy and radiation. 35 The mechanism we propose for the cytoprotective actions of BAG3 involves the stabilization of antiapoptotic Bcl-2 family members, perhaps by impeding their proteasomal or autophagic turnover. 30 Recent reports suggest that BAG3 is an important factor in carcinogenesis and tumor cell viability. [36][37][38] Growing recognition of a role for BAG3 in cancer suggests that its induction mediated by HNE
or other electrophiles derived from oxidative stress may facilitate or promote tumorigenesis. We are currently investigating the importance of BAG3 and other HSF1-regulated genes in mediating tumor cell viability and resistance to apoptosis.
The results of our combined studies of HNE-induced protein modification and gene expression reveal the complexity of cellular responses to electrophile stress. There are many proteins modified and many signaling pathways engaged. The ultimate cellular response reflects the input of multiple signaling networks and does not result from a key "molecular target" or dedicated signaling pathway. This contrasts with studies of cellular responses to electrophilic natural products (e.g., fumagillin) where specific molecular targets dictate the cellular response. A key difference between lipid electrophiles and natural products is the relative structural simplicity of the former. Molecules such as HNE contain electrophilic centers that are relatively unhindered and can, therefore, react with a variety of protein targets, and in some cases at multiple sites. In contrast, most natural products contain a high degree of structural complexity that limits access of the electrophilic center to relatively few proteins that possess complementary binding pockets. Although HNE reacts with a wide range of proteins, it and other simple electrophiles still modify only a small fraction of the total cellular proteome. Moreover, at concentrations that induce programmed cell death, many targets are modified at only one or two sites. 3 Thus, considerable selectivity in the reactivity of the proteome is observed even with such "nonspecific" electrophiles as HNE. It also appears that stability of the protein adduct is an important determinant of the ultimate cellular response to electrophiles: more stable adducts are associated with greater toxicity. 39 Judging from our experience with HNE, cellular responses to electrophiles appear to be graduated based upon concentration and time of exposure. The gene expression changes induced by HNE treatment of RKO cells indicate that responses at low concentrations are mainly adaptive (e.g., heat shock response, antioxidant response) and are designed to protect against further damage. At higher concentrations, the adduct load on protein and DNA overwhelms these protective mechanisms and the cells undergo apoptosis. Finally, at extreme concentrations of electrophile, cells undergo a necrotic cell death. Despite the complex nature of the cellular response to electrophiles, the approach outlined in our case study proved effective in defining individual elements in the process of cellular adaptation. siRNA knockdown of key gene products (e.g., HSF1) coupled with detailed follow-up studies not only highlight the importance of the pathway of inter-est but can be used to decipher the mechanistic details of the overall cellular response.
Our decision to closely examine the heat shock response was motivated by two observations. First, we found that many heat shock genes were dramatically induced in HNE-treated RKO cells.
Second, analysis of luciferase reporter constructs containing heat shock response elements confirmed that HNE promotes HSF1dependent gene expression. Our hypothesis that HNE modification of Hsp70 or Hsp90 leads to the release and activation of HSF1 is supported by a reasonable body of in vitro and in vivo results. 19,[27][28][29]40,41 However, it is also possible that modification of other cellular proteins leads to their unfolding, which attracts Hsp90 away from HSF1. Experiments are underway to test our hypothesis that modification of Hsp70 and Hsp90 is responsible for the HNE-induced heat shock response. A corollary of our hypothesis is that abundant cellular proteins such as Hsp70 and Hsp90 represent important targets that, when modified, can exert a profound cellular response. This is particularly important because, in addition to HSF1, Hsp90 has legion client proteins that control a wide variety of cellular functions. If the modification of Hsp90 indeed liberates client proteins, the potential for HNE and other relatively simple electrophiles to influence cellular function is great.
Our studies also highlight potential mechanisms by which tumor cells protect themselves from cytotoxic challenge presented by the innate immune system. Recent studies show that tumor cells can adapt to reactive oxidants and unfolded protein stress. 42 This adaptation manifests as a resistance to the cytotoxic insults of neutrophils and macrophages, plus a variety of chemotherapeutic agents. The heat shock response, mediated by HSF1, appears to be a major contributor to survival during such types of stress. This observation is substantiated by experiments performed in mice, where genetic deletion of HSF1 dramatically reduces the induction of skin tumors by the mutagenic and tumor-promoting combination of dimethylbenzanthracene and tetradecanoylphorbolacetate. 43 Further work must be performed to reveal the processes that make HSF1 both cytoprotective and protumorigenic. Our demonstration that BAG3 is essential in the HSF1mediated resistance of RKO cells to electrophile-mediated cell death stress suggests BAG3 may be a feasible target to sensitize cancer cells to radiation, chemotherapy, and immune-mediated toxicities. However, additional genes are undoubtedly involved in HSF1-mediated resistance to electrophiles. We anticipate that some will be revealed through our ongoing global analysis of electrophile responses, and should contribute further to our understanding of cellular adaptations to stress.
The research summarized in this account is funded as part of a program project from the National Institutes of Health (ES13125) and a center grant from the National Foundation for Cancer Research. We are deeply appreciative of the interactions with and efforts of our coinvestigators and colleagues on this program, specifically Ned Porter, Daniel Liebler, Jack Roberts, Bing Zhang, Keri Tallman, Simona Codreanu, Colleen McGrath, Jody Ullery, Rebecca Connor, and Mariana Boiani. We also acknowledge the efforts of former members of the Marnett laboratory, Chuan Ji, James West, and Andrew Vila, who built the foundation for this program.
Cellular Responses to Electrophiles Jacobs and Marnett
Vol. 43, No. 5 May 2010 673-683 ACCOUNTS OF CHEMICAL RESEARCH 675
Vol. 43, No. 5 May 2010 673-683 ACCOUNTS OF CHEMICAL RESEARCH 679
Vol. 43, No. 5 May 2010 673-683 ACCOUNTS OF CHEMICAL RESEARCH 681
BIOGRAPHICAL INFORMATION Aaron T. Jacobs received his B.S. degree in biology from UC Irvine in 1993 and a Ph.D. in pharmacology from UCLA in 2003 under the guidance of Louis J. Ignarro. He then pursued postdoctoral studies with Lawrence J. Marnett at Vanderbilt University until recently joining the faculty at the University of Hawaii, Hilo College of Pharmacy. Lawrence J. Marnett received a B.S. in Chemistry from Rockhurst College in 1969 and a Ph.D. in Chemistry from Duke University in 1973 under the direction of Ned Porter. After postdoctoral research with Bengt Samuelsson at Karolinska Institute and A. Paul Schaap at Wayne State University, he joined the faculty in Chemistry at Wayne in 1975. He moved to Vanderbilt University in 1989 where he is currently University Professor, Mary Geddes Stahlman Professor of Cancer Research, Professor of Biochemistry, Chemistry, and Pharmacology, and Director of the Vanderbilt Institute of Chemical Biology.
▶ 8.4% of Thai adults reported a transport related injury in the previous 12 months. ▶ Risk was higher for males and young adults and motorcycles were commonly involved. ▶ Males were much more likely to report drink driving than females. ▶ The prevalence of seat belt and helmet wearing was higher than previously reported. ▶ We will monitor changes in transport injury risk and related behaviour in this cohort.
Traffic injuries are an important contributor to the national disease burdens of middleincome countries, and are estimated to cost 2% of the Gross Domestic Product (Roberts, 2004). The increasing death and disability from road trauma in middle-income countries can be largely attributed to massively increasing motorisation in the context of inadequate infrastructure, vehicular safety and safe system regulation (Peden et al., 2004). It can be viewed primarily as a development issue, as it is both a direct consequence of increasing industrialisation and modernisation and a huge constraint on development itself (Hyder and Ghaffar, 2004;McMahon and Ward, 2006;World Health Organization, 2009).
In Thailand, a middle-income "transitioning" country, there has been a dramatic upward trend in transport-related injuries and deaths that parallels rapid economic development and increasing motor vehicle ownership and use. The population injury rate from road traffic collisions has increased from 17 per 100,000 in 1984 to 152 per 100,000 in 2005 (Wibulpolprasert, 2008). Concurrently, traffic injury has become a leading cause of death in Thai males and females aged between 15 and 45 years, with over 5000 young adults dying each year. Road traffic injuries also cause the most permanent disability in Thailand (Sitthiamorn et al., 2007).
A major barrier to addressing transport injury in low to middle-income countries is the lack of information about the nature and extent of non-fatal trauma and the prevalence of behavioural safety indicators and compliance with legislation, e.g. helmet and seat belt use and drink driving (Odero et al., 1997;WHO, 2009).
While Thailand has a stronger injury surveillance system than many middle-income countries, these issues still apply to the road safety data. Data on road traffic collisions in Thailand comes from three sources: hospital data which are collected intermittently from a small number of hospitals; police data which lack a standardised recording system and key information; and data from the Traffic Engineering Division, Department of Highways which cover only a quarter of national roads and largely rely on police reports (Suriyawongpaisal and Kanchanasut, 2003). The large community-based Thai National Injury Survey of 2003/2004 reported on over 300,000 Thais nationwide, however it recorded only injuries needing medical treatment and/or 3 days or more off work. It also relied on one respondent reporting for the whole household (Sitthi-amorn et al., 2007). Consequently, many minor injuries may have been missed.
The aim of this study is to determine the baseline frequency and distribution of transport injury and the prevalence of various road safety behaviours in a newly recruited cohort of Thai adults. This information is intended to serve as a baseline for monitoring changes in traffic crash risks and risk behaviours in response to ongoing implementation of policy and programs to address the problem of transport-related injury in Thailand.
The Thai Health-Risk Transition Study includes an ongoing Thai Cohort Study (TCS) of 87,134 adult Open University students residing across all regions of Thailand. The cohort comprises distance-learning students enrolled at Sukhothai Thammathirat Open University (STOU) which is referred to as an open university because it does not require high school graduates to pass an entrance test. The baseline TCS data were collected in 2005 and include information on transport injury and a wide array of demographic, socio-economic, behavioural and transportation factors that could be linked to Thailand's transport risks.
Details on population selection and methodology have been reported elsewhere (Sleigh et al., 2008). Briefly, the 2005 student register listed approximately 200,000 names and addresses: a 20-page questionnaire was mailed out to each student and 87,134 (44%) responded.
2.2 Measures 2.2.1 Transport injury-All respondents were asked, 'In the last 12 months how many injuries have you had that were serious enough to interfere with daily activities and/or required medical treatment?' For their most serious injury, respondents were asked, 'where were you when you were injured' and 'was this injury related to transport'. Location of the most serious injury was coded as home, road, sports facility, agricultural workplace, nonagricultural workplace or other. If the most serious injury was related to transport then we ascertained the respondent's role (driver, passenger, pedestrian), and, for drivers and passengers, the vehicle they were in (bicycle, motorbike, bus-van, car-pickup, other vehicle).
Transport-related risk behaviours-We asked the entire cohort whether they had driven a motor vehicle after 3 or more glasses of alcohol in the previous 12 months. We categorised respondent use of safety devices for motorbikes (helmets) and cars (seat beltsback and front seats), as never, sometimes or always for the whole cohort.
Data scanning, verifying, and correcting were conducted using Scandevet, a program developed by a research team from Khon Kaen University. Further data editing was completed using SQL and SPSS software and for analyses we used SPSS and Stata. Descriptive data regarding the distribution of transport injury by sex, age and residence are presented. For those respondents that reported transport injury, their role and the vehicle involved in the injury event are presented by sex and age. Descriptive data regarding the age-sex distribution of transport-related risk behaviours for the whole cohort were also derived.
Ethics approval was obtained from Sukhothai Thammathirat Open University Research and Development Institute (protocol 0522/10) and the Australian National University Human Research Ethics Committee (protocol 2004344). Free and informed written consent was obtained from all participants.
The TCS cohort was 54.7% female, with a median age of 29 years. Compared to the overall STOU student population, the respondents were slightly older, but similar in terms of sex, education level, monthly income and geographic residence. In comparison to the Thai adult population, the cohort had a slightly higher proportion of females, a greater proportion of young to middle aged adults (21-40) and less in the younger (<21) and older (>40) agegroups. They were substantially more educated. Average monthly income compared to the Thai population is difficult to judge due to the large proportion that had no income or did not report their income in the Thai population survey. It appears that the TCS may be a little better off, however their average income is still quite modest. The geographic distribution of the cohort, however, was similar to the Thai population (Seubsman et al., in press) (Table 1).
Overall 7279 (8.4% or 8354 per 100,000) of cohort respondents reported a transport-related injury as their most serious injury. In all age groups, males reported more transport injury than females. Young adults were more likely to report transport injury than older adults (Fig. 1). There was little difference in the proportion who reported transport injury between country residents (8.6%) and city/town residents (8.1%).
Fig. 2 explores the roles the injured respondents played on the occasion of their transportrelated injury. Of the 6371 injured persons who reported their role, there were 4514 (70.9%) drivers, 1459 (22.9%) passengers and 398 (6.2%) pedestrians. In all age groups, a higher proportion of injured males were likely to be drivers compared to injured females, with the reverse being true for passengers. A higher proportion of females (7.2%) experiencing transport-related injury were pedestrians compared to males (5.4%) and this injury category was the only one with a notable age-effect, with the proportion of injured that were pedestrians increasing with age for both sexes, reaching 15.6% among women over 50.
Among the injured, 71.9% were riding motorcycles when their transport injury occurred (Fig. 3). This proportion decreased in older age groups but motorcycles were still the most common vehicle involved in transport injury events in all age groups in both sexes (70.8% for men, 73.2% for women). The proportion injured in cars or buses and vans increased with age for both sexes while the proportion injured while riding bicycles decreased with age for men but not for women. Injuries sustained while a driver or passenger in other types of vehicles (e.g. train, boat or airplane) were uncommon.
The distribution of transport-related risk behaviours was determined for the whole cohort, in order to better target interventions. Of the respondents who reported using a motorbike, approximately 6% of males and 10% of females rarely or never wore a helmet (Fig. 4). Overall, females were less likely to wear a helmet than males, with the highest proportion of non-wearers being women over 50 years of age (18.9%).
Of the respondents who had a front seat belt available for use in the vehicle, 72.2% of males and 66.6% of females reported always wearing the front seat belt and this increased with age (Fig. 5). Teenagers had the highest proportion of respondents who reported never wearing a front seatbelt (10% of males, 9% of females).
In contrast, back seat belt use was low (Fig. 6). Of those who said the vehicle had a back seat belt available, over 50% of males and 65% of females never wore it, although as age increased there was a reduction in the proportion who never wore back seat belts. For all age-groups, men were more likely than women to sometimes or always wear the back seat belt.
Male drivers were much more likely to report having driven after 3 or more glasses of alcohol in the previous 12 months than female drivers (56.1% compared to 17.2%; Fig. 7). This comparison excluded respondents who said that they never drank alcohol (39.0% of females and 10.5% of males in the cohort).
In this study, 8.4% (8354 per 100,000) of the cohort reported that their most serious injury was transport-related, with males and teenagers more likely to be injured and motorcycles the most common vehicle involved. The proportion of injured who were pedestrians increased with age.
As anticipated, the transport injury rates were substantially higher than previously reported for Thailand. Estimates of 152 and 650 per 100,000 population were reported from the Thai Ministry of Public Health (Wibulpolprasert, 2008) and the community-based Thai National Injury Survey, respectively (Sitthi-amorn et al., 2007). The overall patterns of factors surrounding the transport injury, however, are similar with the young and males being the most at risk, motorcycles being the most common vehicle involved and the risk of being injured as a pedestrian increasing with age.
Both previous population surveys used higher injury severity thresholds than our study. The Ministry of Public Health data captured hospitalised injuries (Wibulpolprasert, 2008), while the National Injury Survey captured injuries requiring three or more days off work and/or medical attention (Sitthi-amorn et al., 2007). Further, the National Injury Survey relied on the head of the house reporting for all members, compared with our study which involved self-report. Because our TCS included minor injuries and the respondent reported on their own injuries, we would expect our injury estimates to be higher. Very minor injuries may still have been overlooked as respondents may not recall less serious injuries that occurred during the previous 12 months. In addition, we missed capturing data on transport injuries that were not the most serious injury that the respondent sustained. Thus our estimate is probably conservative.
The external validity of the study in terms of extrapolating the results to the Thai adult population is difficult to judge. Our cohort has slightly more females, less young (<21) and older (>40) adults and is more educated than the Thai adult population. Considering that young males are at higher risk of injury, and our cohort includes proportionally less of these than the Thai population, the estimates may be conservative. However, we also have proportionally less older adults, which may serve to increase the transport injury estimates. The higher education level of the respondents may also lead to a decreased injury rate relative to the general population. It is important to note however, that we are reporting the baseline data for an ongoing cohort study, so the real strength in these data lies in the internal validity for monitoring changes in this cohort over time.
The cohort reported a much higher prevalence of the use of safety devices (motorcycle helmets and seat belts) than previously reported for Thailand in the World Health Organization (WHO) Global Status Report on Road Safety (WHO, 2009). Almost two-thirds of motorcycle users (65%) reported always wearing a helmet compared to 27% in 2005. This rate is lower than for other Southeast Asian middle-income countries with similar legislation, e.g. Malaysia and Indonesia, that reported helmet-wearing rates of 90% or more in 2007 (WHO, 2009). Since helmet wearing was made mandatory in Thailand, however, enforcement has been maintained (WHO, 2009) and may be starting to have the desired cumulative effect. The lower helmet use rate reported by females was notable and could reflect concern with hair-style or hygiene but we did not gather data on motives.
Over two-thirds of respondents (69%) reported always using a front seat belt compared to 56% in 2005. This is similar to Malaysia in 2003, where the seat belt wearing rate was 70%, but less than Indonesia in 2005 when the rate was 85% (WHO, 2009). Only 11% reported always using a back seat belt, however this was still more than previous estimates (3% in 2005;WHO, 2009). There is substantial room for improvement if the Thai population is to reach the wearing rates evident in many of the high income countries (e.g. Australia, New Zealand and Germany in 2006/2007) of 95% or more for front seat belts and 87% or more for back seat belts (WHO, 2009). Considerable benefit would be achieved by mandating seat belt wearing for all vehicle occupants, rather than just front seat occupants as at present.
Excluding respondents that reported never drinking alcohol, more than half of the male drivers (56.1%) reported driving after three or more glasses of alcohol compared to 17.2% of female drivers. The high prevalence of drink driving among male Thai drivers has been reported previously with a cross-sectional study of blood alcohol (BAC) content in 4778 Thai drivers finding 8.7% with a blood alcohol content greater than 50 mg/dl (Chongsuvivatwong et al., 1999). Females were excluded from their analysis, however, because they only made up 2% of the sample. To our knowledge, ours is the first study to highlight the differences between men and women in terms of driving under the influence of alcohol in Thailand. Considering that female respondents in our cohort were much more likely to report never drinking at all (39.0% compared to 10% of males), and those that did drink alcohol were much less likely to drink drive than male drivers, it appears that drink driving countermeasures should be targeted heavily towards male drivers.
Our estimates of the prevalence of transport-related risk behaviours may be biased due to the self-report nature of data collection. Respondents may fail to report behaviours that are risky or socially unacceptable, however, many of our respondents did report drink driving or never wearing seat belts so this does not appear to be a major issue. The lower proportion of young adults aged under 21 and higher education level of our cohort compared to the Thai population may explain the higher estimates for helmet wearing and seat belt wearing than have been found previously. Consequently, the prevalence of the use of safety devices reported in this study probably represents the most optimistic estimates of current safety practices in Thailand. It is, however, unclear what effect this might have on the generalisability of estimates of the prevalence of drink driving. Again, however, we must emphasise the value of our data in terms of the internal validity for monitoring changes in this cohort over time.
We found much higher rates of transport injury than previously estimated by any other source, most likely due to the inclusion of less severe injuries. The findings give an overview for Thailand based on a large national sample of transport using adults and can assist with targeting the groups most at risk of transport injury (males, youth and motorcycle users), and the groups that display the most risky behaviour (drink driving for males, helmet wearing for females and back seat belt use for all age-sex strata). These data are an important addition to the evidence on which Thai authorities can base informed policies for addressing one of the major causes of individual and social burden of health in Thailand. The study goes some way towards redressing the data deficiencies identified by the WHO as one of the major barriers to addressing the problem of transport injury in lower and middle income countries (WHO, 2009). Finally, and perhaps of most value, our cohort study provides an opportunity to monitor changes in these risks and behaviours in the same cohort into the future. Stephan et al. Page 9 Published as: Accid Anal Prev. 2011 May ; 43(3): 1062-1067. Sponsored Document Sponsored Document Sponsored Document Role when injured, by sex and age.
Stephan et al. Page 10 Published as: Accid Anal Prev. 2011 May ; 43(3): 1062-1067.
Sponsored Document Sponsored Document Vehicle used by drivers and passengers when injured, by sex and age.
Stephan et al. Page 11 Published as: Accid Anal Prev. 2011 May ; 43(3): 1062-1067.
Sponsored Document Sponsored Document Reported helmet use for motorcycle drivers and passengers in the cohort, by sex and age (excluding 9685 who reported not riding motorbikes).
Stephan et al. Page 12 Published as: Accid Anal Prev. 2011 May ; 43(3): 1062-1067.
Sponsored Document Sponsored Document Reported front seat belt use for the cohort, by sex and age (excluding 1155 who said the vehicle does not have a front safety belt).
Stephan et al. Page 13 Published as: Accid Anal Prev. 2011 May ; 43(3): 1062-1067.
Sponsored Document Sponsored Document Reported back seat belt use for the cohort, by sex and age (excluding 23,813 who said the vehicle does not have a back safety belt).
Stephan et al. Page 14 Published as: Accid Anal Prev. 2011 May ; 43(3): 1062-1067.
Sponsored Document Sponsored Document Reported drink driving during the previous 12 months for drivers in the cohort, by sex and age (excluding those who reported that they did not normally drive and those who reported never drinking alcohol).
Stephan et al. Page 15 Published as: Accid Anal Prev. 2011 May ; 43(3): 1062-1067.
Published as: Accid Anal Prev. 2011 May ; 43(3): 1062-1067.Sponsored DocumentSponsored Document Sponsored Document
Published as: Accid Anal Prev.2011 May ; 43(3): 1062-1067.
We thank the staff at
Sponsored Document Sponsored Document
This study was supported by the
The authors have no competing interests.
Background: Hypophosphatemia occurs in up to 80% of the patients during continuous renal replacement therapy (CRRT). Phosphate supplementation is time-consuming and the phosphate level might be dangerously low before normophosphatemia is re-established. This study evaluated the possibility to prevent hypophosphatemia during CRRT treatment by using a new commercially available phosphate-containing dialysis fluid. Methods: Forty-two heterogeneous intensive care unit patients, admitted between January 2007 and July 2008, undergoing hemodiafiltration, were treated with a new Gambro dialysis solution with 1.2 mM phosphate (Phoxilium) or with standard medical treatment (Hemosol B0). The patients were divided into three groups: group 1 (n 5 14) receiving standard medical treatment and intravenous phosphate supplementation as required, group 2 (n 5 14) receiving the phosphate solution as dialysate solution and Hemosol B0 as replacement solution and group 3 (n 5 14) receiving the phosphate-containing solution as both dialysate and replacement solutions.
Results: Standard medical treatment resulted in hypophosphatemia in 11 of 14 of the patients (group 1) compared with five of 14 in the patients receiving phosphate solution as the dialysate solution and Hemosol B0 as the replacement solution (group 2). Patients treated with the phosphate-containing dialysis solution (group 3) experienced stable serum phosphate levels throughout the study. Potassium, ionized calcium, magnesium, pH, pCO 2 and bicarbonate remained unchanged throughout the study. Conclusion: The new phosphate-containing replacement and dialysis solution reduces the variability of serum phosphate levels during CRRT and eliminates the incidence of hypophosphatemia.
T HE majority of patients on continuous renal replacement therapy (CRRT) will require phosphate supplementation shortly after CRRT initiation. 1 One reason is that critically ill patients present several conditions predisposing hypophosphatemia such as sepsis, alcohol withdrawal, malnutrition, catecholamines, intravenous glucose infusion, hyperventilation, diuretics and rhabdomyolysis. [2][3][4] Another reason is the CRRT technique that achieves high clearance of small solutes, such as phosphate. [5][6][7][8][9] In addition, low serum phosphate levels may also occur in the setting of extracellular to intracellular shifts that occur with respiratory alkalosis, high blood concentrations of stress hor-mones (i.e., insulin, glucagon, adrenalin, cortisol) and with refeeding syndrome.
As phosphate is a constituent of enzymes and intermediate phosphorylated compounds, it plays a key role in cellular metabolism and is essential in several biological processes. Serum phosphate concentration is maintained between 0.81 and 1.45 mmol/l. By convention, hypophosphatemia is often graded as mild ( o0.81 mmol/l), moderate ( o0.61 mmol/l) and severe (o0.32 mmol/l). Severe hypophosphatemia has been linked to increased mortality in surgical intensive care patients 7 and was recently shown to serve as an independent mortality predictor in sepsis. 10 Symptoms of hypophosphatemia are usually only seen in patients with moderate or severe hypophosphatemia and include ventilatory muscle weakness, cardiac failure, insulin resistance, hemolysis, impaired platelet and white blood cell function, rhabdomyolysis, and, in rare cases, neurologic disorders. 3,[11][12][13][14][15][16][17] However, all these alternations have been shown to reverse by simply correcting the phosphate levels. 3,7,12,13,[18][19][20] Phosphate is supplemented intravenously in symptomatic patients, but phosphate has also been added directly to the dialysate and replacement fluids, 1,21,22 with a risk of precipitation with calcium.
The development of many electrolyte disturbances in the intensive care unit (ICU) could be prevented by the use of better adapted dialysis fluids. This study evaluated the possibility to achieve and maintain a normal phosphate balance over time in patients on CRRT by using a new phosphatecontaining dialysis fluid and replacement fluid.
We used a new phosphate-containing solution for dialysis that in addition to standard electrolytes also contains 4.0 mmol of potassium and 1.2 mmol of phosphate (Phoxilium, Gambro Lundia AB, Lund, Sweden, Table 1). As a control, we used our routine dialysis solution that does not contain phosphate (Hemosol B0, Gambro Lundia AB, Table 1). At our ICU at Lund University Hospital we applied three regimes, half a year each, for all patients requiring CRRT treatment. The treatment mode used was CVVHDF. During the first period (group 1), all the patients received dialysate solution and replacement solution that did not contain phosphate (Hemosol B0), during the next half year (group 2), all patients requiring CRRT treatment received the phosphatecontaining solution as dialysis solution and a phosphate-free replacement solution (Hemosol B0) and finally during the last half year period (group 3), the patients received the phosphate-containing solution both as a dialysis solution and as a replacement solution. Blood sampling was performed according to normal routines at our department, but the physicians in charge continuously modified the CRRT treatment settings and the phosphate supplementation according to the patients' ongoing clinical needs. The physicians treating the patients did not have knowledge of the study setting during the treatments, but they were fully aware of the fluids and their contents.
After acceptance by the Regional Ethical Review Board (DNR 570/2008), Lund University, Sweden, we evaluated retrospectively the first 14 patients who did not fulfill the exclusion criteria in each group, a total of 42 consecutive patients (Table 2). Patients were excluded if they had chronic kidney disease, if they had received intermittent dialysis before the ICU stay, if the CRRT treatment lasted o10 h or if they were under the age of 18 years. In all patients, the Gambro Prismaflex CRRT machine with the CVVHDF modality and a Hospal M100 filter (Hospal Industrie, Meyziew, France) was used. The blood flow, the dialysis fluid flow, the replacement fluid flow, the anticoagulation used and the fluid removal were set according to the patients' conditions and requirements. Of the replacement fluid, 500 ml/h was postfilter and the rest was prefilter in each treatment according to the general standard at the department. Intravenous phosphate addition was prescribed when serum phosphate was o0.8 mmol/l, also according to the general standard at the department. Nutrition was given only if the patients were hemodynamically stable as parenteral or enteral or both during the study period.
All patients had regular measurements of serum sodium, potassium, ionized calcium, pH, pCO 2 and bicarbonate either from an arterial line or from a central venous line before the start and regularly every fourth hour all through the CRRT treatment during 1, 2, 3, 4 and maximum 5 days. For phosphate and magnesium analyses, blood samples were taken at 5:00 hours in the morning and at 17:00 hours in the afternoon before the start and during the CRRT treatment days. Blood samples were analyzed at the hospital laboratory (Labora-tory for Clinical Chemistry, Lund University Hospital, Lund, Sweden). The results are presented as means AE SD for normal distributed data, median and range for remaining data.
In addition, baseline characteristics of study patients and delivered CRRT were registered (Table 3). Hypophosphatemia has been defined as condition where the serum phosphate level is o0.81 mmol/l.
One-way repeated measures analysis of variance was used. The statistical program used was Sigma-Stat, version 3.5, for Windows XP. Differences were considered to be significant for Po0.05.
Main diagnoses leading to intensive care of the study groups, baseline characteristics, timing of initiation of CRRT, ultrafiltration rate, duration of the CRRT and anticoagulation were statistically similar between the three groups (Tables 2 and 3).
The incidence of hypophosphatemia occurred in 11 of 14 patients in group 1, where the patients did not receive the phosphate-containing dialysis solution, but received phosphate intravenously based on the serum phosphate values (Fig. 1). In group 2, where the patients received a phosphate-containing solution as the dialysis solution and a phos- zThe dialysis dose in CRRT is expressed in terms of ml of effluent (dialysate1ultrafiltrate) per kg of body weight (BW) per hour (ml/kg/ h). 23 §Statistical differences were as follows: between groups 1 and 2 (P 5 0.046), between groups 1 and 3 (P 0.001), and between groups 2 and 3 (P 5 0.003).
phate-free solution as the replacement solution, five of 14 patients had at least one episode of hypophosphatemia (Fig. 1). No episodes of hypophosphatemia were detected in group 3. The serum phosphate level was 1.90 mmol/l at baseline and 0.99 mmol/l during CRRT treatment in group 1. In group 2, the corresponding values were 1.54 and 1.20 mmol/l, and in group 3, they were 1.83 and 1.43 mmol/l (Table 5). However, due to the simultaneous intake of enteral/parenteral solutions, two of 14 of the patients in group 3 had a temporary increase in serum phosphate (o1.9 mmol/l), but there were no cases of hyperphosphatemia that required a withdrawal of the phosphate-containing dialysis solution from the treatment.
Phosphate intake was from nutrition and from intravenous supplementation. The average phosphate input calculated from the total nutrition was 18 mmol/CRRT treatment/day for all groups. Short-acting insulin was administered as infusion in order to achieve a blood glucose level between 5 and 8 mmol/l. Phosphate was supplemented intravenously if serum phosphate declined to o0.80 mmol/l. In group 1, the average phosphate supplementation was 10 mmol/CRRT treatment/ day. In group 2, the average supplementation was 5 mmol/CRRT treatment and day, and in group 3, there was no intravenous supplementation.
There was a decline in phosphate levels during CRRT treatment in both groups 1 and 2. In group 1, serum phosphate declined constantly, although the patients received phosphate supplementation in-travenously (Fig. 2). In group 2, the phosphate levels were unstable and reached a low level in the end of the study. In group 3, where the patients received a phosphate-containing solution both as the dialysis solution and the replacement solution, none of the patients had episodes of hypophosphatemia (Fig. 2). Phosphate remained stable in this patient population during the entire study.
For group 3, we also evaluated any adverse event in sodium, potassium and magnesium homeostasis. There was a significant increase in sodium (P 0.001) and a slight increase in ionized calcium (P 5 0.029), whereas potassium and magnesium remained stable during the entire study as summarized in Table 4. A comparison between the study groups revealed that there was a significant difference in phosphate (P 0.001), ionized calcium (P 5 0.004) and bicarbonate (P 5 0.045) between the groups (Table 5). The pH and pCO 2 measurements were obtained at baseline and rapidly declined toward normal values after starting the CRRT treatment and remained essentially unchanged during the treatment in all groups (Table 4).
This study demonstrates that the new phosphatecontaining dialysis solution is safe, reduces the variability of serum phosphate levels during
Various protocols for intravenous phosphate supplementation have been studied in the last 30 years. Today, there is a trend toward the use of larger and faster boluses of phosphate because of high failure of repletion (20-70%) and the need for additional phosphate administration. [24][25][26][27][28][29][30][31] Authors usually agree that larger amounts of phosphate are needed to correct total body deficit, but their fear of adverse reactions has prompted a restrained attitude. Recent studies on ICU patients confirm that a relatively rapid infusion of potassium phosphate is safe if baseline serum potassium is below 4.5 mmol/l. 1 The infusion rate is thus consequently limited by the serum potassium levels, and also by serum calcium levels, as high phosphate serum levels could induce hypocalcemia, as phosphate could precipitate with calcium in blood vessels and tissues. Recently, phosphate has been injected into dialysis solutions during treatment, but there is a risk of precipitation. 1,21,22 We evaluated whether phosphate-containing solutions for dialysis and replacement simplify phosphate replacement even further.
The frequency of phosphate disturbances in critically ill patients is high, although the figure varies considerably depending on the study, between 8.8% and 80%. 32,33 In our study, the incidence of hypophosphatemia during CRRT was 79% in the control group. The incidence of hypophosphatemia was lower (35%) in patients receiving the phosphate-containing solution as the dialysis solution and Hemosol B0 as the replacement solution. There were no episodes of hypophosphatemia in patients receiving only the phosphate-containing solution, but there was an incidence of mild (>1.9 mmol/l) hyperphosphatemia in 14% of these patients. This is probably due to the simultaneous intake of nutritional support, but the phosphate levels were only marginally elevated without physiological consequences and did not require withdrawal of the phosphate-containing dialysis solution from the patient.
The amount of phosphate required to correct total body deficit is variable and depends on the cause of hypophosphatemia and the chronicity of the process. 3,11,24 The many physiological rearrangements in ICU patients may explain the need for larger amounts required for repletion. In our study, phosphate addition was necessary for patients in groups 1 and 2. This is representative of our experience, where the majority of patients on CRRT require phosphate supplementation shortly after CRRT initiation despite nutritional support. Malnourished alcoholic patients have a larger deficit due to long-standing negative phosphate balance. Malnutrition induces a re-feeding syndrome that has been reported after only 48 h of fasting in the ICU. 2 The distribution volume for phosphate might be increased, 24 while insulin, carbohydrate, and catecholamine administration act to decrease the serum phosphate concentration. 2,3,11,34 The incidence of hypophosphatemia in this study was 10% of the 42 patients before the start of CRRT. The incidence of hyperphosphatemia was Mean values of each day of all the patients. *P 0.001 between the groups during the CRRT treatment. wP 5 0.004 between the groups during the CRRT treatment. zDifferences between the groups were not statistically significant. §P 5 0.045 between the groups during the CRRT treatment.
52%, which is slightly lower than that found in other studies, where 65-80% of patients present hyperphosphatemia before the start of the CRRT. 8,35 Phosphate was within normal limits in 38% of patients even before the beginning of CRRT. The new dialysis fluid contains 30 mM bicarbonate compared with 32 mM in Hemosol B0. With prescribed CRRT clearances of ! 20 ml/ kg/h, most acid-base disturbances can be managed with bicarbonate compositions of 25-35 mM. 36 The normal serum range for bicarbonate is 22-30 mmol/l, which was achieved in patients in all groups, although the increase in group 1 was most significant. Ionized calcium increased in all groups during treatment time, but the increase was less marked in groups 2 and 3. There is a possibility that the slow increase in bicarbonate and ionized calcium could reflect the differences in the ion content between Hemosol B0 and the phosphate-containing dialysis fluid, although these differences could also be due to the limitations of this study. The fact that this was a retrospective observational study, without other intervention but for the dialysis solutions, means that bias due to the patients' illnesses and due to other treatments could affect ion concentrations. Even the phosphate concentration in parenteral and enteral nutrition cannot be completely excluded, although this would not alter the basic results of the study.
The next objective of the study was to evaluate whether the administered phosphate could alter calcium and potassium homeostasis. Phosphate could theoretically precipitate with calcium, which could lead to hypocalcemia in the patients. We found that the phosphate-containing dialysis fluid did not induce hypocalcemia during CRRT. Another important concern was whether this fluid would cause potassium overload, as potassium phosphate is used instead of sodium phosphate. We found that the phosphate-containing dialysis fluid did not induce hyperkalemia either. Potassium phosphate was favored over sodium phosphate because potassium usually has to be added to CRRT solutions. In addition, there were no disturbances of magnesium levels.
The present study shows that by using the phosphate-containing fluid both as the dialysis fluid and the replacement fluid, we could eliminate the episodes of hypophosphatemia. The new phosphatecontaining dialysis fluid simplified the phosphate control and avoided rapid phosphate fluctuations with intravenous bolus administration.
*Differences between groups were not statistically significant. Numbers in parentheses are range, unless stated otherwise. wWe scored RIFLE-risk (R, risk; I, injury; F, failure; L, loss of kidney function; E, end-stage kidney disease) as 1, injury as 2, and failure as 3.
Competing interests: The authors have not disclosed any potential competing interests. The ICU department of
The impact of science on ethics forms since long the subject of intense debate. Although there is a growing consensus that science can describe morality and explain its evolutionary origins, there is less consensus about the ability of science to provide input to the normative domain of ethics. Whereas defenders of a scientific normative ethics appeal to naturalism, its critics either see the naturalistic fallacy committed or argue that the relevance of science to normative ethics remains undemonstrated. In this paper, we argue that current scientific normative ethicists commit no fallacy, that criticisms of scientific ethics contradict each other, and that scientific insights are relevant to normative inquiries by informing ethics about the options open to the ethical debate. Moreover, when conceiving normative ethics as being a nonfoundational ethics, science can be used to evaluate every possible norm. This stands in contrast to foundational ethics in which some norms remain beyond scientific inquiry. Finally, we state that a difference in conception of normative ethics underlies the disagreement between proponents and opponents of a scientific ethics. Our argument is based on and preceded by a reconsideration of the notions naturalistic fallacy and foundational ethics. This argument differs from previous work in scientific ethics: whereas before the philosophical project of naturalizing the normative has been stressed, here we focus on concrete consequences of biological findings for normative decisions or on the day-to-day normative relevance of these scientific insights.
How do we know right from wrong? Do we dig deep into our intuitions? Should science offer a full picture of human virtue? These questions remain as yet unresolved, and continue to offer ample room for academic debate. Especially the importance of science for ethics proves to be substantially discussed (e.g., Kurtz 2007;Pigliucci 2003). Many authors agree that science can describe morality and that science can go a long way in explaining morality's origins (e.g., Joyce 2006). But there is much disagreement about the relevance of science for the normative domain of ethics.
Normative ethics concerns questions about right and wrong and the criteria to distinguish them. It is not about how the world is, but about how it should be. More accurately, normative theories attempt to delineate what is correct use of actionguiding or prescriptive terms as ought, value, good, should, duty, obligation, right, wrong, permissible or forbidden. This makes normative inquiry different from scientific inquiry. Regarding the latter, what scientists find out about the world is not qualified in terms of 'right' or 'wrong'. Science is deemed devoid of normativity; instead, it is a purely descriptive and explanatory endeavor. As such descriptive ethics is a part of science and does not in itself prescribe: it merely describes how people use normative ethical terms. This notwithstanding, some ethicists defend that scientific findings can be a guide in determining how one should live (e.g., Rottschaefer 2007). Such scientific normative ethics is the topic of this paper. We will hereafter shorten it to scientific ethics since we will only be concerned with the normative domain of ethics.
Scientific ethics generally meets two kinds of criticism. First, the idea that normative statements can be deduced from scientific statements is accused of committing the naturalistic fallacy (e.g., Farber 1994;Woolcock 1999; see Sect. 2.2). Second, when not committing this fallacy, it is claimed that scientific ethics fails in demonstrating the relevance of science for normativity because science cannot offer a foundation for ethics (e.g., Farber 1994;Woolcock 1999;Rosenberg 2000; see Sect. 2.1). While the first criticism is often debated, the second criticism is not systematically discussed in the literature. Still, it is not unusual for critics of scientific ethics to endorse both statements as valid criticisms.
The first aim of this paper is to defend scientific ethics against these two major criticisms. Initially, we show that most contemporary scientific ethicists do not commit the naturalistic fallacy, contrary to what their critics claim. To support this thesis, in Sect. 3 it is illustrated that science can be relevant to ethics without committing the naturalistic fallacy, while in Sect. 4 more general arguments are formed. Additionally, the critics' critique is analyzed and found to be contradictory: the same criticists who refer to the naturalistic fallacy complain that science does not offer a foundation for normative ethics. We refer to this contradiction in Sect. 3.2. To substantiate these arguments, we first revisit George Edward Moore's notion of the naturalistic fallacy; we also explain what a foundation is and how the reasoning behind the naturalistic fallacy is in fact an argument against foundational ethics (see Sect. 2).
The second aim of this paper is to counter further criticisms by explaining and defending the reasoning behind scientific ethics. In Sect. 4.1, it becomes clear that a difference in conception of normative ethics underlies the disagreement between proponents and opponents of scientific ethics. Indeed, the discussed criticisms of scientific ethics all rely on a foundational view of normative ethics, while scientific ethicists-by referring to methodological naturalism-see normative ethics as nonfoundational. Scientific ethicists refer to methodological naturalism as the proper method for scientific inquiry. In Sect. 4.2, we argue that methodological naturalism can be used also for normative inquiry. Since methodological naturalism is a nonfoundational method, we hereby defend nonfoundational ethics in general. In this section, we further explain how science informs normative ethics in a nonfoundational as opposed to a foundational system. We conclude by stating that (1) scientific ethics is best conceived of as an instance of nonfoundational normative ethics; that (2) when scientific ethics is nonfoundational, science is relevant to normative inquiry without committing any fallacy; and (3) that nonfoundational scientific ethics can be preferred over foundational ethics because the former is more successful. 1The presented arguments are different from previous work in scientific ethics. Our understanding of the naturalistic fallacy is in line with diCarlo and Teehan (2007). However, where their paper generally and abstractly concludes that science informs ethics, we take these conclusions further by discussing how scientific insights are relevant to normative inquiry. Therefore we revisit the methodological underpinnings of scientific ethics by discussing methodological naturalism in ethics. Finally, while recent naturalistic accounts attempt to formulate an appropriate moral theory that translates all normative concepts in empirically testable concepts (Casebeer 2003), we focus on the other direction in which scientific findings and methods evaluate normative aims and methods.
But now, let us recapitulate the naturalistic fallacy, explain what is meant with foundations in normative ethics and argue how both themes are related to each other.
The words foundation, grounding and their derivatives are differentially used in the ethical literature. In this paper, foundational normative ethics, shortly foundational ethics, refers to any attempt at deriving a true normative system out of one or several first norms. Grounding ethics, then, refers to the act of finding such first norms. Let us look into these concepts.
In what follows, we will make a distinction between normative and descriptive statements. We use descriptive statement or description very broadly, namely to denote all statements concerning the nature of things in the realm of science, religion or metaphysics. Normative statements or moral norms is used to denote action-guiding statements or prescriptions; i.e., statements that can be in the form of 'X is good, valuable, right' or 'we should do X', 'we ought to do X', etc. If X is an action or state of the world, we will talk about substantive norms. If X describes the form of a judgment (e.g., as in the statement 'judgments that are universalizable are right'), we will denote them as formal norms. If X is a procedure for finding normative statements, as Habermas' discourse principle is, we will denote it as a procedural norm. Throughout this paper we are concerned with the quest for moral norms that are descriptively determined, i.e., we will be concerned with grounds for normative ethics. What does this mean?
Some philosophers have attempted to find one or a very limited amount of moral norms that are grounded in a non-normative theory, mostly a descriptive theory. This means that the descriptive theory in itself, without the help of any purely normative statement, determines at least one moral norm. That moral norm can be refuted on the basis of new descriptive information but it cannot be refuted on the basis of other moral norms. We will denote such premised determined moral norms as first moral norms. The quest for such first moral norms accordingly will be called grounding ethics. 2 Grounding ethics results in a foundational ethics. We will now give illustrations to further clarify these concepts.
Natural law theories in ethics can serve as examples of foundational ethics. According to Feser (2010), natural law theories evolving from the classical tradition (e.g., Thomas Aquinas) ground moral rules in nature by making no strict distinction between descriptive and normative statements: moral norms, including their moral force, are part of nature and can be described as such. For classical natural law theorists, a description of nature also determines general moral norms from which specific rules can be inferred. According to Thomas Aquinas' natural law theory, for instance, the precepts of moral law theory are given by God and are to be found in nature. They are universally binding and universally knowable (Murphy 2008). The content of Aquinas' moral theory is that good should be done and evil avoided. This is an abstract 'first moral norm' and it is conceivable that many agree with it. The content of the moral norm however is not important in deciding if it is a foundational ethics or not: for this we must ask how the moral norms in the system relate to each other. In this case, all moral norms are derived from this moral norm. Moreover, the norm is determined by nature and cannot be refuted by moral norms that follow from it. This, then, is a clear instance of foundational ethics.
Another example is a new natural law theory as developed by Walsh (2008). In his theory, friendship, offspring and life are first identified as ends in themselves, as basic human goods. These first values cannot be questioned within the moral system that follows from them. Also, they are the touchstone against which all acts must be evaluated. Acts can be chosen because of the act itself, or because of its consequences. Either way, if the choice to perform an act entails the choice of an appropriate human good, then this act is morally good; if not, it is morally bad. As such, Walsh (2008) argues, sexual acts are only morally good if the choice to perform them entails the choice of a basic human good. According to Walsh' new natural law theory, sex in itself is not a basic human good, but procreation is. Hence, a sexual act must entail the choice to procreate. Following this reasoning, Walsh considers homosexual sex to be morally wrong because the choice for homosexual sex cannot entail the choice to procreate. This shows that specific basic human goods here function as independently derived foundations of a moral system. Walsh' religiously inspired new natural law theory is also a foundational normative ethic. Contrary to the former example though, the content of its first moral norms is much more concrete and more likely to be controversial. However, in deciding if the system is foundational or not, one has to consider the procedure for finding substantive moral norms and not the content of the resulting substantive moral norms.
Other examples of foundational ethics pertain to the work of certain nineteenth century intellectuals who developed a normative system, attempting to ground ethics in biological evolution. Herbert Spencer's (1820-1903) evolutionary ethics is a case in point. He reasoned that evolution by natural selection results in adaptations that are morally superior. Whatever is further evolved by natural selection is therefore better. This implies that everything following from this first principle must be true, and that one should promote evolution by natural selection (Moore 1993(Moore / 1903)). Whether one agrees with the content of this moral norm or not, the basic idea is again that it is a first moral norm. Precisely because it was a foundational system, philosophers instantly refuted Spencer's ethics. George Edward Moore (1873Moore ( -1958) ) dedicated substantial parts of his Principia Ethica to Spencer's evolutionary ethics (Moore 1993(Moore /1903, §33), §33). According to Moore, Spencer committed a crucial fallacy, which he coined the naturalistic fallacy. This fallacy is often invoked to argue that one cannot ground ethics in nature. 3 But a close reading of the Principia Ethica reveals that Moore in fact argued that one cannot 'ground' ethics at all, hence one cannot ground it in anything.
In the next section, we discuss Moore's reasoning that leads to the naturalistic fallacy argument. It is important to know that we do not purport to discuss the validity of this reasoning. We aim to make its reasoning clear in order to ask if scientific ethicists are indeed committing the naturalistic fallacy, as its critics suggest, and in order to evaluate the coherence of critics' arguments in Sect. 3.2. Though diCarlo and Teehan (2007) put forward a similar argument, here we specifically stress that the naturalistic fallacy relates to 'grounding' ethics. Since this is crucial to evaluate the criticisms of scientific ethics, we will highlight the relevant parts in Moore's reasoning.
In his explication of the naturalistic fallacy, Moore built on the insights of Henry Sidgwick . Sidgwick, a British utilitarian moral philosopher in turn was influenced by David Hume's (1711-1776) work. Hume noticed that the author of every moral system seems to make prescriptive or normative conclusions from descriptive statements (Hume 1739(Hume -1740)). Since by now many interpretations of Hume's and Moore's reasoning exist (Curry 2006), it is helpful to consider both their arguments in more detail.
Take the following reasoning (cf. Ferguson 2001): Premise 1: Humans are evolutionary disposed to act altruistically. Conclusion: It is good to act altruistically.
According to Hume, this is a wrong kind of reasoning because the conclusion does not logically follow from the premise: there is a difference in meaning between 'we are evolutionary disposed to' and 'it is good to'. This difference in meaning between a descriptive statement and a prescriptive statement is known as Hume's is/ought gap. Accepting this gap has direct consequences for any 'scientific ethics'. If scientists find that something is the case, it does not logically follow that the descriptive statement, or parts of it, ought to be the case. There is no such simple logical connection between scientific statements and ethical statements. According to Hume, ''a reason should be given'' (Hume 1739(Hume -1740) ) for why a moral statement follows from descriptive statements. This can be done by adding a second premise, as is done below: Premise 1: Humans are evolutionary disposed to act altruistically. Premise 2: It is good to do everything humans are disposed to by their evolution. Conclusion: It is good to act altruistically.
Here, the conclusion does follow logically from the premises. However, it comes at the cost of premise 2 being a prescription instead of a description. As a result, one has not derived a moral principle from descriptive statements only. In other words, it is not demonstrated that one can go from an 'is' to an 'ought'. As Hume's reasoning is applicable to all descriptive theories and all moral statements, the is/ought gap precludes the possibility of demonstratively deriving first moral principles from descriptive statements.
Moore's reasoning is somewhat different, but has similar implications. Also Moore deemed it impossible to find a demonstratively true first moral principle that cannot be doubted from within the moral realm. Supported by the arguments that led him to the formulation of the naturalistic fallacy, Moore rejected the possibility of a first moral principle. More correctly, he rejected a certain class of first principles, namely those that are considered to be analytically true. Before clarifying this, let us first revisit Moore's reasoning in the Principia Ethica.
Ethics-in Moore's terminology-is about moral truth, not about practice (Moore 1993(Moore /1903, §3- §5, §14), §3- §5, §14). It is about finding a first statement upon which Ethics-including the discussion of our everyday normative judgments (ibidem, §1)-can be built. This first statement provides an answer to Ethics' first question, i.e., ''What is good?'' (ibidem, §2). Moore adds: ''Unless this first question be fully understood, and its true answer clearly recognized, the rest of Ethics is as good as useless from the point of view of systematic knowledge'' (ibidem, §5). In other words, to save Ethics, one must find a first moral statement-such as the second premise in the example above. This principle must define what is good and it must be true by definition. This means that it must be analytically true (cf. infra).
So far so good, were it not that Moore insisted that finding a first moral principle that truly defines what is good is impossible. This has to do with the fact that he has an analytic definition of the word 'good' in mind (ibidem, §6). In general, a true analytic definition describes the real nature of a notion denoted by the word; it enumerates the simple notions that are already in the meaning of the complex notion (ibidem, §7). Analytic statements hence only explicate what is already in the meaning of the subject. The meaning of 'good' then describes its true nature. How does one find this meaning according to Moore? One does not need any observation to establish the real nature of a notion. Every normal user of a certain language, when thinking clearly, instantly grasps when an analytic statement is true. Hence one can derive the true meaning of 'good' by clear thinking alone. Now 'good' is indefinable, says Moore: it is already a simple notion, meaning that there is nothing in the meaning of 'good' than 'good' itself. Those who define 'good' as something else and claim this definition to be true all commit the naturalistic fallacy (ibidem, §1- §15). Moreover, Moore continues, we intuitively acknowledge that we cannot define 'good' in that for any definition of 'good' as something else we can meaningfully ask whether this 'something else' is indeed 'good'. This means that we never instantly see such a statement to be true, thus it can never be analytically true. This argument is since known as the 'open question argument' (ibidem, §13).
Moore's idea that all of Ethics should be built upon an analytic truth, logically implies that nothing that follows from this truth can refute this first definition-otherwise it would not be an analytic truth. Hence Moore was looking for a 'first norm'. The core idea of Moore's reasoning is thus that one cannot 'ground' a first moral principle: not in nature, not in metaphysics, and not in ethics itself. Only analysis of the meaning of a moral concept like 'good' would provide a solution, but this is impossible. According to Moore, 'naturalists'-up to his time-made this very mistake. They tried to identify 'good' with something else. Contrary to what the term 'naturalistic fallacy' seems to imply, Moore's argument hence also applies to metaphysical properties (ibidem, §66- §85). Similarly, religiously grounded normative systems are equally debunked if they rely purely on conceptual analysis for their foundations (cf. diCarlo and Teehan 2007).
In this interpretation, both Hume's 'is/ought' gap and Moore's naturalistic fallacy preclude the possibility of foundational ethics, and the derivation of a first normative principle from descriptive theories. Because the subtle differences between these two fallacies are less important for our argument, we will use them interchangeably in the remainder of this paper.
Let us now illustrate that science can be relevant for ethics without committing the naturalistic fallacy and explain why some critics of naturalistic ethics contradict themselves.
Though Moore denounced all 'naturalist' moral systems, there were numerous early approaches in evolutionary ethics that did not commit the naturalistic fallacy (e.g., by T.H. Huxley and G.G. Simpson). Also from the last decades of the twentieth century onwards, several accounts proliferate in defence of a closer and argumentatively sound interplay between science and normative ethics (e.g., Binmore 2005;Ruse 2008). What typifies these approaches is the argument that science is relevant for ethics, without their being an attempt to start from a first moral principle. Neither is there the attempt to derive such a principle.
Proposals in which scientific findings are claimed to play an important role for normativity vary from being uncontroversial and allegedly 'trivial' to supposedly reductionist accounts. Most authors stress the philosophical question of how moral and empirical concepts are connected (or unconnected); rarely do they make their proposals concrete, e.g., by exemplifying how science informs ethics in everyday issues. A refreshing exception, though in the field of ethics broadly conceived, can be found in Pigliucci (2003).
Our aim here is to discuss how scientific findings have an impact on normative ethics and ethical practice, even if they do not yield demonstratively true ethical principles. Scientific ethics' deviates indeed from Moore's 'Ethics', in being preoccupied less with absolute truth and more with practice. This aligns with current conceptions on ethics as an orienting tool to reflect on individual and societal practices (e.g., Kurtz and Koepsell 2007). In the third section, we look closer at the philosophical assumptions underpinning this view of ethics. For now, it suffices to point out that scientific information is conditionally relevant for ethics. That is, if we accept certain moral principles, then everything known can be used to infer rules that help us to reach these moral ends. In this scenario, scientific knowledge is instrumental to ethics (Rosenberg 2000), or science can help us to infer hypothetical imperatives only (Binmore 2005). This is not controversial, and both foundational and nonfoundational systems can accept this procedure. Hence science is important for ethics in general. However, scientific ethics relies merely on this conditional procedure, while foundational ethics further relies on the inference of first moral norms. Here we demonstrate that its conditional procedure does not commit the naturalistic fallacy: first we illustrate how science informs ethics; then we explicate the line of reasoning.
A clarifying example is provided by the Kibbutzim in Israel, modern communities that are unique in their organization of production, ownership, consumption and child care (Agassi 1989). From the start these communities aimed to create a society where all would be equal and free from exploitation. Property was common. Every member received an equal wage depending on his or her needs. Men and women were expected to participate equally in all kinds of work: household chores, childcare, politics, farming and so on. Trained nurses and teachers raised children away from their parents. It was hoped that this would liberate women from their traditional mother roles. However, after one generation already this organizational structure weakened. Women were found to be more active in teaching and child care, while men participated more in politics and field work. Men also took up the majority of leading and managing positions. Because of these 'role patterns', men had easier access to some assets such as a car, an office and an apartment in town.
Some commentaries (e.g., Agassi 1988) remained convinced that these gender differences could and should be eradicated. To do so, it would be helpful-or even necessary-to identify the precise factors causing the gender differences. Other commentaries (e.g., Palgi et al. 1983) saw in the unique constellation of the Israeli Kibbutzim a test case for social theories explaining gender inequality as a consequence of the unequal social organization of production, ownership and so on. Since gender differences were not eradicated in the Kibbutzim, where social organization started out equal for men and women, these theories are not supported. Maybe then one can consider biology as a factor accounting for at least some gender differences?
Let us zoom in on explanations of childcare asymmetries (yet without claiming these explanations to apply to other aspects of role patterns-indeed, therefore more scientific information would be needed).
Concerning child care asymmetries, in all cultures mothers spend more time with their children than fathers do (Lamb 2003;Owen Blakemore et al. 2008). This can be modified partly by the social environment. For example, pregnant women who had more prior childcare experience (for example due to baby-sitting) feel more positive about caretaking, children and their own fetus (Fleming et al. 1997); and women may be asked to baby-sit more than men. But biology also plays a role in 'moulding' mothers into this role. Pregnancy hormones seem to influence nurturing behaviour: a pregnant woman's body experiences a change in the estrogen/ progesterone ratio. The change in this ratio during pregnancy correlates with maternal behaviour immediately after birth (Fleming et al. 1997). Lactation as well may influence mothering behaviour due to lactation-induced hormonal changes. As tested in nonhuman primates, breastfeeding heightens the concentration of blood hormones like oxytocin, which has a motivating role in nursing and grooming behaviour (Maestripieri et al. 2009). In addition, women have a lower threshold for responding to babies than most men (Silk 2002) and feel more protective towards infants (Alley 1983). More recently, it was found that women are more interested than men in babies and caretaking (Maestripieri and Pelka 2002) and that women feel somewhat more motivated than men to take care for babies when these have (manipulated) very baby-like faces (Glocker et al. 2009). It is suggested that these biological factors induce nursing behaviour in females (Hrdy 2005) and make it satisfying for mothers to nurture their children. However, this does not mean that men cannot be induced to demonstrate caretaking behaviour. That the social environment can induce paternal care is for instance suggested by the finding that men engage in more paternal care when couple intimacy is high (Belsky et al. 1991). Also biology helps in inducing paternal care: expectant mothers and fathers both experience an increase in prolactin levels and, in humans, higher prolactin levels in men are correlated with more paternal behaviour (Storey et al. 2000;Fleming et al. 2002). Experienced fathers are more reactive towards cries of babies than first-time or less experienced fathers: they show a more enhanced prolactin response and they feel a greater need to respond to the infant's cries (Fleming et al. 2002).
In other words, while men can be induced to be more responsive to children, it is plausible that many mothers-not necessarily women in general, maybe only those who have been pregnant or are lactating-will still want to spend more time with their children compared to fathers. If these differences in desires are-even partlycaused by hormonal changes during pregnancy and lactation, then we may expect these differences in desires to exist over a vast range of social environments. Along this line of thought, one can expect that completely eradicating the resulting 'role patterns' would demand that many men and women constantly act against their internal desires. This could be very hard to do, and even could be dissatisfying. Of course, it is exactly the point of moral behavior to act against certain tendencies for moral reasons. 4 However, enforcing the total eradication of all gender differences not only conflicts with strong spontaneous tendencies, it can therefore also conflict with specific values humans have. Since people differ in their basic outlook of life, we value freedom of choice and life satisfaction; in general women also value familial intimacy more than men do. We also consider these values as moral reasons for acting. As a consequence, a more coherent solution could allow for role patterns to exist without forcing people into a certain role and without disvaluing one or the other role in e.g., economic terms. This implies that one takes into account the inherent desires people have; 5 men who prefer child care over politics may as well fulfil this role; women who prefer politics over child care may pursue their ambitions. But if a substantial amount of mothers spontaneously want to specialize in child care and service work, their choice can be allowed as well.
Then the question becomes how to accommodate the possibility that several women want to have both employment and children. Indeed, studies show that across Europe, the US and Japan, a relative majority of women prefers combining employment and family work above either a work-centred life (focused on a career and where family-life is fitted around their paid work) or a home-centred life (giving priority on private life and family over paid work). Significantly, men tend to prefer a work-centred life more than women do (Hakim 2008). This makes one expect that several women wanting to combine employment or a career with having children cannot easily rely on the willingness of their partner to contribute equally in the household. 4 We thank an anonymous reviewer for this remark. 5 One can remark that taking into account the inherent desires people have enforces us to equally consider the inherent desires of pedophiles, psychopaths, sexists, and so on. However, 'taking into account' is not the same as legitimating. It is better to know about these desires and their origins if one wants to eradicate malicious behavior. Second, those desires would be unproblematic if they did not conflict with the desires of other people. It is exactly because they do conflict with the desires of other people that we do not agree with these activities. Here science is of great help in pointing out what harm it does to small children if they are manipulated into sexual activities, what harm it does to people if they are denied certain positions due to their sex and so on. Third, consistently with the rest of our account, scientific agreement alone cannot solve the discussion: we need to find an agreement on some values to have a basis for discussion.
Here science provides us unforeseen options. For instance, in modern societies grandparents often invest heavily in their grandchildren (e.g., Pollet 2007). In extant hunter-gatherer societies as well, children clearly benefit from the help of others than their parents, especially of maternal grandmothers (Sear and Mace 2008). It is suggested that during long periods of our evolution, children's survival depended on the additional care they received from others than their mothers (Hrdy 2005). On the basis of this knowledge, one can consider promoting institutionalized childcare or familial assistance, benefiting those mothers who pursue demanding occupations. Moreover, fathers can be induced to feel more attentive towards infants as well. We can use this and similar information to optimally promote paternal care, although realizing that since differences in desires remain, an equal role pattern will be very hard to achieve. In sum, to promote women's professional aspirations, a narrow focus on paternal care will not help as much in reaching this aim as other possibilities would. A more optimal and desired solution is to keep the possibilities open by promoting or facilitating familial care, institutionalized childcare and paternal care.
What this account illustrates is that scientific knowledge about children's needs and our evolved nature incites us to consider more successful alternatives to the enforced paternal care one tried to implement in the original Kibbutzim. Fathers should have the possibility to go on paternity leave, but science teaches us that this possibility alone will not be enough to free ambitious mothers from their mother roles. Promoting childcare facilities and familial assistance may be a more fruitful option.
Scientific findings play a double role in this example. First, they make us realize that people hold unforeseen values. Scientific findings make us take seriously the fact that women in general value childcare more than men in general do because. according to the scientific information we have, this difference is unlikely to be eradicated by upbringing. Also, familial solidarity appeared a possible and partial solution for childcare regulations. If we care about freedom of life choice and more economic equality, then science informs us that we could promote familial childcare systems. Hence, science is conditionally relevant for normative conclusions. Second, science guides away from certain value systems when, as in the example, its values cannot be realized because for instance they conflict too much. Total equality conflicts with the fact that men and women generally value different things and want to make other life choices. Hence, if we accept that we want a practically coherent normative system, then we have to downgrade the importance of either total equality or of freedom of choice. If we want a coherent system that takes deeply ingrained desires into account, then we should not aim for total equality. Again, science is conditionally relevant for our normative conclusions. Now, when science guides us away from value sets or imports new moral options, do we then commit the naturalistic fallacy? In both cases, one can ask if we are not deriving a first moral norm from a pure description of the world. Let us consider the case where science guides us away from a normative system based on total equality. The structure of the reasoning was as follows:
Moral premises: Freedom of life choice, equality and practical feasibility are all morally good.
Factual premises: In general and over a broad range of situations (varying in upbringing, culture, etc.) women value childcare more then men do.
Conclusion: Sexual differences in time spent in caring for children ought not to be totally eradicated.
Clearly, the conclusion is not derived independently from normative rules. It is therefore not a first moral norm. But one might ask where the moral premises come from. Is any of them a first moral norm? Some of these norms (e.g., freedom of life choice) came into play because scientific findings made us realize they were important. However, we did not try to establish their truth: they were used as an assumption. We could have rejected these norms and used different ones, for example when they conflict with other values we hold or scientific information about their feasibility Therefore, no naturalistic fallacy has been commited. However, seeing the status of moral norms as mere assumptions invites the criticism that science does not offer a definite justification, obligation or 'foundation' for any normative statement. To this, we can only say that we could not agree more: we fully endorse that science guides ethics conditionally, not absolutely. The Kibbutzim do not have to be organized that way, this is conditional on whether we accept these values or not. Science does guide ethics though, not by inferring true moral principles but by pointing us to which values we do hold and which value sets are incoherent. In the following section we will also argue that this quest for foundations is often misguided.
How do critics oppose the sketched conditional procedures? To answer this question we draw on the clarifications made in Sect. 2.2. There we argued that Moore's concept of the naturalistic fallacy is an argument against ethical foundations. Hence Moore's critique was aimed towards early evolutionary ethicists like Spencer who did commit the naturalistic fallacy; it is not used to criticize ethicists who do not provide such a foundation. Contemporary critics however, accuse current scientific ethicists (and more specifically, evolutionary ethicists) of committing the naturalistic fallacy, while at the same time critiquing them for not providing a foundation for ethics. Let us dig deeper in this request for foundations as done by contemporary critics of scientific ethics.
Several scientific ethicists have argued that scientific information can be used to argue for and against specific values (e.g., Flanagan 1996;Casebeer 2003). Some of these scholars grant a special role to evolutionary theories (e.g., Ruse 1995). The idea is that information about our evolved nature is particularly relevant to ethics because it highlights general human possibilities and constraints. Hence, evolutionary theories, together with empirical data that corroborate these theories, can guide normative ethics in the most general way. As Rosenberg (2000, 9) asserts, of all sciences evolutionary theory ''maximally combines relevance to human affairs and well-foundedness.'' Among scientific ethics, it is mostly this kind of evolutionary ethics that is under attack. This is understandable from a historical perspective. Some evolutionary ethicists did try to ground ethics in evolution by inferring a first moral principle from our evolved nature (Richards 1986;E. O. Wilson 1984). Most evolutionary inspired scientific ethicists however mainly indulge in the reasoning as sketched in the example (Ruse and Wilson 1986;Binmore 2005). Nonetheless, both accounts have been criticized.
As one of the established critics of scientific ethics, especially Paul Farber (1994) reviewed accounts of evolutionary ethics throughout history. His work demonstrates the same reasoning behind recent criticism against scientific ethics. Therefore Farber's The Temptations of Evolutionary Ethics is used as a template to analyze this criticism. According to Farber, sociobiology-which relates animal and human behavior to its evolutionary history-''offers no new hope, no new foundation'' for ethics (ibidem, 156). With this statement, Farber warns against reintroducing the naturalistic fallacy in evolutionary ethics, which is the most famous way of grounding ethics. However, should one abandon hope together with foundations?
Although Farber acknowledges the existence of nonfoundational accounts, he is little enthusiastic about them. He discusses a range of programs in twentieth-century evolutionary ethics, of which several do not commit the naturalistic fallacy and make no attempt at grounding anything. One of them is the strong program, which attempts to provide moral guidance by informing us about our biological nature. Farber rejects this program because ''an established picture of human nature from which to derive useful lessons is far away'' (ibidem, 160). About the weak program, which aims at an understanding of what morality is, Farber argues that it does not provide moral guidance. Still, he recognizes it as ''a possible source of relevant information'' (ibidem, 160) and adopts the ambitions of the weaker program in using scientific information ''in order to avoid misguided moralizing'' (ibidem, 160). This seems to hint at a contradiction, especially because 'the avoidance of misguided moralizing' can be taken at least as some kind of moral guidance. In the Kibbutzim example, we concluded that scientific discussions can lead to conditional moral guidance. Evolutionary information is a helpful guide for moral practice, exactly because it constrains the desirable possibilities, while it suggests otherwise unnoticed options.
Farber finds these approaches wanting and concludes pessimistically that ''the newest program for an evolutionary ethics looks […] unpromising as a theory of ethics'' (ibidem, 166-7). The only option he considers for evolutionary science is to provide a foundation for ethics (ibidem, 163-165). However, as argued in the discussion about the naturalistic fallacy, nothing can offer a foundation for ethics. Indeed, also Farber (ibidem,165) is aware that all attempts to construe a unified rational ethics have ''hit on hard times''. Consequently, if a foundationalist ethics proves to be impossible, why strive for it and not seek other alternatives?
Only at the end of his book, Farber briefly speculates on another possibility: ''perhaps if philosophers develop an ethical theory […] that is nonfoundationalist, evolutionary considerations may enter the philosophical arena'' (ibidem, 165). He tentatively mentions pragmatism and Rawls' Theory of Justice. But, then again, he adds, these ethical philosophers rarely mention evolutionary ethics. The possibility that their ethics could benefit from evolutionary findings is not even considered by Farber. He simply concludes that evolutionary ethics looks unpromising as a theory of ethics. We think that, given Farber's opposition towards committing the naturalistic fallacy, he should either consider a nonfoundationalist approach for evolutionary and scientific ethics or make clear what he intends with a theory of ethics.
Criticism like Farber's is well spread. Peter Woolcock, for example, argues that all the work in evolutionary ethics he studied committed the naturalistic fallacy. But he also claims that ''in order to have some normative relevance, a descriptive theory would seem to have to be able to leap the ''is/ought'' gap'' (Woolcock 1999, 290). And since evolutionary theory cannot leap this gap, he concludes that the naturalistic fallacy invalidates all efforts at an evolutionary ethics (ibidem, 282). In between lines, he does suggest that there can be other ways to ground ethics. For instance, he argues that ethical terms may not be ''identical in meaning with some natural property, nonetheless they might be identical in fact with some natural property, just as water does not mean ''H 2 O,'' even though in fact it is identical with H 2 O'' (ibidem, 284). But Woolcock does not consider this a serious option for science. Therefore, his argument is similar to that of Farber's: there is the impossible demand that a descriptive theory should leap the is/ought gap if it is to be relevant to ethics. At the same time, ethics that are inspired by scientific theories (in casu evolutionary theory) are accused of committing the naturalistic fallacy. This is inconsistent, unless Woolcock explains how the is/ought gap is different from the naturalistic fallacy in this regard (which he does not). Moreover, if nothing can ground ethics, considering grounding to be a criterion for ethical relevance is highly questionable.
Last but not least, Alexander Rosenberg acknowledges that science can inform ethics in the ways described here in Sect. 3.1. But he also claims that this is not enough: ''for a theory of human nature to have ramifications for moral philosophy itself, it will have to do more than any of these things'' (Rosenberg 2000, 120). According to Rosenberg, to be morally interesting, a theory of human nature must at least be able to derive some moral statement-a principle, value, obligation, etc.from a descriptive theory. One cannot begin with assumptions with normative content because then ''these assumptions are doing all the real work, and […] the biological theory makes no distinctive contribution to the derivation'' (ibidem, 120). Indeed, the normative assumptions in the Kibbutzim example do some of the work-but the scientific information is relevant, both for eliminating certain value sets because they are less consistent than others, as for pointing us towards certain values. Still, Rosenberg demands an independent derivation of moral statements from a descriptive theory if this descriptive theory is to be truly relevant to ethics. But why would he demand this? Even more so when taking that he, too, explicitly connects the derivation of first principles with the illegitimate bridging of the is/ ought gap: ''the possibility of deriving […] the existence of some moral principle […] rests on two preconditions. The first is that we can derive ''ought '' from ''is''''(ibidem, 120). Even though Rosenberg does not express his opinion on whether he accepts the reasoning behind the naturalistic fallacy or not, that this first precondition cannot be realized ''seems to me [Rosenberg] at least as widely held a view as any other claim in moral philosophy or meta-ethics'' (ibidem, 120). As Woolcock, perhaps he does not follow Moore's original interpretation of the naturalistic fallacy. Perhaps he too has some kind of foundation in mind that is not refuted by it. Unfortunately, once again, there is no indication that he really is considering such an alternative.
In sum, according to the discussed authors, scientific ethicists either commit the naturalistic fallacy or fail to make their descriptive theory morally relevant. This also counts when using evolutionary theory in order to ground ethics, as has been the case in several sociobiological and evolutionary epistemological approaches. Questioning when science would be relevant for normative ethics, these critics suggest that it should provide either a new foundation (Farber), leap the is/ought gap (Woolcock) or derive moral statements from a descriptive theory (Rosenberg). In light of the naturalistic fallacy, these suggestions are all impossible. This leads one to ask whether the authors either accept Moore's interpretation of the naturalistic fallacy or have a foundational ethics in mind that does not commit to Moore's reasoning. Only Farber suggested a way out of these impossibilities, namely that in a nonfoundational ethical theory, evolutionary considerations may be of relevance. While Farber never examined this option further, we already illustrated in Sect. 3.1 that scientific ethics can be promising even if one is not trying to 'ground' ethics. In what follows, we will argue that scientific ethics is also a philosophically underpinned theory. As such, there is no use to abandon hope together with 'foundations', as Farber does.foot_3 4 Naturalistic Ethics and the Methods of the Sciences It appears that science can inform ethics without committing the naturalistic fallacy and that common arguments against scientific ethics are misguided: critics demand scientific ethicists to provide a foundation for ethics while at the same time opposing an analytic ground for ethics. This, then, leaves to question what arguments we have for preferring nonfoundational over foundational ethics. We will first show that scientific ethicists-endorsing nonfoundational ethics-defend their nonfoundational normative system by appealing to methodological naturalism. This entails that we are interested in the proper method of normative inquiry. In defending methodological naturalism for normative inquiry, we follow a slightly modified reasoning than that pursued by the discussed scientific ethicists.
Certain scientific ethicists support their argument for ethics informed by science or 'ethical naturalism' by referring to methodological naturalism. As Flanagan et al. (2008, 5) argue: ''Ethical naturalism is not chiefly concerned with ontology but with the proper way of approaching moral inquiry'' (see also Flanagan 1996).
What does this method of moral inquiry consist of? On the one hand, Casebeer (2003, 9) refers to ''methodological naturalism'' as stating that ''the methodological and epistemological assumptions of the natural sciences should serve as standards for this inquiry.'' He asserts that ''robust moral norms […] can be constrained by and derived from the sciences'' (Casebeer 2003, 34). Consequently, he aims to develop a theory that helps to delineate those values that are conducive to human flourishing. Flanagan et al. (2008, 5) on the other hand argue that ''the claims of ethical naturalism cannot be shielded from empirical testing. […] ethical science must be continuous with other sciences''. They describe the method of naturalistic ethics as consisting of two components: a descriptive-genealogical component consisting of scientific descriptions and explanations of the moral phenomenon (normative practices, judgments and so on) (Flanagan et al. 2008) and a normative component drawing upon this information and either extracting successful normative practices from unsuccessful ones (Flanagan 1996) or considering which moral practices are part of what humans need and desire (Flanagan et al. 2008).
Does the concept of 'foundation' play a role in accounts of naturalistic ethics? Casebeer (2003) explains that one cannot analytically 'ground' ethics or find true moral principles by conceptual analysis. In other words, he recognizes that one cannot find an analytically true first principle-not because 'good' is a simple notion, but because the notion of finding truth by pure analysis (i.e., analytic truth) is flawed. His reasoning largely builds on Quine's Two Dogma's of Empiricism (1951) and is in contrast with Moore's reasoning which relied on the possibility of finding analytic truths. Also according to Flanagan et al. (2008, 5), ''moral philosophy should not employ a distinctive a priori method of yielding substantive, self-evident and foundational truths from pure conceptual analytical testing''.
Consequently and importantly, naturalists like Casebeer and Flanagan do not rely on analytic statements when backing up their moral principles with facts or when proposing certain universal moral values. Their arguments are not about the very meaning of a moral word or about the true nature of a moral notion. If equal worth is good, it means that there are scientific and moral arguments to endorse equal worth and that you can disagree and give counterarguments: ''With regard to the alleged is-ought problem, the smart naturalist makes no claims to establish demonstratively moral norms. He or she points to certain practices, values, virtues and principles as reasonable based on inductive and abductive reasoning'' (Flanagan et al. 2008, 14).
Despite subtle differences, Flanagan and Casebeer share the same basic picture (see also Casebeer 2003, 34). We interpret both as stating that, if we accept certain concrete values, then we can use scientific methods and findings to distinguish right from wrong conduct. This is in accord with the example where science was conditionally relevant for ethics without offering a foundation for ethics. Hence one can never fully determine which values are worth pursuing. But science can give arguments for or against them. As such, the naturalist method of normative inquiry is not about building normative theories on independently derived first moral principles. Instead, it draws on the existing pool of moral practices and values and all the scientific information to be found about them. These practices and values are evaluated in the light of other values and in the light of what can reasonably be valued by human beings. 7 Still, to answer why this method of normative inquiry is preferential, we need to look into the rationale behind naturalism. This will also lead to further clarifications of how science is relevant for scientific ethics in ways it is not for foundational ethics.
Naturalism is committed to the methods and findings of science (Rosenberg 2000;Casebeer 2003). To find out if these methods and findings can be applied to normative ethics, let us look into the basic idea behind scientific inquiry. Basically, it is considered legitimate to engage oneself to a specific constellation of methods and aims when this constellation has been shown to be more productive-that is, more successful in leading to a predetermined aim-than another constellation. According to Rosenberg for instance, naturalism implies that the methods of the natural sciences are to guide philosophy because of the contingent historical fact that science has been more successful than any other approach in predicting new phenomena and exerting control over the physical world (Rosenberg 2000). This successful constellation of methods and aims hence became the standard for scientific inquiry.
An example can clarify the notion of success. Fred Wilson (2007, 251-252) has reviewed methods and aims used throughout the history of natural philosophy. Before the sixteenth, seventeenth century, for instance, 'rational intuition' was thought to give one direct access to natural laws. Some patterns in nature were supposed to reflect natural laws or motions, others to reflect unnatural motions. Natural motions were thought to be essential to a particular substance (e.g., falling down is essential to an earthy object), unnatural motions were thought to be induced by an external substance (e.g., the parabolic motion of a projectile is not essential to the object; someone or something-an external substance-must have thrown it to give the object its forward thrust). Natural laws, so it was believed, could not be observed; they were to be found by the method of rational intuition. Science was to deduce these natural laws. However, this conviction did not lead to great progress in questions such as projectile motion. Galileo changed the aims: one should not seek to distinguish the natural laws versus the unnatural motions. One should try to find exceptionless patterns in the observable world and forget about whether they are essential or not to the object. Galileo also changed the method: these patterns can be found by observation and experiments on the behaviour of changing things. This new science was very successful (F. Wilson 2007, 254). Therefore, observation came to have a more prominent role in the scientific method while the aim of distinguishing natural versus unnatural motions was abandoned.
This leaves the question whether the modern method and aim of science can serve as standard for normative inquiry. According to Rosenberg, science aims to predict and control the natural world (Rosenberg 2000). According to Ernst Nagel, science aims to provide systematic and supported explanations (E. Nagel 1961, 15), enabling the explanation and prediction of new phenomena that were not yet in the evidence on which the explanation was built (ibidem, 64). Are these aims the same as those of normative inquiry? In the literature, several objects have been postulated as the aim of ethics. We already saw that, according to Moore (Moore 1993(Moore /1903, §14), §14), 'Ethics' must aim at truth. Others, like Warnock, situate the object of morality in the amelioration of the human predicament (Warnock 1971, 16) while Thomas Nagel identifies morality as the combination of a personal perspective with an objective perspective (T. Nagel 1985, 3). While many other proposals exist, most of them do not consider it the aim of normative inquiry to explain, predict or control what will happen. Hence, we consider it problematic to take the aim of science and make this into the aim of normative ethics.
What about the methods of science? The natural sciences typically test hypotheses against observations. When inconsistencies are discovered, hypotheses or theories are adjusted. Data from observations are only seldom adjusted because the existing methods allow obtaining reliable data from observation. Reliable data are the same when gathered under the same experimental circumstances, and they are objective in that everybody is able to see or (re)confirm the same raw data. But even when taking that values are amenable to observation, we do not (or not yet) have an experimental method or theory to gather raw data in a way that makes everybody see, or be convinced by, the same values. As a result, as things stand now, one cannot simply copy the aim and method of science to normative inquiry. So how can normative ethics proceed? What are the criteria for successful normative ethics, analogous to the criteria for successful science?
Casebeer (2003) asks a similar question and suggests that we naturalize normativity. In his proposal, all moral terms can be reduced to functional terms (Casebeer 2003, 38): ''To live the life informed and motivated by practical reason and wisdom is to live a functional life'' (ibidem, 42). Furthermore we can understand all functional facts within a materialist ontological framework (ibidem, 54). Casebeer goes on developing a theory of functions that is scientific and useful in biology as well as in normative theory: ''Value properties […] are scientifically tractable in the same way that biological notions of function are'' (ibidem, 55). Hence, he develops an encompassing moral theory that is amenable to scientific testing. It follows that according to Casebeer's naturalized normativity, the aim and method of the natural sciences can be applied to morality. It must be stressed that his theory is not deduced from pure analytical statements that are demonstrated to be true. His theory consists of statements with conceptual and empirical content; it is also deemed internally consistent and supported by empirical knowledge. Hence, his theory must not be discussed by reference to rationality only; one can give empirical and conceptual arguments for and against it.
Our approach may be compatible with Casebeer's but does not suggest an allencompassing general theory that translates normative terms into descriptive or factual terms. We want to focus on how to tackle concrete day-to-day moral questions. Hereto, we take the previously sketched reasoning behind naturalism in science and apply it to ethics. We hence ask the empirical question what constellation of aims and methods until now has been most successful in normative inquiry. We consider a method of inquiry to be successful if its methods lead to her predetermined purpose. Two questions of interest to our project here can be considered:(1) how successful foundational ethics has been, in solving specific moral problems compared to the method in Sect. 3.1 and (2) how science is relevant to normative ethics in a nonfoundational system. Let us turn to the first question.
The twentieth century was dominated by analytical ethics, which gained attention thanks to Moore's Principia Ethica. As a field, it grew out of a strong rebuttal of the possibility of analytical normative ethics Analytical ethicists did not primarily aim to discuss normative questions, but rather examined the meaning of moral terms and moral judgments and aimed for analytic truths in ethics. Analysis hence was mainly used in the domain of meta-ethics and not in the domain of normative ethics. Thus we can at least conclude that analytical ethics was never meant to lead to normative progress. However, the focus on analysis in the twentieth century seemed to suggest that this was the preferred method for all ethics. Moreover, practical moral choices always side with or against certain theoretical positions. Still, the relevance and merits of analytical ethics for normative ethics is contested. Holmes (1990), for instance, discusses the relevance of analytical ethics for bioethics. He argues that analytical ethics can only clarify normative issues and cannot provide moral wisdom. Similarly, while agreeing that conceptual analysis can clarify the logical connections between moral concepts, he doubts that it can resolve which normative theory is true or a better solution. Therefore he advised that bioethicists do not turn to conceptual analysis to solve their problems (Holmes 1990). A similar pessimism towards foundational normative ethics is found in Farber's work. Farber mentions that philosophers since Sidgwick have tried to systematize morality, but without success (Farber 1994, 165). Also Edward O. Wilson (1975, 562) described the result of analytical ethics in the twentieth century as ''several oddly disjunct conceptualizations''. Naturalism does not reject analysis per se, but it rejects the possibility of finding true statements by means of pure conceptual analysis. It thus rejects the suitability of this particular method for the specific aim of finding true statements; or stronger, it rejects the plausibility of ever finding analytic truths.
This supports the conclusion of the naturalistic fallacy, namely that one cannot ground norms in facts. Indeed, naturalism offers a genuine reason for why one should not 'ground' ethics 8 and practically neutralizes the criticism against scientific ethics that it would commit the naturalistic fallacy. One can reasonably expect that 8 According to Casebeer (2003), Quine's (1951) argument also shows that Moore's reasoning behind the naturalistic fallacy is incorrect, even though its conclusion holds. This is because Moore's reasoning behind the naturalistic fallacy assumes that we can find true statements by analyzing the meaning of the words, without any observational input. For a more elaborate discussion, see Casebeer (2003) and Quine (1951). Naturalists like Flanagan, Casebeer and Ruse hence accept the conclusions of Moore's naturalistic fallacy without necessarily accepting the reasoning behind it. scientific ethicists who explicitly endorse naturalism as here presented do not rely on analytical statements or first principles. In fact, this is the case with some authors who have been accused of committing the naturalistic fallacy. Ruse, for example, claims that he is grounding ethics and is consequently refuted by Woolcock for committing the naturalistic fallacy. But Ruse explicitly endorses the 'is/ought' gap. A closer look teaches us that with 'grounding' Ruse certainly does not aim to analytically derive a first principle (Ruse 1995). If however naturalistic scientific ethicists do commit the naturalistic fallacy, we can poignantly accuse them of contradicting explicitly endorsed naturalistic commitments.
Finally, the historical reasoning as proposed here can continuously and empirically be applied to the question of which local aim and method in ethics is most successful. It is here that the relevance of scientific findings for normative ethics has to be laid out. In effect, naturalists can interpret values and value sets as local aims. One can try to promote these by means consistent with our values. In the example of the Kibbutzim, the aim of total equality conflicts with our values of freedom life choice and life satisfaction. This can explain why the Kibbutzim did not reach their goal of total equality and it is unlikely given experience and scientific information that it ever would. Hence a decision was made to try something else. This is consistent with Flanagan's (1996) principle of drawing successful practices from unsuccessful ones. 9What practices do we choose from? Here we saw that science informs us about other options. When the method of promoting paternal care alone hardly relieves working mothers, other possibilities could be considered based on recent findings about the evolution of childcare. As Flanagan holds, we import our values from the values we already hold; scientific descriptions and explanations of the moral phenomenon-such as naturalistic descriptions of normative practices, judgments and so on-can help us with this (Flanagan et al. 2008).
Important is that all adaptations to our value systems are conditional on other values. When arguing for a moral rule we always draw on the pool of values we hold, rejecting and strengthening norms as the resulting system is more or les consistent and successful. As such, all values can be revised in the light of new evidence. This means that moral decisions are never absolute but change as knowledge about the situation grows. This dynamic view on morality here differs from a foundational account: when introducing a foundation this value and all that follows from it cannot be revised in the light of new evidence about other values we (want to) hold.
Today, many philosophers still aim at establishing a normative system built on an unimpeachable foundation; or they demand such a foundation from others. At the same time, they refer to the naturalistic fallacy as a legitimate criticism against instantiations of scientific ethics, mostly evolutionary ethics. We have shown that both arguments when used together contradict each other; we argued that no fallacy is committed in the work they criticize. We also countered the assumption that ethics needs to be foundational and that science is not relevant for normative ethics. Though agreeing that science loses some of its relevance for foundational ethics, we claim that science is highly relevant in nonfoundational ethics. Crucially, we reasoned that scientific ethics is best conceived of as an instance of such a nonfoundational normative ethics. We believe the debate between proponents and opponents of scientific ethics would benefit from recognizing scientific ethics as nonfoundational. Much of the discussed disagreement occurred because nonfoundationalist proponents were debated within a foundationalist framework; therefore the discussion should be focused on this difference.
In the last sections, we discussed and argued for the nonfoundationalist view of ethics. Defenders of scientific ethics refer to naturalism to support their view. Naturalists take the implausibility of a foundational ethics at face value and endorse another approach. We proposed a slightly modified naturalistic reasoning in support of scientific ethics. Our approach does not aim at building a grand philosophical theory but suggests that a hands-on method for normative inquiry can give more direct success. Normative inquiry can be aimed at local and concrete problem solving, wherein a moral problem is never absolutely solved. It is thereby a challenging approach that demands regular reassessment of a moral problem while science proceeds and offers new information. As naturalists, analytic truth is not our aim and the search for first foundations is rejected in favour of conditional moral judgments that can be tested on their practical success. 10
We discuss the notion of success in Sect. 4.2.
If a moral norm would be determined by something else than a descriptive theory (e.g. direct intuition), but irrefutable in the light of other moral norms, we would still call it a 'first moral norm'. It would just not be a descriptively determined first moral norm.
The term 'nature' here is referring to a modern, non-teleological view of a mechanistic observable and physical world. This is different from 'nature' in natural law theories, where purpose and normativity are taken to be part of the world, hence of physical nature.
Unwarranted criticism of scientific ethics, as laid out here, is in fact more widespread than this discussion of scholarly criticists may show. There seems to be the idea that normative ethics has to be foundational among the foundational theorists we discussed in Sect. 2.1. Also Blancke and Quintelier (under review) illustrate that a similar kind of criticism is enthusiastically propagated by creationist propaganda. Specifically, the creationist movement accuses evolutionary ethicists of committing the naturalistic fallacy while at the same time demanding evolutionary ethicists to provide a foundation for ethics. Because of the social relevance of this criticism and the widespread conception of normative ethics as foundational, we think it is important to defend nonfoundational ethics wherever we find it under attack.
Naturalists' views on normative ethics share similarities with the pragmatic tradition in ethics. For exampleRorty (2007), a recent pragmatist, also rejects the quest for foundations for a historically contingent epistemology. Several of the here described naturalists are influenced by and explicitly refer to the work of one classical pragmatist, John Dewey (diCarlo and Teehan 2007; Casebeer 2003). David B. Wong, another naturalistic ethicist elaborates his naturalistic ethics by contrasting it with the work of, among others, Rorty. Nonetheless, even though there is mutual interest between pragmatic and naturalistic ethicists, it would be interesting if recent pragmatist and naturalistic ethics would be more intertwined. One can imagine a project where a naturalist elaborates on pragmatist theories in the light of a naturalist framework, or the other way around. This could stimulate discussion and integrate both views with each other.
Our argument for nonfoundational ethics here is that nonfoundational ethics is bound to be more successful than foundational ethics. While this is an advantage of nonfoundational ethics, there may also be disadvantages. It can be argued that the open-endedness of this endeavor is a drawback. However, this only holds compared to a successful foundational ethics, meaning that the first norm would finally be accepted by a large majority. In reality though, we see the same open-endedness in foundational ethics: since no foundational ethics ever reached the point where the first norm was accepted by a large majority, we had and have to give arguments for and against all norms in a foundational system as well. Another disadvantage of our scientific ethics is that it requires a shift from universal rules to the values all individuals hold: it is democratic. Therefore scientific ethics are less likely to be accepted in nondemocratic societies. However, nondemocratic societies should still defend why they value antidemocracy more than holding a successful normative ethic.
This is not a new foundation for ethics. Nowhere in this paper did we deductively infer this principle from other information; nor did we state that it absolutely fixed. We did give arguments in favor but it can be refuted in the light of new knowledge or moral prescripts that are found to follow from it.
Acknowledgments
Open Access This article is distributed under the terms of the Creative Commons Attribution Noncommercial License which permits any noncommercial use, distribution, and reproduction in any medium, provided the original author(s) and source are credited.
The Pi sampling method is derived from the incomplete factorial approach to macromolecular crystallization screen design. The resulting 'Pi screens' have a modular distribution of a given set of up to 36 stock solutions. Maximally diverse conditions can be produced by taking into account the properties of the chemicals used in the formulation and the concentrations of the corresponding solutions. The Pi sampling method has been implemented in a web-based application that generates screen formulations and recipes. It is particularly adapted to screens consisting of 96 different conditions. The flexibility and efficiency of Pi sampling is demonstrated by the crystallization of soluble proteins and of an integral membrane-protein sample.
A crucial aspect of macromolecular crystallographic studies is finding suitable conditions for the crystallization of a sample. This can be difficult because many factors alter the crystallization behaviour of macromolecules, including the type and the concentration of the chemicals employed to formulate the conditions (McPherson, 1990). A condition includes at least a precipitant and most conditions also include a buffer and an additive. During the initial crystallization experiments, the structure of the macromolecule is not known and hence the most efficient formulation cannot be predicted. As a consequence, one should be cautious when making initial assumptions and limiting choices in subsequent optimizations (Rupp, 2003). Nonetheless, the number of initial crystallization conditions cannot be unreasonably large since purified protein is often difficult and expensive to produce in large quantities.
There are essentially two approaches to restrict an initial screen to a limited number of crystallization conditions. Firstly, a sparse-matrix formulation can be used, which consists of an empirically derived combination of components based on known or published crystallization conditions (Jancarik & Kim, 1991). Secondly, an incomplete factorial formulation can be generated in which selected components are combined to prepare new conditions in accordance with principles of randomization and balance (Carter & Carter, 1979). Numerous commercial screens based on these two main approaches are available. Automated systems have been implemented at the Medical Research Council (MRC) Laboratory of Molecular Biology (LMB) to test these as routine initial screens using the 96-well crystallization plate format (Stock et al., 2005). However, for various reasons, many laboratories opt for a minimal screen (Kimber et al., 2003) and still perform at least some aspects of the work manually (Bergfors, 2007).
Here, we present a development based on the incomplete factorial formulation: the Pi sampling method. The name of the method was inspired by the story of Archimedes, who used the 'method of exhaustion' (i.e. an empirical approach) with a 96-sided polygon in order to reach the first good numerical approximation of (Smith, 1958). Pi sampling uses modular arithmetic to form combinations of three stock solutions across a 96-condition grid. Maximally diverse conditions can be produced by taking into account the properties of the chemicals used in the formulation and the concentrations of the corresponding stock solutions. We have implemented this approach in a web-based application called Pi Sampler: user input consists of the details of up to 36 stock solutions, from which the application generates the formulations for a 96-condition screen. The Pi sampling method is intended to help laboratories to test new crystallization-screen formulations on a day-to-day basis based on the properties of the macromolecules investigated, as has been performed previously with RNA (Doudna et al., 1993).
Firstly, we tested Pi sampling with ten commercially available soluble proteins. For this, the 'Pi minimal screen' was employed including a wide variety of well known chemicals frequently used for macromolecular crystallization.
We then investigated the impact of Pi sampling on the crystallization of a G-protein-coupled receptor (GPCR) that had been difficult to crystallize: the adenosine A 2A receptor (construct A 2A R-GL31). We formulated another Pi screen, the 'Pi-PEG screen', taking into consideration general observations made about crystallization of integral membrane-protein samples. Previous crystallization experiments on another GPCR (the 1 -adrenergic receptor) had indicated that the use of simple proprietary screens formulated with poly(ethylene glycol) (PEG) and buffers gave a greater yield of crystals than all commercially available screens, including those geared towards membrane proteins (Warne et al., 2009), and the 2.7 A resolution structure was solved using conditions optimized from a proprietary screen essentially based on PEGs (Warne et al., 2008). This has been observed previously with other membrane-protein targets (Lemieux et al., 2003). In addition, mixtures of polyethylene glycols have been used successfully to develop a minimal screen (Brzozowski & Walton, 2001) and to study crystal structures of the Kir potassium channel (Clarke et al., 2010). Such mixtures were incorporated into the Pi-PEG screen.
Pi sampling begins with up to 36 stock solutions, divided into three sets of 12. The first set of solutions is used in the screen at constant concentration. The second and third sets are added according to a gradient between specified minimum and maximum concentrations. Typically, the first set is composed of buffers and the second and third sets are precipitants/ additives.
The combinations of three stock solutions (one from each set) are generated according to Fig. 1, where 1-12 refer to the IDs for solutions of the first set, A-M to those of the second set and N-X to those of the third set. The number in each cell shows which solution of the first set will be combined with the corresponding solutions of the second and third sets. Blank spaces show when no such combinations are generated.
Fig. 2 summarizes the distribution of the stock solutions in a standard 96-condition plate layout (i.e. 12 columns and eight rows).
Set 1: each solution (ID 1-12) is seen in the eight conditions forming a column of the plate. A variable Á should be associated with the stock solutions. The variable Á corresponds to a property of the solution selected (e.g. pH, molecular weight of the main chemical, absorption properties or others). Á values increase from left to right in the screen layout.
Set 2: each solution (ID A-L) is represented once in each row. The final concentrations decrease gradually from the top to the bottom of the screen layout, forming a gradient. The distribution of solutions A-L is based on the sequence of Á values established for set 1: the positions of the solutions shift across five columns and down one row. Solutions A-L should also be associated with a variable Á and hence a sequence is formed for the distribution of the third set of solutions.
Set 3: each solution (ID M-X) is also represented once in each row. The final concentrations increase gradually from the top to the bottom of the screen, forming another gradient. The solutions M-X are distributed with the same modulo arithmetic operation as previously, but with respect to the Á values of solutions A-L. For example, solution M is mixed with solution A in the first row, solution F in the second row, solution K in the third row and so on, as shown in Fig. 2. This means that both the second and third sets are arranged according to the same modulo arithmetic operation (5 modulo 12); however, when looking at the plate layout, the positions of solutions M-X shift across ten columns and down one row.
Pi Sampler can be accessed via the internet at http:// pisampler.mrc-lmb.cam.ac.uk/. Users can enter the details of up to 36 stock solutions, including stock concentrations, desired screen concentration ranges and Á values. The application then generates a 96-condition screen formulation following the Pi sampling method described above. Formulations, recipes and total required volumes of stock solutions are presented and may conveniently be downloaded in comma-separated variable format (CSV), allowing the user to import them into other software for automated screen making (Cox & Weber, 1987), formulation analysis (Hedderich et al., 2011) and data mining (Kantardjieff & Rupp, 2004). The parameters used to generate the screen can also be saved and uploaded in the same format. Further details and instructions can be found on the website.
The final formulation of the Pi minimal screen can be found in Table 1. There are 36 starting stock solutions overall. Each solution composing the first set (ID 1-12) is a mixture of an acid with its corresponding base (e.g. HEPES pH 7.5: 1 M HEPES solution mixed with 1 M HEPES sodium salt in order to reach pH 7.5), except for buffer 11 (AMPD mixed with Tris base). Note that this is also true for the precipitant phosphate (phosphate system: sodium dihydrogen phosphate/dipotassium hydrogen phosphate). Values of pH (4.0-9.5) were chosen as the variable Á for the first set, whilst arbitrary values were chosen for additives of various natures composing the second set (ID A-L). Eventually, a few conditions were made without additive/buffer because of chemical incompatibilities (Table 1).
Highest purity grade chemicals (Molecular Biology grade when available) were purchased from Sigma-Aldrich to prepare 36 stock solutions. The solutions were mixed in 96 Falcon tubes. The screen was dispensed into 'MRC original plates' (96-well, two-drop, Swissci; Stock et al., 2005).
Commercial proteins that had been crystallized before were chosen to prepare test samples. Protein concentrations were chosen randomly between 7 and 150 mg ml À1 (Table 2). Vapour-diffusion experiments were set up at 295 K, mixing two different sample: condition ratios (1:3 and 3:1) to give a final volume of 400 nl. The plates were then stored at 291 K. A condition was considered to be a hit when at least one of the two corresponding drops contained crystals with well known morphology after one week. Table 3 shows the 'hits per condition' observed and the corresponding results expected for the binomial distribution (see x4.2).
The final formulation of the Pi-PEG screen can be found in Table 4. The formulation can also be generated using Pi Sampler by loading the Pi-PEG example data. The pH values (4.8-8.8) were chosen as the variable Á for the buffers composing set 1 (ID 1-12), whilst molecular weight was chosen for set 2 (PEGs A-L, final concentration range 0-22.5%). The same 12 PEGs were used for set 3 (PEGs M-X, final concentration range 0-45%). General details of the preparation are similar to x2.3, but there are 24 stock solutions at the start (instead of 36). Vapour-diffusion experiments were set up at 277 K, mixing sample and condition in a 1:1 ratio to give a final volume of 200 nl. The preparation of A 2A R-GL31 will be published elsewhere (Lebon et al., submitted work). Crystal X-ray screening was performed at the Diamond
Table 1
Final formulation of the Pi minimal screen.
ADA, N-(2-acetamido)iminodiacetic acid; AMPD, 2-amino-2-methyl-1,3-propanediol; CAPSO, 3-(cyclohexylamino)-2-hydroxy-1-propanesulfonic acid; HEPES, 4-(2-hydroxyethyl)piperazine-1-ethanesulfonic acid; MOPS, 3-(N-morpholino)propanesulfonic acid; PEG, poly(ethylene glycol); TAPS, N-[Tris(hydroxymethyl)methyl]-3-aminopropanesulfonic acid.
Set 1 Set 2 Set 3 Well ID Name Conc. Unit ID Name Conc. Unit ID Name Conc. Unit A1 1 Formate pH 4.0 0.15 M A Potassium bromide 0.160 M M Phosphate 0.6 M A2 2 Acetate pH 4.5 0.15 M B PEG 300 8.000 %(v/v) N PEG MME 550 24.00 %(v/v) A3 3 Malate pH 5.0 0.15 M C Magnesium sulfate 0.160 M O Ammonium nitrate 2.0 M A4 4 Citrate pH 5.5 0.15 M D Sodium fluoride 0.032 M P PEG 20 000 10.0 %(w/v) A5 5 MES pH 6.0 0.15 M E Potassium thiocyanate 0.080 M Q PEG 1000 30.0 %(w/v) A6 6 Cacodylate pH 6.5 0.15 M F Sodium iodide 0.160 M R Sodium chloride 1.6 M A7 7 MOPS pH 7.0 0.15 M G Propanediol 8.000 %(v/v) S PEG 4000 24.0 %(w/v) A8 8 HEPES pH 7.5 0.15 M H T Lithium sulfate 0.8 M A9 9 Tris pH 8.0 0.15 M I Ethylene glycol 8.000 %(v/v) U PEG MME 5000 20.0 %(w/v) A10 10 TAPS pH 8.5 0.15 M J Sodium potassium tartrate 0.080 M V Glycerol 36.0 %(w/v) A11 11 AMPD/Tris pH 9.0 0.15 M K MPD 8.000 %(v/v) W Ammonium sulfate 1.4 M A12 12 CAPSO pH 9.5 0.15 M L 2-Butanol 8.000 %(v/v) X PEG 8000 20.0 %(w/v) B1 1 Formate pH 4.0 0.15 M H Calcium chloride 0.070 M O Ammonium nitrate 2.3 M B2 2 Acetate pH 4.5 0.15 M I Ethylene glycol 7.000 %(v/v) P PEG 20 000 12.0 %(w/v) B3 3 Malate pH 5.0 0.15 M J Sodium potassium tartrate 0.070 M Q PEG 1000 35.0 %(w/v) B4 4 Citrate pH 5.5 0.15 M K MPD 7.000 %(v/v) R Sodium chloride 1.8 M B5 5 MES pH 6.0 0.15 M L 2-Butanol 7.000 %(v/v) S PEG 4000 28.0 %(w/v) B6 6 Cacodylate pH 6.5 0.15 M A Potassium bromide 0.140 M T Lithium sulfate 0.9 M B7 7 MOPS pH 7.0 0.15 M B PEG 300 7.000 %(v/v) U PEG MME 5000 23.0 %(w/v) B8 8 HEPES pH 7.5 0.15 M C Magnesium sulfate 0.140 M V Glycerol 42.0 %(w/v) B9 9 Tris pH 8.0 0.15 M D Sodium fluoride 0.028 M W Ammonium sulfate 1.6 M B10 10 TAPS pH 8.5 0.15 M E Potassium thiocyanate 0.070 M X PEG 8000 23.0 %(w/v) B11 11 AMPD/Tris pH 9.0 0.15 M F Sodium iodide 0.140 M M Phosphate 0.7 M B12 12 CAPSO pH 9.5 0.15 M G Propanediol 7.000 %(v/v) N PEG MME 550 28.00 %(v/v) C1 1 Formate pH 4.0 0.15 M C Magnesium sulfate 0.120 M Q PEG 1000 39.0 %(w/v) C2 2 Acetate pH 4.5 0.15 M D Sodium fluoride 0.024 M R Sodium chloride 2.1 M C3 3 Malate pH 5.0 0.15 M E Potassium thiocyanate 0.060 M S PEG 4000 31.0 %(w/v) C4 4 Citrate pH 5.5 0.15 M F Sodium iodide 0.120 M T Lithium sulfate 1.0 M C5 5 MES pH 6.0 0.15 M G Propanediol 6.000 %(v/v) U PEG MME 5000 26.0 %(w/v) C6 6 Cacodylate pH 6.5 0.15 M H Calcium chloride 0.060 M V Glycerol 47.0 %(w/v) C7 7 MOPS pH 7.0 0.15 M I Ethylene glycol 6.000 %(v/v) W Ammonium sulfate 1.8 M C8 8 HEPES pH 7.5 0.15 M J Sodium potassium tartrate 0.060 M X PEG 8000 26.0 %(w/v) C9 9 Tris pH 8.0 0.15 M K MPD 6.000 %(v/v) M Phosphate 0.8 M C10 10 TAPS pH 8.5 0.15 M L 2-Butanol 6.000 %(v/v) N PEG MME 550 31.00 %(v/v) C11 11 AMPD/Tris pH 9.0 0.15 M A Potassium bromide 0.120 M O Ammonium nitrate 2.6 M C12 12 CAPSO pH 9.5 0.15 M B PEG 300 6.000 %(v/v) P PEG 20 000 13.0 %(w/v) D1 1 Formate pH 4.0 0.15 M J S PEG 4000 35.0 %(w/v) D2 2 Acetate pH 4.5 0.15 M K MPD 5.000 %(v/v) T Lithium sulfate 1.1 M D3 3 Malate pH 5.0 0.15 M L 2-Butanol 5.000 %(v/v) U PEG MME 5000 38.00 %(v/v) D4 4 Citrate pH 5.5 0.15 M A Potassium bromide 0.100 M V Glycerol 52.0 %(w/v) D5 5 MES pH 6.0 0.15 M B PEG 300 5.000 %(v/v) W Ammonium sulfate 2.0 M D6 6 Cacodylate pH 6.5 0.15 M C Magnesium sulfate 0.100 M X PEG 8000 29.0 %(w/v) D7 7 MOPS pH 7.0 0.15 M D Sodium fluoride 0.020 M M Phosphate 0.9 M D8 8 HEPES pH 7.5 0.15 M E Potassium thiocyanate 0.050 M N PEG MME 550 34.00 %(v/v) D9 9 Tris pH 8.0 0.15 M F Sodium iodide 0.100 M O Ammonium nitrate 2.9 M D10 10 TAPS pH 8.5 0.15 M G Propanediol 5.000 %(v/v) P PEG 20 000 15.0 %(w/v) D11 11 H Calcium chloride 0.050 M Q PEG 1000 43.0 %(w/v) D12 12 CAPSO pH 9.5 0.15 M I Ethylene glycol 5.000 %(v/v) R Sodium chloride 2.3 M E1 1 Formate pH 4.0 0.15 M E Potassium thiocyanate 0.040 M U PEG MME 5000 32.0 %(w/v) E2 2 Acetate pH 4.5 0.15 M F Sodium iodide 0.080 M V Glycerol 57.0 %(w/v) E3 3 Malate pH 5.0 0.15 M G Propanediol 4.000 %(v/v) W Ammonium sulfate 2.2 M E4 4 Citrate pH 5.5 0.15 M H X PEG 8000 32.0 %(w/v) E5 5 MES pH 6.0 0.15 M I Ethylene glycol 4.000 %(v/v) M Phosphate 0.9 M E6 6 Cacodylate pH 6.5 0.15 M J Sodium potassium tartrate 0.040 M N PEG MME 550 38.0 %(v/v) E7 7 MOPS pH 7.0 0.15 M K MPD 4.000 %(v/v) O Ammonium nitrate 3.1 M E8 8 HEPES pH 7.5 0.15 M L 2-Butanol 4.000 %(v/v) P PEG 20 000 16.0 %(w/v) E9 9 Tris pH 8.0 0.15 M A Potassium bromide 0.080 M Q PEG 1000 48.0 %(w/v) E10 10 TAPS pH 8.5 0.15 M B PEG 300 4.000 %(v/v) R Sodium chloride 2.5 M E11 11 AMPD/Tris pH 9.0 0.15 M C Magnesium sulfate 0.080 M S PEG 4000 38.0 %(w/v) E12 12 CAPSO pH 9.5 0.15 M D T Lithium sulfate 1.3 M F1 1 Formate pH 4.0 0.15 M L 2-Butanol 3.000 %(v/v) W Ammonium sulfate 2.4 M F2 2 Acetate pH 4.5 0.15 M A Potassium bromide 0.060 M X PEG 8000 35.0 %(w/v) F3 3 Malate pH 5.0 0.15 M B PEG 300 3.000 %(v/v) M Phosphate 1.0 M F4 4 C Magnesium sulfate 0.06 M N PEG MME 550 42.00 %(v/v) F5 5 MES pH 6.0 0.15 M D Sodium fluoride 0.012 M O Ammonium nitrate 3.4 M synchrotron light source (microfocus beamline I24 equipped with a Pilatus 6M detector).
There were 116 crystallization hits overall for the experiments with the Pi minimal screen (Table 2). Some conditions produced hits for several samples (Table 3).
The Pi-PEG screen yielded crystals that diffracted to 3.0 A ˚resolution for A 2A R-GL31 with bound agonist. Fig. 3 shows the crystals of A 2A R-GL31 obtained in well E9 [50 mM Tris-HCl pH 7.6, 9.6%(v/v) PEG 200, 22.9%(v/v) PEG 300] and an example of the corresponding diffraction pattern (no cryoprotectant was required).
In order to understand the rationale behind the modular arithmetic employed for the Pi sampling, it may help to imagine, on a 12 h clock, a series of events occurring every 5 h. The first event is at noon, the second at 5 pm, then 10 pm, then 3 am etc.
Eventually, there is a succession of 12 events occurring at different hours, with as much time as possible in between each event. If we now look at combinations of three components, there are originally 12 3 or 1728 possibilities. Pi Sampler generates 96 of these combinations that correspond to conditions that are distant in properties. The variety between conditions is then accentuated using a number of different concentrations of solutions (Fig. 2). If the first and second sets of solutions are ordered according to physico-chemical properties, the generated screen will be an incomplete factorial sampling of interactions between chemicals with these properties. If the chemicals selected have completely different natures, they can be arranged randomly (see x2.3).
The ordering of the third set of solutions can be used to avoid obvious chemical incompatibilities (e.g. mixing phosphate and magnesium salts). It is also possible to design simpler screens with only two sets of stock solutions.
In order to check the homogeneity of the hits across the screen with the ten samples, we compared the results obtained with what would be expected if each condition had the same probability of hits overall (Table 3). This can be approximated by a binomial distribution. The probability of success for the binomial distribution is the observed probability for ten attempts: 116/ (10 Â 96) = 0.12083. The 2 statistic for the data is 3.48. This can be compared with the quantiles of a 2 distribution with two degrees of freedom, which gives a p value of 0.18 (calculations not shown). This 2 test indicates that no Figure 3
Crystals of A 2A R-GL31 obtained with the Pi-PEG screen (Table 4) and an example of a corresponding diffraction pattern.
conditions are obvious outliers with regard to success or failure. There are, however, a multitude of possible biases implied when proceeding with crystallization experiments (which would be even more accentuated with the use of novel samples); hence, any statistical analysis should be taken with precaution. Nonetheless, it is interesting to see that the analysis of the distribution is in accordance with the original approach based on balanced randomization (Carter & Carter, 1979;Rupp, 2003).
In addition, the conditions of the Pi minimal screen show no identities to the extensive list of conditions (7230) from commercial screens stored in the 'PICKScreens' database (Hedderich et al., 2011).
The extent of effects on crystallization for precipitants such as PEGs is correlated with their concentrations (McPherson, 1976) and molecular weights (Forsythe et al., 2002). The Pi-PEG screen covers a wide range of parameters (kinetics of equilibrium, protein stabilization etc.). In addition, the concentrations of the two different PEGs in a condition can be adjusted for condition optimization (Stock et al., 2005) and for crystal cryoprotection (Berejnov et al., 2006). Furthermore, the PICKScreens database shows that the Pi-PEG screen is unique (as for the Pi minimal screen; see x4.2).
Samples of A 2A R-GL31 purified in a number of different detergents rarely crystallized in commercially available screens used at the LMB (Stock et al., 2005) and when they did the crystal quality was not sufficient for structure determination. The first quality crystals were recently obtained using the Pi-PEG screen.
We have demonstrated that the Pi sampling is a methodical and flexible approach to initial screening for macromolecular crystallization. Two unique screens produced de novo have resulted from this strategy. The Pi minimal screen potentially has an ideal formulation for crystallization of novel soluble protein samples. The Pi-PEG screen is a tailor-made screen for GPCRs and potentially other membrane proteins generated by biasing the formulation towards components known to be essential.
Further screens can be formulated with the Pi Sampler on a day-to-day basis in order to test chemicals and techniques, with the aim of increasing the yield of quality crystals. Also, new crystallization techniques are constantly emerging for macromolecular targets such as membrane proteins and hence formulations with special considerations are required: one may want to formulate screens compatible with the lipidic cubic phase (LCP) concept (Landau & Rosenbusch, 1996) or make extensive use of detergents (Koszelak-Rosenblum et al., 2009).
In order for laboratories to be able to handle many Pi screen formulations and the flow of resulting data, we are working on the integration of Pi Sampler into the 'xtalPiMS' Laboratory Information Management System (LIMS; Morris et al., 2011; see http://www.pims-lims.org).
Thanks to Simon Byrne (Cambridge University Statistics Clinic; http://www.statslab.cam.ac.uk/clinic/) for discussions. Thanks to the LMB members Jan Lo ¨we, John Kendrick-Jones, Christopher Aylett, Chris Tate, Jake Grimmett and Graham Lingley for various contributions. Finally, thanks to Karen Law (MRC Technology), Chris Morris (STFC, funded by CCP4) and Tanja Hedderich (Max Planck Institute). Conflicting commercial interest: we hereby state that we have a conflicting commercial interest in that MRC Technology (http://www.mrctechnology.org/) will commercialize Pi screens under an exclusive licence to Jena Bioscience (http:// www.jenabioscience.com/).
Acta Cryst. (2011). D67, 463-470
Acta Cryst. (2011). D67, 463-470 Gorrec et al. Pi sampling 465
Acta Cryst. (2011). D67, 463-470 Gorrec et al. Pi sampling 467
Acta Cryst. (2011). D67, 463-470 Gorrec et al. Pi sampling 469
The molecular structure of the title compound, C 28 H 20 N 4 O 6 , consists of three fused six-membered rings (A,B,C) and one five-membered ring (D). The latter is linked to an isoxazole ring (E) via a methylene unit. A 4-nitro-phenyl substituent (F) is attached to the isoxazole. The fused five and six-membered rings (C,D) are almost coplanar with an r.m.s. deviation of 0.0345 A ˚and make a dihedral angle of 9.40 (8) with ring A. The isoxazole and 4-nitro-phenyl rings (E,F) are also almost coplanar with the imidazole and the fused adjacent ring (C,D), forming a dihedral angle of 11.4 (6) . The crystal packing displays intermolecular C-HÁ Á ÁO hydrogen bonding. An intramolecular C-HÁ Á ÁO interaction also occurs.
For the biological activity of anthraquinone derivatives, see: Agarwal et al. (2000); Barnard et al. (1995) (2000). For a derivative of the title compound, see: Afrakssou et al. (2010). For the use of related compounds as synthetic dyes, see: Simi et al. (1995). For puckering parameters, see : Cremer & Pople (1975).
Crystal data
Data collection: APEX2 (Bruker, 2009); cell refinement: SAINT-Plus (Bruker, 2009); data reduction: SAINT-Plus; program(s) used to solve structure: SHELXS97 (Sheldrick, 2008); program(s) used to refine structure: SHELXL97 (Sheldrick, 2008); molecular graphics: ORTEP-3 for Windows (Farrugia, 1997); software used to prepare material for publication: SHELXL97. Supplementary data and figures for this paper are available from the IUCr electronic archives (Reference: IM2283).
Recently, a number of pharmacological tests revealed that anthraquinone derivatives present various biological activities including antifungal (Agarwal et al., 2000), antimicrobial (Wu et al., 2005), anticancer (Koyamaa et al., 2002;Su et al., 2005;Chen et al., 2007), antioxidant (Yen et al., 2000;Iizuka et al., 2004), and antihuman cytomegalovirus activity (Barnard et al., 1995).
Aminoanthraquinone derivatives are a class of compounds largely used as phytotherapeutic drugs (laxatives, sedatives and antikidney and antibladder stones) and colouring agents (in the food, cosmetics and textile) (Simi et al., 1995). Anthraquinone derivatives have also been utilized for the activation of human telomerase reverse transcriptase expression (Haug et al., 2003), and they act as telomerase inhibitors or activators.
Due to their importance, in a previous study we have synthesized 1,3-diallyl-1H-anthra[1,2-d]imidazole-2,6,11(3H)-trione (Afrakssou et al., 2010). Here we have focused on the reactivity of the exocyclic C=C bond of the allyl substituents towards nitriloxides. The latter are produced as intermediates in the dehydrohalogenation of 4-nitrobenzaldoxime by a solution of sodium hypochlorite. The oxime then reacts with 1,3-diallyl-1H-anthra[1,2-d] imidazole-2,6,11(3H)-trione in a biphasic medium (water-chloroform) at 0°C during 4 h to a unique cycloadduct 3-allyl-1-[3 -(4-nitro-phenyl)-4,5-dihydro-isoxazole-5-ylmethyl] -1,3-dihydro-anthra [1,2 d] imidazole-2,6,11-trione (Scheme 1). Due to this reaction sequence a racemate of the title compound which shows a stereogenic center at C(17) is formed.
Fig. 1 shows the molecular plot of the crystal structure of the title compound. The isoxazole (E) adopts an envelope conformation on C(17) as indicated by the Cremer & Pople (1975) puckering parameters Q2 = 0.2117 (19) Å and φ2 = 141.6 (5)°. Moreover, ring (B) has a twisted conformation, with puckering parameters Q = 0.1480 (19) Å, θ = 114.6 (8) ° and φ = 139.7 (8) °, whereas all other rings are planar. The dihedral angles between fused five and six-membered rings (C,D) and the isoxazole ring (E) is 11.4 (6)°. The torsion angle between C27-C26-N2-C1 is 87.42 (0.22)°. In the crystal, adjacent molecules are linked by intermolecular C-H•••O hydrogen bonding as shown in Fig. 2 and Table 1.
To a solution of 1,3-diallyl-1H-anthra[1,2-d] imidazole-2,6,11(3H)-trione (0.3 g, 0.87 mmol) and 4-nitrobenzaldoxime (0.36 g, 2.17 mmol) in chloroform (16 ml) was added dropwise a 24% sodium hypochlorite solution (8 ml) at 273 K. Stirring was continued for 4 h. The organic layer was dried over Na 2 SO 4 and the solvent was evaporated under reduced pressure. The residue was then purified by column chromatography on silica gel using a mixture of hexane/ethyl acetate (1/1) as eluent. The title compound is formed as a racemate (Yield: 35%). Orange crystals are isolated after the solvent was allowed to evaporate.
supplementary materials sup-2 Refinement All H atoms were located in a difference map and treated as riding with C-H = 0.93 Å for all aromatic H atoms, 0.97 Å for methylene and 0.98Å for methine with U iso (H) = 1.2 U eq aromatic, methylene and methine. The C-bound H atoms of the allyl group were positioned geometrically and treated as riding with C-H = 0.93 Å (H27, H28A and H28B) and 0.97 Å (methylene) with U iso (H) = 1.2U eq .
The reflections (0 2 0) and (0 1 1) were omitted because the difference between their calculated and observed intensities are very large. They are affected by the beamstop. Fig. 1. : Molecular plot of the title compound with the atom-labelling scheme. Displacement ellipsoids are drawn at the 50% probability level. H atoms are represented as small circles. 3-Allyl-1-{[3-(4-nitrophenyl)-4,5-dihydro-1,3-oxazol-5-yl]methyl}-1H-anthra[1,2-d]imidazole-2,6,11(3H)-trione Crystal data C 28 H 20 N 4 O 6 F(000) = 1056 M r = 508.48 D x = 1.431 Mg m -3 Monoclinic, P2 1 /n Melting point: 471 K Hall symbol: -P 2yn Mo Kα radiation, λ = 0.71073 Å a = 10.0780 (3) Å Cell parameters from 7395 reflections b = 22.7094 (6) Å θ = 2.4-21.5°c = 11.2729 (3) Å µ = 0.10 mm -1 β = 113.809 (1)°T = 296 K V = 2360.41 (11) Å 3 Prism, yellow Z = 4 0.40 × 0.14 × 0.11 mm Data collection Bruker APEXII CCD diffractometer 4828 independent reflections Radiation source: fine-focus sealed tube 2998 reflections with I > 2σ(I) graphite R int = 0.048 φ and ω scans θ max = 26.4°, θ min = 2.5° supplementary materials sup-3 Absorption correction: multi-scan (SADABS, Bruker, 2009) h = -12→12 T min = 0.704, T max = 0.745 k = -28→28 47772 measured reflections l = -12→14 Refinement Refinement on F 2 Primary atom site location: structure-invariant direct methods Least-squares matrix: full Secondary atom site location: difference Fourier map R[F 2 > 2σ(F 2 )] = 0.040 Hydrogen site location: inferred from neighbouring sites wR(F 2 ) = 0.111 H-atom parameters constrained S = 1.00 w = 1/[σ 2 (F o 2 ) + (0.0455P) 2 + 0.4447P] where P = (F o 2 + 2F c 2 )/3 4828 reflections (Δ/σ) max = 0.001 343 parameters Δρ max = 0.12 e Å -3 0 restraints Δρ min = -0.15 e Å -3 Special details Geometry. All s.u.'s (except the s.u. in the dihedral angle between two l.s. planes) are estimated using the full covariance matrix. The cell s.u.'s are taken into account individually in the estimation of s.u.'s in distances, angles and torsion angles; correlations between s.u.'s in cell parameters are only used when they are defined by crystal symmetry. An approximate (isotropic) treatment of cell s.u.'s is used for estimating s.u.'s involving l.s. planes. Refinement. Refinement of F 2 against ALL reflections. The weighted R-factor wR and goodness of fit S are based on F 2 , conventional R-factors R are based on F, with F set to zero for negative F 2 . The threshold expression of F 2 > 2σ(F 2 ) is used only for calculating Rfactors(gt) etc. and is not relevant to the choice of reflections for refinement. R-factors based on F 2 are statistically about twice as large as those based on F, and R-factors based on ALL data will be even larger.
Fractional atomic coordinates and isotropic or equivalent isotropic displacement parameters (Å 2 ) x y z U iso */U eq C1 0.6321 (2) 0.82375 (9) 0.19097 (19) 0.0612 (5) C2 0.85135 (19) 0.84194 (8) 0.35084 (17) 0.0552 (4) C3 0.81023 (17) 0.89317 (8) 0.27367 (15) 0.0502 (4) C4 0.90411 (17) 0.94182 (7) 0.30538 (15) 0.0475 (4) C5 1.04177 (18) 0.93398 (7) 0.40936 (15) 0.0505 (4) C6 1.0776 (2) 0.88249 (8) 0.48157 (16) 0.0593 (5) H6 1.1683 0.8794 0.5498 0.071* C7 0.9829 (2) 0.83582 (8) 0.45530 (17) 0.0622 (5) H7 1.0064 0.8018 0.5055 0.075* C8 1.1555 (2) 0.97997 (8) 0.44241 (17) 0.0574 (4) C9 1.1231 (2) 1.03401 (8) 0.36364 (17) 0.0556 (4) C10 0.98508 (19) 1.04342 (7) 0.26723 (16) 0.0527 (4) C11 0.86591 (19) 1.00068 (8) 0.24546 (16) 0.0524 (4) C12 1.2303 (2) 1.07627 (9) 0.3858 (2) 0.0712 (5) supplementary materials sup-4 H12 1.3224 1.0703 0.4505 0.085* C13 1.2012 (3) 1.12668 (10) 0.3127 (2) 0.0793 (6) H13 1.2730 1.1550 0.3292 0.095* C14 1.0663 (3) 1.13554 (9) 0.2151 (2) 0.0754 (6) H14 1.0482 1.1692 0.1641 0.090* C15 0.9574 (2) 1.09440 (8) 0.19255 (19) 0.0651 (5) H15 0.8658 1.1008 0.1275 0.078* C16 0.59135 (18) 0.91052 (8) 0.05248 (16) 0.0576 (5) H16A 0.5982 0.9527 0.0669 0.069* H16B 0.4900 0.8995 0.0226 0.069* C17 0.64489 (18) 0.89559 (9) -0.05152 (17) 0.0588 (5) H17 0.6223 0.8546 -0.0797 0.071* C18 0.58172 (18) 0.93738 (8) -0.16582 (17) 0.0617 (5) H18A 0.5628 0.9177 -0.2474 0.074* H18B 0.4935 0.9560 -0.1695 0.074* C19 0.70447 (18) 0.98060 (8) -0.13153 (16) 0.0553 (4) C20 0.69973 (18) 1.03873 (8) -0.18912 (16) 0.0546 (4) C21 0.5711 (2) 1.06028 (9) -0.28290 (18) 0.0640 (5) H21 0.4882 1.0370 -0.3120 0.077* C22 0.5655 (2) 1.11605 (9) -0.33315 (19) 0.0701 (5) H22 0.4792 1.1305 -0.3955 0.084* C23 0.6886 (2) 1.15008 (8) -0.2904 (2) 0.0662 (5) C24 0.8184 (2) 1.12962 (10) -0.1991 (2) 0.0727 (6) H24 0.9011 1.1530 -0.1718 0.087* C25 0.8234 (2) 1.07410 (9) -0.14903 (19) 0.0655 (5) H25 0.9106 1.0599 -0.0875 0.079* C26 0.7429 (2) 0.74050 (9) 0.3414 (2) 0.0749 (6) H26A 0.6442 0.7270 0.3176 0.090* H26B 0.7923 0.7385 0.4351 0.090* C27 0.8169 (2) 0.70030 (9) 0.2829 (2) 0.0776 (6) H27 0.8222 0.6608 0.3062 0.093* C28 0.8743 (3) 0.71473 (12) 0.2031 (3) 0.0922 (7) H28A 0.8718 0.7537 0.1768 0.111* H28B 0.9182 0.6862 0.1720 0.111* N1 0.67395 (14) 0.88059 (7) 0.17487 (13) 0.0554 (4) N2 0.73969 (16) 0.80148 (7) 0.30080 (15) 0.0633 (4) N3 0.82302 (15) 0.96221 (7) -0.04143 (14) 0.0606 (4) N4 0.6801 (3) 1.20974 (9) -0.3438 (2) 0.0907 (6) O1 0.52115 (15) 0.79872 (6) 0.12021 (14) 0.0768 (4) O2 0.73990 (14) 1.01619 (6) 0.18463 (13) 0.0677 (4) O3 1.27476 (15) 0.97262 (6) 0.53085 (13) 0.0835 (4) O4 0.80057 (12) 0.90603 (6) 0.00001 (12) 0.0663 (4) O5 0.5638 (2) 1.22624 (8) -0.4262 (2) 0.1111 (6) O6 0.7864 (2) 1.24061 (9) -0.3026 (3) 0.1445 (9) Atomic displacement parameters (Å 2 ) U 11 U 22 U 33 U 12 U 13 U 23 supplementary materials sup-5 C1 0.0496 (11) 0.0657 (13) 0.0690 (12) -0.0039 (10) 0.0245 (10) -0.0025 (10) C2 0.0513 (11) 0.0581 (11) 0.0572 (10) -0.0025 (9) 0.0231 (9) -0.0016 (8) C3 0.0439 (10) 0.0582 (11) 0.0491 (9) 0.0041 (8) 0.0193 (8) -0.0003 (8) C4 0.0464 (10) 0.0516 (10) 0.0463 (9) 0.0034 (8) 0.0204 (8) -0.0033 (7) C5 0.0488 (10) 0.0559 (10) 0.0464 (9) 0.0008 (8) 0.0188 (8) -0.0058 (8) C6 0.0528 (11) 0.0668 (12) 0.0495 (9) 0.0057 (9) 0.0115 (8) 0.0018 (9) C7 0.0623 (12) 0.0614 (12) 0.0578 (11) 0.0031 (10) 0.0190 (9) 0.0077 (9) C8 0.0540 (11) 0.0654 (12) 0.0470 (9) -0.0015 (9) 0.0145 (9) -0.0099 (8) C9 0.0584 (11) 0.0547 (11) 0.0557 (10) -0.0037 (9) 0.0251 (9) -0.0144 (8) C10 0.0584 (11) 0.0489 (10) 0.0569 (10) 0.0057 (8) 0.0294 (9) -0.0085 (8) C11 0.0495 (11) 0.0581 (11) 0.0513 (9) 0.0077 (9) 0.0219 (8) -0.0055 (8) C12 0.0700 (13) 0.0670 (13) 0.0737 (13) -0.0131 (11) 0.0258 (11) -0.0153 (10) C13 0.0819 (16) 0.0621 (14) 0.1010 (17) -0.0139 (12) 0.0444 (14) -0.0140 (12) C14 0.0913 (17) 0.0502 (11) 0.0998 (16) 0.0048 (11) 0.0543 (14) 0.0004 (11) C15 0.0718 (13) 0.0559 (11) 0.0748 (12) 0.0107 (10) 0.0371 (11) -0.0013 (9) C16 0.0391 (9) 0.0674 (11) 0.0611 (11) 0.0029 (8) 0.0147 (8) 0.0008 (9) C17 0.0413 (10) 0.0693 (12) 0.0606 (10) 0.0026 (9) 0.0151 (8) -0.0027 (9) C18 0.0412 (10) 0.0799 (13) 0.0575 (10) 0.0015 (9) 0.0131 (8) -0.0025 (9) C19 0.0377 (10) 0.0742 (12) 0.0508 (9) 0.0038 (9) 0.0147 (8) -0.0049 (9) C20 0.0420 (10) 0.0692 (12) 0.0507 (9) 0.0026 (8) 0.0168 (8) -0.0098 (8) C21 0.0500 (11) 0.0681 (13) 0.0609 (11) -0.0024 (9) 0.0090 (9) -0.0084 (9) C22 0.0600 (13) 0.0716 (13) 0.0649 (12) 0.0062 (11) 0.0109 (10) -0.0085 (10) C23 0.0684 (14) 0.0582 (12) 0.0739 (12) 0.0016 (10) 0.0309 (11) -0.0113 (10) C24 0.0553 (13) 0.0743 (14) 0.0889 (14) -0.0077 (11) 0.0293 (11) -0.0138 (12) C25 0.0417 (10) 0.0818 (14) 0.0683 (12) -0.0003 (10) 0.0174 (9) -0.0063 (10) C26 0.0668 (13) 0.0692 (13) 0.0855 (14) -0.0114 (11) 0.0273 (11) 0.0125 (11) C27 0.0647 (14) 0.0634 (13) 0.0915 (16) -0.0019 (11) 0.0179 (12) 0.0047 (11) C28 0.0823 (16) 0.0928 (17) 0.1015 (18) -0.0009 (14) 0.0369 (15) -0.0112 (14) N1 0.0402 (8) 0.0646 (10) 0.0581 (8) -0.0012 (7) 0.0164 (7) 0.0008 (7) N2 0.0554 (10) 0.0604 (10) 0.0703 (10) -0.0051 (8) 0.0216 (8) 0.0070 (8) N3 0.0408 (8) 0.0796 (11) 0.0598 (9) 0.0054 (8) 0.0188 (7) 0.0010 (8) N4 0.0901 (16) 0.0661 (13) 0.1166 (17) 0.0044 (12) 0.0426 (14) -0.0096 (12) O1 0.0551 (8) 0.0801 (10) 0.0868 (9) -0.0157 (7) 0.0197 (7) -0.0064 (7) O2 0.0528 (8) 0.0662 (8) 0.0803 (9) 0.0133 (6) 0.0230 (7) 0.0037 (6) O3 0.0639 (9) 0.0907 (11) 0.0692 (8) -0.0144 (8) -0.0008 (7) 0.0015 (7) O4 0.0420 (7) 0.0861 (10) 0.0672 (8) 0.0113 (6) 0.0182 (6) 0.0102 (7) O5 0.1209 (16) 0.0787 (12) 0.1199 (14) 0.0141 (11) 0.0344 (13) 0.0101 (10) O6 0.1056 (16) 0.0775 (12) 0.233 (3) -0.0174 (11) 0.0507 (16) 0.0068 (14) Geometric parameters (Å, °) C1-O1 1.221 (2) C16-H16B 0.9700 C1-N2 1.371 (2) C17-O4 1.456 (2) C1-N1 1.392 (2) C17-C18 1.518 (2) C2-C7 1.379 (2) C17-H17 0.9800 C2-N2 1.384 (2) C18-C19 1.502 (2) C2-C3 1.411 (2) C18-H18A 0.9700 C3-C4 1.404 (2) C18-H18B 0.9700 C3-N1 1.405 (2) C19-N3 1.286 (2) supplementary materials sup-6 C4-C5 1.420 (2) C19-C20 1.463 (3) C4-C11 1.477 (2) C20-C21 1.390 (2) C5-C6 1.387 (2) C20-C25 1.396 (2) C5-C8 1.483 (2) C21-C22 1.379 (3) C6-C7 1.376 (2) C21-H21 0.9300 C6-H6 0.9300 C22-C23 1.373 (3) C7-H7 0.9300 C22-H22 0.9300 C8-O3 1.224 (2) C23-C24 1.378 (3) C8-C9 1.472 (3) C23-N4 1.471 (3) C9-C12 1.390 (3) C24-C25 1.374 (3) C9-C10 1.393 (2) C24-H24 0.9300 C10-C15 1.392 (2) C25-H25 0.9300 C10-C11 1.486 (2) C26-N2 1.455 (2) C11-O2 1.2264 (19) C26-C27 1.490 (3) C12-C13 1.371 (3) C26-H26A 0.9700 C12-H12 0.9300 C26-H26B 0.9700 C13-C14 1.375 (3) C27-C28 1.294 (3) C13-H13 0.9300 C27-H27 0.9300 C14-C15 1.384 (3) C28-H28A 0.9300 C14-H14 0.9300 C28-H28B 0.9300 C15-H15 0.9300 N3-O4 1.408 (2) C16-N1 1.460 (2) N4-O6 1.206 (3) C16-C17 1.514 (2) N4-O5 1.224 (2) C16-H16A 0.9700 O1-C1-N2 126.88 (18) O4-C17-H17 110.8 O1-C1-N1 126.32 (18) C16-C17-H17 110.8 N2-C1-N1 106.80 (16) C18-C17-H17 110.8 C7-C2-N2 128.61 (17) C19-C18-C17 99.82 (14) C7-C2-C3 123.44 (17) C19-C18-H18A 111.8 N2-C2-C3 107.93 (15) C17-C18-H18A 111.8 C4-C3-N1 134.81 (15) C19-C18-H18B 111.8 C4-C3-C2 119.47 (15) C17-C18-H18B 111.8 N1-C3-C2 105.71 (15) H18A-C18-H18B 109.5 C3-C4-C5 116.42 (15) N3-C19-C20 119.80 (16) C3-C4-C11 124.85 (15) N3-C19-C18 113.39 (16) C5-C4-C11 118.57 (15) C20-C19-C18 126.82 (15) C6-C5-C4 121.74 (16) C21-C20-C25 118.73 (18) C6-C5-C8 117.06 (16) C21-C20-C19 120.55 (17) C4-C5-C8 121.17 (15) C25-C20-C19 120.70 (16) C7-C6-C5 121.98 (16) C22-C21-C20 120.41 (18) C7-C6-H6 119.0 C22-C21-H21 119.8 C5-C6-H6 119.0 C20-C21-H21 119.8 C6-C7-C2 116.76 (17) C23-C22-C21 119.46 (19) C6-C7-H7 121.6 C23-C22-H22 120.3 C2-C7-H7 121.6 C21-C22-H22 120.3 O3-C8-C9 120.70 (17) C22-C23-C24 121.52 (19) O3-C8-C5 120.92 (17) C22-C23-N4 118.7 (2) C9-C8-C5 118.36 (16) C24-C23-N4 119.8 (2) C12-C9-C10 119.54 (18) C25-C24-C23 118.87 (19) supplementary materials sup-7 C12-C9-C8 120.01 (17) C25-C24-H24 120.6 C10-C9-C8 120.45 (16) C23-C24-H24 120.6 C15-C10-C9 119.50 (17) C24-C25-C20 120.99 (18) C15-C10-C11 119.50 (17) C24-C25-H25 119.5 C9-C10-C11 120.97 (16) C20-C25-H25 119.5 O2-C11-C4 122.39 (16) N2-C26-C27 113.34 (18) O2-C11-C10 119.36 (16) N2-C26-H26A 108.9 C4-C11-C10 118.15 (15) C27-C26-H26A 108.9 C13-C12-C9 120.4 (2) N2-C26-H26B 108.9 C13-C12-H12 119.8 C27-C26-H26B 108.9 C9-C12-H12 119.8 H26A-C26-H26B 107.7 C12-C13-C14 120.4 (2) C28-C27-C26 126.6 (2) C12-C13-H13 119.8 C28-C27-H27 116.7 C14-C13-H13 119.8 C26-C27-H27 116.7 C13-C14-C15 120.2 (2) C27-C28-H28A 120.0 C13-C14-H14 119.9 C27-C28-H28B 120.0 C15-C14-H14 119.9 H28A-C28-H28B 120.0 C14-C15-C10 120.0 (2) C1-N1-C3 109.62 (14) C14-C15-H15 120.0 C1-N1-C16 117.89 (14) C10-C15-H15 120.0 C3-N1-C16 131.04 (15) N1-C16-C17 112.46 (14) C1-N2-C2 109.85 (15) N1-C16-H16A 109.1 C1-N2-C26 122.97 (16) C17-C16-H16A 109.1 C2-N2-C26 126.32 (16) N1-C16-H16B 109.1 C19-N3-O4 109.50 (14) C17-C16-H16B 109.1 O6-N4-O5 123.0 (2) H16A-C16-H16B 107.8 O6-N4-C23 118.8 (2) O4-C17-C16 108.55 (14) O5-N4-C23 118.2 (2) O4-C17-C18 104.60 (14) N3-O4-C17 107.87 (12) C16-C17-C18 111.04 (15) Hydrogen-bond geometry (Å, °) D-H•••A D-H H•••A D•••A D-H•••A C7-H7•••O1 i 0.93 2.60 3.516 (2) 169 C18-H18B•••O2 ii 0.97 2.37 3.333 (2) 170 C26-H26B•••O1 i 0.97 2.55 3.379 (3) 144 C16-H16A•••O2 0.97 2.10 2.902 (2) 141 Symmetry codes: (i) x+1/2, -y+3/2, z+1/2; (ii) -x+1, -y+2, -z. supplementary materials sup-8 supplementary materials sup-9
The nectin family of Ca 2+ -independent immunoglobulin-like cell-cell adhesion molecules contains four members. Nectins, which have three Ig-like domains in their extracellular region, form cell-cell adherens junctions cooperatively with cadherins. The whole extracellular regions of nectin-1 (nectin-1-EC) and nectin-2 (nectin-2-EC) were expressed in Escherichia coli as inclusion bodies, solubilized in 8 M urea and then refolded by rapid dilution into refolding solution. The refolded proteins were subsequently purified by three chromatographic steps and crystallized using the hanging-drop vapour-diffusion method. The nectin-1-EC crystals belonged to space group P2 1 3 and the nectin-2-EC crystals belonged to space group P6 1 22 or P6 5 22.
Cell adhesion molecules are the primary mediators for various types of cell-cell junctions and play essential roles in various cellular processes, including morphogenesis, differentiation, proliferation and migration (Harris & Tepass, 2010;Ogita & Takai, 2006). In polarized epithelial cells, cell-cell junctions comprise several adhesive apparatuses including tight junctions (TJs), adherens junctions (AJs), desmosomes and gap junctions. AJs regulate TJ formation and the establishment of the apical-basal polarity at cell-cell adhesion sites (Takai, Ikeda et al., 2008). Nectins, which are immunoglobulin-like cell adhesion molecules, and cadherins, which are Ca 2+ -dependent cell adhesion molecules, localize at AJs and have essential cooperative roles in AJ formation. Nectins initiate AJ formation before cadherins form cell-cell adhesions. Initial cell-cell contacts are formed between two neighbouring cells by nectins, and cadherins are then recruited to the nectin-based adhesion sites to form strong cellcell adhesions. TJs are formed at the apical side of AJs. Together, nectins and cadherins mediate TJ formation by recruiting junctional adhesion molecules (JAMs), followed by claudins and occludins, to the apical side of AJs.
Nectins comprise a family of four members, nectin-1, nectin-2, nectin-3 and nectin-4, each with multiple isoforms (Takai et al., 2003;Ogita et al., 2010). Most nectin-family members are membrane glycoproteins with an extracellular N-terminal variable region-like domain, two extracellular constant region-like domains, a transmembrane region and a cytoplasmic tail. All nectins, except for nectin-1 , nectin-1 , nectin-3 and nectin-4, contain a PDZ-binding motif (E/A-X-Y-V) in their cytoplasmic tail. Through this sequence the nectins can bind afadin, which itself binds to an actin filament and -catenin.
Two nectin molecules on the surface of the same cell first form cis-dimers, which is followed by the formation of trans-dimers of the cis-dimers on apposing cells, resulting in the formation of cell-cell adhesions (Takai, Miyoshi et al., 2008). Heterophilic trans-interactions have been detected between nectin-1 and nectin-3, between nectin-2 and nectin-3 and between nectin-1 and nectin-4. In addition to their cell adhesion activity, nectin-1 and nectin-2, but not nectin-3 or nectin-4, serve as entry receptors for -herpesviruses by binding the virus-envelope glycoprotein gD with different specificities (Cocchi et al., 1998;Sakisaka et al., 2001;Lopez et al., 2000;Geraghty et al., 1998;Warner et al., 1998). While nectin-1 shows activity as a receptor for herpes simplex virus type 1 (HSV-1), HSV-2, pseudo-rabies virus (PRV) and bovine herpesvirus type 1, nectin-2 mediates the entry of HSV-2 and PRV.
To further understand how nectin-family members form cisinteractions and trans-interactions, structural information on the extracellular region of nectin-family members is fundamental. Here, we report the refolding, crystallization and preliminary X-ray crystallographic analyses of the extracellular regions of nectin-1 and nectin-2. To our knowledge, this is the first report of the crystallization of the whole extracellular regions of nectins.
Genes encoding the extracellular region of human nectin-1 (nectin-1-EC; residues 30-335) or mouse nectin-2 (nectin-2-EC; residues 32-339) were PCR-amplified using the Expand High Fidelity PCR system (Roche) from respective full-length cDNAs using primer A (5 0 -TAGGATCCGTCCCAGGCGTCCACTCC-3 0 ; nectin-1, forward) and primer B (5 0 -TAGCGGCCGCTTCTGTGATATT- GACCTCCACC-3 0 ; nectin-1, reverse) or primer C (5 0 -TAGGAT-CCGCAGGATGTGCGAGTTCAAGTGC-3 0 ; nectin-2, forward) and primer D (5 0 -TAGCGGCCGCGTCTCGCACCAGGATGAC-CT-3 0 ; nectin-2, reverse). The PCR products were inserted into the pGEM-T Easy vector (Promega), which contains 3 0 -T overhangs at the insertion site. Target genes were isolated by digestion of the plasmids with the restriction enzymes BamHI and NotI and were confirmed by DNA sequencing. This was followed by ligation into the T7 promoter expression vector pET21b (Novagen) in frame with a C-terminal 6ÂHis tag. Each recombinant plasmid was transformed into Escherichia coli BL21 (DE3) cells (Novagen).
The cells were cultured in Luria-Bertani broth containing 100 mg ml À1 ampicillin at 310 K until the OD 600 reached 0.5-0.7. Protein expression was induced at 298 K by the addition of isopropyl -d-1-thiogalactopyranoside to a final concentration of 0.4 mM. At 16 h post-induction the cells were harvested by centrifugation, suspended in phosphate-buffered saline containing 40 mg ml À1 hen egg-white lysozyme, subjected to two cycles of freezing and thawing and sonicated until the lysate was homogeneous. Centrifugation of these lysates at 10 000g for 20 min yielded inclusion bodies.
The inclusion bodies were washed four times with a buffer consisting of 20 mM Tris-HCl pH 7.5, 300 mM NaCl, 1 mM EDTA, 0.5% Triton X-100 and 1 mM DTT and were dissolved in a buffer consisting of 50 mM MES-NaOH pH 6.0, 8 M urea, 1 mM EDTA and 1 mM DTT. This mixture was slowly rotated overnight at 277 K before centrifugation at 10 000g for 20 min to remove insoluble materials.
The yields of inclusion bodies were $0.1 g for nectin-1-EC and $0.2 g for nectin-2-EC per litre of culture.
Unfolded proteins were refolded by 300-fold dilution into refolding solution A [500 mM l-arginine, 100 mM Tris-HCl pH 9.0, 2 mM oxidized glutathione (GSSG) and 1 mM reduced glutathione (GSH)] in the case of nectin-1-EC or refolding solution B (500 mM l-arginine, 100 mM Tris-HCl pH 9.0, 10 mM GSSG, 0.1 mM GSH) in the case of nectin-2-EC, followed by incubation for 48 h at 277 K. After concentration using a 10 000 molecular-weight cutoff ultrafiltration membrane (GE Healthcare), the samples were subjected to size-exclusion chromatography on a HiLoad 16/60 Superdex 200 pg column (GE Healthcare) to separate correctly folded proteins from aggregated forms. These fractions were dialyzed against 20 mM MES pH 6.0 to precipitate almost-misfolded proteins, filtered using an Ultrafree-MC GV 0.22 mm (Millipore) and applied onto a HiTrap SP HP column (5 ml; GE Healthcare) followed by a Mono Q column (1 ml; GE Healthcare). The protein yields were $0.5 mg for nectin-1-EC and $10 mg for nectin-2-EC from 100 mg inclusion bodies.
The standard conditions for refolding nectin-1-EC and nectin-2-EC were as follows: 400 mM l-arginine, 100 mM Tris-HCl pH 9.0, 2.5 mM GSH, 2.5 mM GSSG and 100 mg ml À1 unfolded protein. Small-scale refolding assays (1 ml) were performed to investigate the effects of changing the l-arginine concentration from 100 to 600 mM (for nectin-1-EC), the pH from 7.0 to 9.0 (for nectin-1-EC), the GSSG:GSH ratio from 10.0:0.1 mM to 0.1:10.0 mM (for both nectin-1-EC and nectin-2-EC) and the concentration of unfolded protein from 25 to 200 mg ml À1 (for nectin-1-EC). In each case, the unfolded protein solutions were diluted at least 300-fold into each of the refolding solutions such that only one parameter was varied while the other parameters were kept at the standard conditions.
The solutions were incubated at 277 K for 48 h and then subjected to size-exclusion chromatography on a Superdex 200 10/300 GL column using an A ¨KTA FPLC system (GE Healthcare; Figs. 1a-1d). A peak at an elution volume of $15.5 ml corresponded to the correctly folded protein containing native intramolecular disulfide bonds. Aggregates containing intermolecular disulfide bonds eluted in the void volume ($8.5 ml).
Purified nectin-1-EC and nectin-2-EC were dialyzed against solutions C (20 mM Tris-HCl pH 7.5, 150 mM NaCl) and D (20 mM Tris-HCl pH 9.0, 150 mM NaCl), respectively, and then concentrated to 5 and 4 mg ml À1 , respectively, using a Vivaspin 6 10k (GE Healthcare). The homogeneous proteins were analyzed by screening additives using dynamic light scattering with a Zetasizer Nano ZS (Malvern Instruments) to determine their suitability for crystallization. In the presence of 0.2 M NDSB201 nectin-1-EC and nectin-2-EC were monodisperse. Initial crystallization trials were performed using a Phoenix liquid-handling system (Art Robbins Instruments) at 296 K using SaltRx 1, SaltRx 2, 50%(v/v) (half concentration) PEGRx 1, 50%(v/v) PEGRx 2, 50%(v/v) PEG/Ion 1 and 50%(v/v) PEG/Ion 2. The volume of the reservoir solution was 60 ml. The drops consisted of 0.2 ml of both the protein and reservoir solution.
In the presence of 0.4 M NDSB201, nectin-1-EC and nectin-2-EC crystals were obtained using both polyethylene glycol and salt conditions as the reservoir. The initial crystallization conditions for nectin-1-EC and nectin-2-EC were further refined by changing the pH, precipitant concentration and additives. The most promising crystals of nectin-1-EC were observed in drops comprised of equal volumes of nectin-1-EC solution [20 mM Tris-HCl pH 7.5, 150 mM NaCl and 6%(w/v) 1,6-hexanediol] and precipitant solution [50 mM citric acid, 50 mM bis-Tris propane and 1-3%(v/v) PEG 3350] at 296 K (Fig. 2a). Crystals of nectin-1-EC could be obtained even if the crystallization conditions contained no NDSB201, and the NDSB201 did not influence the diffraction quality of the crystals. The most promising crystals of nectin-2-EC were observed in drops comprised of equal volumes of nectin-2-EC solution (20 mM Tris-HCl pH 9.0, 150 mM NaCl and 0.35 M NDSD201) and precipitant solution (45 mM citric acid, 55 mM bis-Tris propane and 3.6 M sodium nitrate) at 296 K (Fig. 2b).
To improve the diffraction quality of the nectin-1-EC crystals, the crystals were subjected to dehydration with increasing concentrations of PEG 300. The crystals were transferred in a large number of steps from harvesting buffer [20 mM Tris-HCl pH 7.5, 150 mM NaCl, 6%(w/v) 1,6-hexanediol, 50 mM citric acid, 50 mM bis-Tris propane and 5%(v/v) PEG 3350] to harvesting buffer including 25%(v/v) PEG 300 at 277 K. Before freezing with liquid nitrogen, the crystals were equilibrated in the final buffer for 3 d. This procedure markedly improved the resolution limit of the crystals from $5 to $2.8 A ˚. The nectin-2-EC crystals were soaked in a cryoprotection solution consisting of 20 mM Tris-HCl pH 9.0, 150 mM NaCl, 0.4 M NDSB201, 45 mM citric acid, 55 mM bis-Tris propane, 4.0 M sodium nitrate and 14%(v/v) ethylene glycol by stepwise transfer at room temperature.
Diffraction data sets were collected from nectin-1-EC and nectin-2-EC crystals on the BL44XU beamline at the SPring-8 synchrotron facility (Harima, Hyogo, Japan) at 100 K using a DIP6040 imagingplate detector (MAC Science/Bruker AXS). A total of 60 frames of data were collected for nectin-1-EC in three runs with a translation of 70 mm along the rotation axis. The nectin-1-EC data-collection parameters included a crystal-to-detector distance of 540 mm, an oscillation angle of 0.5 and an exposure time of 20 s per frame at a wavelength of 0.9000 A ˚(Fig. 3a). For nectin-2-EC, a total of 70 frames of data were collected with a crystal-to-detector distance of 400 mm, an oscillation angle of 0.5 and an exposure time of 2 s per frame at a wavelength of 0.9000 A ˚(Fig. 3b). All data sets were processed and scaled with the HKL-2000 program package (Otwinowski & Minor, 1997). Data-collection statistics are summarized in Table 1.
The extracellular regions of nectin-1 and nectin-2 fused with a C-terminal 6ÂHis tag were expressed as inclusion bodies in E. coli BL21 (DE3). After solubilizing the inclusion bodies in 8 M urea, nectin-1-EC and nectin-2-EC proteins were successfully refolded by rapid dilution with a glutathione redox couple. To increase the yields of refolded nectin-1-EC protein, the refolding conditions (i.e. pH, GSSG:GSH ratio, l-arginine concentration and nectin-1-EC concentration) were optimized in a series of small reactions (1 ml). Correct folding was assessed by size-exclusion chromatography on a Superdex 200 10/300 GL column using an A ¨KTA FPLC system (GE Healthcare; Figs. 1a-1d). Notably, a lack of the C-terminal 6ÂHis tag significantly decreased the yield of correctly folded protein; we could not obtain crystals of nectin-1-EC without the tag. Based on the results for nectin-1-EC, only the GSSG:GSH ratio was optimized for nectin-2-EC (Fig. 2e). The optimized refolding solutions for nectin-1-EC and nectin-2-EC are described in x2. Nectin-1-EC crystals belonged to the cubic space group P2 1 3, with unit-cell parameters a = b = c = 164.9 A ˚. Nectin-2-EC crystals belonged to the hexagonal space group P6 1 22 or P6 5 22, with unit-cell parameters a = b = 79.3, c = 235.4 A ˚. Molecular-replacement calculations with MOLREP and Phaser (Collaborative Computational Project, Number 4, 1994) using the structure of a homologous protein [CD155, which has the maximum sequence identity to nectin-1-EC (48.4%) and nectin-2-EC (25.3%); PDB code 3eow; Zhang et al., 2008] as a search model were unsuccessful. Therefore, heavy-atom derivatives of nectin-1-EC and nectin-2-EC crystals have been prepared for phase determination. Diffraction data sets for the heavy-atom derivatives are currently being collected on BL44XU at the SPring-8 synchrotron facility for phase determination.
hkl P i jI i ðhklÞ À hIðhklÞij= P hkl P i I i ðhklÞ, where hI(hkl)i is the mean intensity of symmetry-equivalent reflections.
Acta Cryst.(2011). F67, 344-348
We thank
The objective of this study is to examine the efficacy and tolerability of miglitol with respect to improving glycemic control in Chinese patients with type 2 diabetes mellitus inadequately controlled by diet and sulfonylurea treatment. This was a randomized, double-blinded, placebo-controlled, multicenter study. A total of 105 patients were randomized to receive 24 weeks of treatment with miglitol (n = 52; titrated from 50 mg to 100 mg 3 times daily) or placebo (n = 53). Concomitant sulfonylurea treatment and diet remained unchanged. The primary endpoint was change in glycated hemoglobin (HbA1c) from baseline at 24 weeks. Secondary endpoints were changes in fasting plasma glucose (FPG), postprandial plasma glucose (PPG), and postprandial serum insulin (PSI). The miglitol treatment group showed significantly greater reductions in HbA1c and PPG levels compared with the placebo group. With respect to adverse events, abdominal discomfort, diarrhea, and hypoglycemia occurred with similar frequency in both groups. Results of this study indicate that miglitol significantly improves metabolic control in Chinese patients with type 2 diabetes mellitus. Miglitol is safe and well tolerated, with the exception of abdominal discomfort. Therefore, miglitol may be a useful adjuvant therapy for Chinese patients with type 2 diabetes mellitus inadequately controlled by diet and sulfonylurea treatment.
Diabetes mellitus, a rapidly growing health problem in many countries, is an important cause of morbidity and mortality. According to the US Centers for Disease Control and Prevention, approximately 14.7 million people in the United States were diagnosed with diabetes as of 2004, with type 2 diabetes accounting for approximately 90% of those cases [1]. Type 2 diabetes has resulted in an extremely large and growing economic burden. Despite the availability of effective diabetes-specific therapies, achievement of glycemic goals by patients is far from adequate in the United States. Less than half of adults with diabetes are reported to attain a glycated hemoglobin (HbA1c) level of \7% [2]. From 1960 to 1988 in Taiwan, mortality ascribed to diabetes increased 6.3-fold [3]. In another study from Taiwan, with a total of 1,124,348.4 person-years of follow-up, 43,888 patients with diabetes died, and the crude mortality rate was 39.0/1,000 person-years [4].
Maintaining a normal plasma glucose level is key for reducing the risk of developing complications of diabetes [5]. Current recommendations from the UK Prospective Diabetes Study Group emphasize lifestyle management, diet, and exercise as the first-line approach, followed by therapy with oral antidiabetic drugs, administered alone or in combination [5]. Recent reviews of the literature confirm the salutary effects of exercise in individuals with type 2 diabetes [6,7]. Other standard therapies include the use of peroxisome proliferator-activated receptor (PPAR)-gamma agonists (thiazolidinediones), which decrease levels of glycated hemaglobin, fasting plasma glucose, insulin, and free fatty acids in patients with type 2 diabetes [8,9], and insulin, together with metformin, which acts predominantly to inhibit hepatic glucose production, or with insulin secretagogues [10,11]. Antidiabetic drug monotherapy eventually necessitates the use of increasing dosage and/or a second antidiabetic medication because type 2 diabetes worsens over time as a result of declining pancreatic b-cell function [12].
The class of a-glucosidase inhibitors has a unique mode of action. These drugs block oligosaccharide catabolism, delay carbohydrate digestion and absorption, and smooth and lower postprandial plasma glucose (PPG) peaks [13,14]. Miglitol is the first pseudomonosaccharide a-glucosidase inhibitor derived from 1-deoxynojirimycin and is structurally a glucose analogue [15]. Its efficacy, in monotherapy [16] or in combination with sulfonylureas [17], as a glucose-lowering agent in Chinese patients with type 2 diabetes has not been determined in clinical studies. The aim of this study was to investigate the efficacy and tolerability of miglitol in combination with sulfonylureas, compared to sulfonylurea monotherapy, for the improvement of glycemic control in Chinese outpatients with type 2 diabetes mellitus inadequately controlled by diet and sulfonylurea treatment.
The study design was a randomized, double-blinded, placebo-controlled, multicenter comparison of miglitol treatment compared with placebo administration over a 24-week period. Patients with a confirmed diagnosis of type 2 diabetes mellitus whose previous treatment with diet and sulfonylureas had proved inadequate according to medical chart monitoring were recruited at 4 medical centers in Taiwan. Inclusion criteria included age [20 years; fasting plasma glucose (FPG) concentration of 100 mg/dL to 240 mg/dL; HbA1c value of 6.5% (based on the glycemic goal for adults of B6.5% as specified by the Diabetes Association of Taiwan) to 10.0%; history of uncontrolled type 2 diabetes mellitus despite prior nutrition therapy; and stable dosing with a sulfonylurea for at least 8 weeks before randomization.
Exclusion criteria included the following: suggested diagnosis of type 1 diabetes mellitus; active insulin therapy, known lactose intolerance, or treatment with medication that significantly alters gastrointestinal motility and/ or absorption; concomitant glucocorticoid therapy, other medication affecting glucose homeostasis, or treatment with investigational drugs; serum transaminase level [2.5 times the upper normal limit or serum creatinine level [1.5 mg/dL; presence of significant disease or condition (including emotional disorder or substance abuse) that would likely alter the course of diabetes or the patient's ability to complete the study; documented gastrointestinal disease associated with marked disorder of digestion or absorption, or condition that may worsen as a result of increased gas formation in the intestine; pregnant or lactating women or women of childbearing age without a medically approved method of contraception; and patients participating in another clinical trial within 90 days of screening.
This study was conducted in accordance with the European Community guidelines for Good Clinical Practice and the Declaration of Helsinki and its amendments. The protocol was approved by the corresponding Joint Institutional Review Board or the Department of Health and Ethics Committee of each investigational site. Written informed consent was obtained from each patient.
Patients with uncontrolled diabetes despite nutrition therapy and sulfonylurea treatment were assigned to a 2-week baseline work-up and dietary run-in period. During this period, demographic data were obtained, and vital signs and routine laboratory variables were measured. To confirm adherence to nutrition therapy, patients were instructed to complete a 3-day diet record before visit 2. Eligible patients were randomized to receive miglitol (Migbose; Standard Chem. & Pharm. Co., Ltd.) 50 mg 3 times daily for 12 weeks, titrated to 100 mg 3 times daily for 12 weeks, or placebo. After randomization, patients were instructed to complete a 3-day diet record before each visit.
Patients were asked to adhere to a dietary plan tailored to their energy requirements and metabolic control, according to current American Diabetes Association recommendations: carbohydrates up to 60%, fat \30%, and protein 12-20%. Concomitant sulfonylurea treatment remained unchanged throughout the study.
Patients were instructed to take 1 miglitol or placebo tablet with the first mouthful of food at each of 3 main daily meals. Drug compliance was determined by tablet count at each visit. After randomization (week 0), patients were assessed at weeks 4, 8, 12, 16, 20, and 24. A physical examination, assessment of adverse events, dietary counseling, and measurement of HbA1c were carried out at each visit. All secondary efficacy variables and routine laboratory variables were measured at baseline obtained before week 0 during the run-in period and at week 24. In cases of premature termination, routine laboratory variables were measured at the last visit. Dose titration at week 12 was performed at the discretion of the investigator. Patients with good tolerance to miglitol 50 mg 3 times daily were titrated to 100 mg 3 times daily. Patients with unsatisfactory, but acceptable, tolerance to the treatment were maintained at 50 mg 3 times daily. Patients unable to tolerate the treatment were discontinued from the study. Patients reported to the study station between 08:00 and 08:30 AM after a 12-h fast. After emptying the bladder, body height and weight were measured, and body mass index (BMI) was calculated as body weight (kg) divided by height squared (m 2 ). An electrocardiogram was performed for all patients to evaluate cardiac function. The antecubital vein of the arm was cannulated for blood sampling. Baseline or fasting blood samples were obtained after approximately 10 min of rest after cannula placement.
Venous blood samples were placed into individual tubes with ethylenediaminetetraacetic acid. Aliquots of serum and plasma were stored at -80°C. Samples from each patient were measured in the same assay to reduce interassay variation. Hematology, biochemical assays (sodium, potassium, serum creatinine, alkaline phosphatase, alanine aminotransferase [ALT], aspartate aminotransferase [AST], total protein), and lipid assays (triglyceride, total cholesterol, low-density lipoprotein cholesterol [LDL-C], highdensity lipoprotein cholesterol [HDL-C]) were carried out by routine automated methods. Plasma glucose was detected by the glucose oxidase method with a 2300 STAT glucose analyzer (Yellow Springs Instrument Inc., Yellow Springs, OH). Serum insulin was determined by microparticle enzyme immunoassay with an AxSYM system (Abbott Laboratories, Abbott Park, IL). Measurement of HbA1c was performed with a DCA 2000 analyzer (Bayer Diagnostics, Elkhart, IN).
Intent-to-treat (ITT) analysis was performed for assessment of efficacy. Patients were included in the ITT analysis if they had efficacy data at baseline (week 0) and at least 1 postbaseline efficacy measurement. The primary efficacy variable was HbA1c concentration. The endpoint was defined as the last available measurement. Secondary efficacy variables included FPG, PPG, and postprandial serum insulin (PSI). Venous blood for postprandial measurements was taken 2 h after a standard breakfast.
Patients were included in the safety analysis if they had taken at least 1 dose of medication and had at least 1 postbaseline safety measurement. The baseline for safety analysis was defined as measurements taken at visit 1 (week 2), and the endpoint was defined as the last measurements taken at visit 8 (week 24). Safety variables were analyzed descriptively. All adverse events were defined according to the Coding Symbols for a Thesaurus of Adverse Reaction Terms (COSTART) glossary and body system categories (http://hedwig.mgh.harvard.edu/biostatistics/files/costart. html).
Safety and tolerance were assessed primarily from spontaneously reported adverse events and described by the patients at each visit, with special attention to the severity of hypoglycemia; occurrences of symptoms suggestive of hypoglycemia were recorded in the patient's diary. These symptoms were rated as grade 1 (mild and transient), grade 2 (transient inability to pursue usual activities), grade 3 (need for external assistance), or grade 4 (need for medical assistance). Hypoglycemia was defined as at least 1 episode of symptoms suggestive of hypoglycemia during the study period. Other adverse events were also recorded in the patient's diary. Serious adverse events were defined as events resulting in persistent or significant disability or incapacity, new hospitalization or prolongation of current hospitalization, severe hypoglycemia, and life-threatening events or death. Acute intoxication, important medical events, and pregnancy were considered serious adverse events.
Data are shown as mean ± standard deviation (SD) for continuous variables and as n (%) for categorical variables for demographics and adverse events follow-up. Data are shown as mean ± standard error for change from baseline for primary efficacy endpoints. The last observation carried forward (LOCF) approach was used for evaluation of primary efficacy endpoints in the ITT population. For comparison of baseline demographics and change in efficacy endpoints from baseline, a 2-sample t test was performed for continuous variables, and chi-square or Fisher exact test was performed for categorical variables. Nonparametric Wilcoxon rank-sum test was also performed if the continuous data were not normally distributed. Data analysis was performed with SAS version 9.0 (SAS Institute Inc., Cary, NC). Differences were considered statistically significant at P \ 0.05.
Patient disposition is detailed in Fig. 1. A total of 138 patients with type 2 diabetes mellitus inadequately controlled by diet and sulfonylurea treatment were screened. A total of 105 patients were eligible for randomization; 52 were assigned to receive miglitol treatment, and 53 were assigned to receive placebo. Efficacy endpoints were analyzed in the ITT population, regardless of protocol compliance, and adverse events were followed-up in the safety population. Of the 105 patients, 100 (49 in the miglitol group and 51 in the placebo group) comprised the ITT population; 5 patients failed to return for postclinical assessment. All 105 randomized patients received treatment and were followed up for safety.
Baseline demographic and clinical characteristics of patients in the ITT population, listed by treatment group, are presented in Table 1. There were no significant differences in demographic or other baseline characteristics between the miglitol and placebo groups. Table 2 shows results for changes from baseline for the efficacy variables HbA1c, FPG, PPG, and PSI at week 24 in the ITT population. The change in HbA1c from baseline for the miglitol group was -0.85% ± 0.12% compared to -0.19% ± 0.11% for the placebo group (P \ 0.001). There was also a significant difference in the change in PPG between groups (P \ 0.001). No significant difference in change in FPG (P = 0.052) or PSI (P = 0.364) was found between groups. The change in ALT from baseline for the miglitol group was 8.40 ± 7.20 U/L compared to 2.29 ± 6.66 U/L for the placebo group (P = 0.009). For both groups, findings for ALT were similar between baseline and week 24 (visit 8) in terms of median, SD, and range; however, the mean value was significantly increased in the miglitol group, owing to a single patient with underlying chronic hepatitis and fatty liver.
Glycemic control in the ITT population, as measured by HbA1c level, showed significant improvement in both the miglitol and placebo groups compared to baseline (week 0) after 12 weeks (P \ 0.01), 16 weeks (P \ 0.01), 20 weeks (P \ 0.001), and 24 weeks (P \ 0.001) (Fig. 2). In addition, the decrease from baseline to 24 weeks in HbA1c was significantly higher in the miglitol group than in the placebo group (P \ 0.001).
Among the 105 patients, 49 (94.2%) in the miglitol group and 42 (79.3%) in the placebo group experienced at least 1 adverse event during the study period. A total of 59 and 39 adverse events occurred in the miglitol and placebo groups, respectively. Table 3 shows the most frequent adverse events, which included abdominal discomfort, diarrhea, hypoglycemia, and other. Patients in the miglitol group reported other adverse events significantly more often than did those in the placebo group (P = 0.036). No major episodes requiring external assistance were reported. No clinically significant changes in any of the hematologic or clinical biochemistry variables were identified, and all changes were within normal range, with the exception that 4 patients (2 in the miglitol group, and 2 in the placebo group) had elevated liver enzymes at the end of the study. However, these same patients already had abnormal liver enzyme levels at baseline. There were no significant differences between treatment groups with respect to vital signs or results of physical examination, urinalysis, or electrocardiography.
The objective of the present study was to investigate the efficacy and tolerability of miglitol in combination with sulfonylureas, compared to sulfonylurea monotherapy, for the improvement of glycemic control in Chinese outpatients with type 2 diabetes mellitus inadequately controlled Results showed a greater than fourfold difference between the change in HbA1c from baseline in the miglitol group compared with the placebo group, with the difference between the 2 groups reaching statistical significance at week 12 (P \ 0.01). Clinically significant effects of miglitol treatment were found from weeks 12 to 24. This result corresponded well with decreases in HbA1c reported in another study of miglitol adjuvant therapy [18]. The HbA1c concentration reflects long-term glycemic control [19], and a decrease in HbA1c results in a reduced risk of microvascular adverse events [5]. A major problem in the use of oral antidiabetic drugs, such as sulfonylureas, is the gradual decrease in the ability of these drugs to satisfactorily control blood glucose level [20]. Pharmacologic agents are available that modify primarily the PPG level to reduce serum HbA1c [21,22]. a-Glucosidase inhibitors produce an antihyperglycemic effect and do not induce weight gain, a common problem encountered with sulfonylureas and insulin [23]. A recent meta-analysis of 41 randomized trials examined the efficacy of a-glucosidase inhibitors in patients with type 2 diabetes and showed no evidence of a beneficial effect on morbidity or mortality [24]. However, statistically significant effects on HbA1c (by acarbose), FPG (by miglitol), postload glucose, insulin level, and BMI (by acarbose) were found. The study found no effects on PSI or lipids and only minor effects on body weight.
The clinical significance of the regulation of postprandial hyperglycemia in reducing the risk of microvascular and macrovascular complications has been established in several epidemiologic studies [25][26][27]. Results of the present study are of interest because the PPG level was significantly decreased in the miglitol group compared with the placebo group (P \ 0.001). With respect to lipid variables, we found no significant differences between groups in the present study, similar to previous results [24].
The major adverse effects of a-glucosidase inhibitors, such as acarbose, include gastrointestinal symptoms; these arise mainly from the fermentation of undigested carbohydrates by colonic bacteria [28]. In contrast, miglitol is absorbed systemically, but is not metabolized, and is excreted into the urine within a relatively short period of time. Consequently, systemic adverse effects are not anticipated. Indeed, no systemic adverse effects occurred in the present study. There was also no incidence of severe hypoglycemia. The most common complaints in this study were mild hypoglycemia, diarrhea, and abdominal discomfort, resulting in premature termination of miglitol treatment by 4 patients. Approximately half of the patients in the miglitol group and 15.7% of the patients in the placebo group experienced these symptoms.
Several case reports from Europe and Japan have indicated that acarbose, a commonly prescribed a-glucosidase inhibitor, can result in severe, but reversible, hepatotoxicity, as indicated by markedly increased levels of AST and ALT [29][30][31]. In contrast, treatment with miglitol at the dosages used in the present study increased AST and ALT to a much lesser extent (B1.8 times to upper limit of normal) [17]. Because of the relatively high prevalence of hepatitis B and hepatitis C in Asian countries, the use of miglitol may prove to be a better choice compared to acarbose with respect to minimizing hepatic effects.
The small number of cases and the relatively short, 24-week study period are potential limitations of the present study. Future large-scale studies are needed to assess the long-term cardiovascular and glucose-control effects of miglitol.
Results of the present study indicate that miglitol improved PPG level and glycemic control, as reflected by decreased HbA1c concentration in Chinese patients with type 2 diabetes mellitus. Miglitol was well tolerated, with no unusual changes in safety profiles, with the exception of abdominal discomfort during the 24-week treatment period. Therefore, miglitol may be a useful adjuvant therapy for Chinese patients with type 2 diabetes inadequately controlled by diet and sulfonylurea treatment.
Changes in efficacy endpoints from baseline by study group in the intention-to-treat population (n = 100) Characteristics a, b Miglitol (n = 49) Placebo (n = 51) P value c HbA1c (%) -0.85 ± 0.12 -0.19 ± 0.11 \0.001 d FPG (mg/dL) -13.44 ± 5.48 -0.20 ± 3.91 0.052 PPG (mg/dL) -44.8 ± 10.43 14.07 ± 9.47 \0.001 d PSI (lU/mL) -4.53 ± 3.91 -3.78 ± 5.34 0.364 AST (SGOT, U/L) 1.96 ± 3.66 -1.5 ± 1.a Data were presented as mean ± standard error (SE) b HbA1c Glycated hemoglobin, FPG fasting plasma glucose, PPG postprandial plasma glucose, PSI postprandial serum insulin, AST aspartate aminotransferase, ALT alanine aminotransferase, HDL highdensity lipoprotein cholesterol, LDL low-density lipoprotein cholesterol, TG triglyceride c
This study was supported in part by
The authors declare no financial interests.
Localization of alkaline phosphatase (ALP) and cathepsin D (CAPD) in primary cultures of fetal rat hepatocytes was examined using double immunofluorescent staining in order to investigate the relationship between lysosome movement and the fate of ALP during cell restoration after microtubule disruption by colchicine. At 3 hr and 24 hr after colchicine treatment, numerous coarse dots containing ALP were observed throughout the cytoplasm, and some of these showed colocalization with CAPD. At 48 hr and 72 hr after colchicine treatment, although most of the dots containing ALP in the cytoplasm disappeared, dots containing CAPD remained. The present results suggest that the denatured ALP proteins remaining in the cytoplasm of hepatocytes during cell restoration after colchicine treatment are digested by lysosomes.
Alkaline phosphatase (ALP) is predominantly localized in the bile canalicular membrane in adult rat hepatocytes and microtubules are involved in the transport of ALP to the bile canalicular membrane [1,2,4]. We previously reported that, in primary cultures of fetal rat hepatocytes, colchicine inhibits ALP transportation to the plasma membrane of bile canaliculus-like structures, and, consequently, numerous coarse dots containing ALP appear in the cytoplasm [6]. In addition, electron microscopy has revealed that autophagolysosome-like granules containing ALP are present in the cytoplasm of colchicine-treated rat hepatocytes [1,4]. However, the relationship between lysosome movement and the fate of the ALP-containing dots in the cytoplasm remains to be fully elucidated. In the present study, the localization of ALP and the lysosome marker cathepsin D (CAPD) during cell restoration, after colchicine treatment in primary cultures of fetal rat hepatocytes, was examined in order to elucidate the relationship between lysosome movement and the fate of the ALP-containing dots in the cytoplasm.
Hepatocytes were cultured as described previously [6]. Three days after the start of culture, 10 -5 M colchicine was added to the medium and hepatocytes were incubated for 1 hr. Medium was then replaced with fresh normal medium and cells were further cultured for 3, 24, 48 or 72 hr. After the termination of culture, cells were fixed for 5 min at room temperature (RT) in 4% paraformaldehyde in 0.1 M phosphate buffer, pH 7.4, followed by absolute methanol for 5 min at -20°C. Cells were washed with 0.01 M phosphate buffer, pH 7.2, containing 0.85% NaCl and 0.05% saponin (PBSS) at 4°C overnight and immersed for 5 min at RT in 0.1% Triton X-100 solution in PBSS. After washing with PBSS, cells were incubated for 1 hr at RT in 1:50 anti-rat β-tubulin monoclonal antibody (Chemicon International, Temecula, CA, USA) in PBSS. Cells were washed with PBSS and reacted for 30 min at RT with 1:150 fluorescein isothiocyanate (FITC)-labeled anti-mouse IgG antibodies (Medical and Biological Laboratories, Nagoya, Japan) in PBSS.
After termination of culture, some cells were fixed in the same manner using Zamboni fixative solution in place of 4% paraformaldehyde. After washing with PBSS and Triton X-100 treatment, cells were incubated for 1 hr at RT in mixed solution containing 1:50 anti-ALP rabbit serum [3] and 1:100 anti-CAPD goat antibody (Santa Cruz Biotechnology, Inc., Santa Cruz, CA, USA) in PBSS. Cells were then washed with PBSS and reacted for 30 min at RT with 1:50 rhodamine-labeled anti-rabbit IgG antibody (Medical and Biological Laboratories) and 1:150 FITClabeled anti-goat IgG antibody (Medical and Biological Laboratories) in PBSS.
After immunostaining, all cell nuclei were stained with 4',6-diamidino-2-phenyl-indole. Samples were examined under a fluorescence microscope (ECLIPSE E-600; Nikon, Tokyo, Japan) and photographed using a digital camera (DS-L2; Nikon) equipped with a fluorescein figure analysis system (LuminaVision; Mitani Co., Tokyo, Japan).
The present study was approved by the Ethics Committee for Animal Experiments of Kitasato University.
In normal hepatocytes, microtubules were observed to distribute radially from the cytoplasm around the nuclei to the peripheral cytoplasm (Fig. 1A), and ALP was localized in the plasma membrane along bile canaliculus-like cell borders (Fig. 1B). We previously demonstrated that the long stretches of plasma membrane exhibiting ALP localization between cell borders of normal hepatocytes are those surrounding the bile canaliculus-like structure, as the bile canaliculus marker occludin is colocalized in these long stretches of plasma membrane [5]. On the other hand, CAPD was observed in small scattered dots in the cytoplasm (Fig. 1B). Three hours after colchicine treatment, the microtubule structures in hepatocytes were destroyed, and numerous dots or fine nets showing specific fluorescence were observed throughout the cytoplasm (Fig. 2A).
Microtubular structures are known to reappear by wash-out with normal medium after treatment with antimicrotubular agents. It has been reported that, in Madin-Darby canine kidney (MDCK) cells treated with colcemid for 4 hr and cultured in normal medium for 2 hr, microtubule networks were reconstructed and reverted to normal shape [8]. On the other hand, we found that, in McA-RH 7777 cells incubated for 8 hr in basal medium after 4-hr colchicine treatment, microtubular structures did not reappear in the cytoplasm, as observed in the present study [9]. The difference in these results may be the result of differences in anti-microtubular agent or cell type.
On double staining for ALP and CAPD, at 3 hr after colchicine treatment, ALP was observed along cell borders between adjacent hepatocytes and in numerous coarse dots within the cytoplasm, and CAPD was colocalized in some of these coarse dots (Fig. 2B). These dots, in which ALP and CARD were colocalized, appear to be secondary lysosomes showing fusion of autophagosomes containing ALP and granular primary lysosomes containing CAPD and other proteinases. ALP in autophagosomes may be digested by these lysosomal enzymes. At 24 hr after colchicine treatment, microtubule structures showed a random distribution in the cytoplasm, but coarse dots showing positive reactions for ALP and autophagolysosome-like dots containing ALP and CAPD remained in the cytoplasm (data not shown). This indicates that transportation of ALP to the plasma membrane is not fully restored at 24 hr after colchicine treatment.
At 48 hr and 72 hr after colchicine treatment, abundant microtubule structures were distributed from the nuclei to the peripheral cytoplasm (Fig. 3A). ALP was localized in the plasma membrane of bile canaliculus-like structures but was scarcely seen in the cytoplasm, while CAPD was localized in dots of various sizes in the cytoplasm (Fig. 3B). In the same culture system, we observed small granules containing early endosomal antigen 1 (EEA1) along the plasma membrane showing positive reactions for ALP at 3 hr and 24 hr after colchicine treatment, but these granules were located in the peripheral cytoplasm in normal hepatocytes or hepatocytes at 48 hr and 72 hr after colchicine treatment (unpublished data). This suggests that ALP may be transported to the plasma membrane of bile canaliculuslike structures from the cytoplasm via endosomes during cell restoration after colchicine treatment. Accordingly, denatured ALP enzyme proteins in the cytoplasm may be taken up by autophagosomes and digested by lysosomes.
It has been reported that the apical membrane protein B10 in rat hepatocytes is localized in numerous vesicles in the cytoplasm when hepatocytes are cultured in the presence of colchicine for 24 hr, and that it is mainly localized in vacuoles resembling lysosomal structures when hepatocytes are cultured in the presence of nocodazole for 48 hr [7]. After microtubule disruption, bile canalicular membrane proteins such as ALP and B10 may be taken up by autophagosomes and endosomes and digested by lysosomes or transported to the plasma membrane via endosomes during cell restoration.
The deformation-induced nanostructure developed during high-pressure torsion of B2 long-range ordered FeAl is shown to be unstable upon heating. The structural changes were analyzed using transmission electron microscopy, differential scanning calorimetry and microhardness measurements. Heating up to 220 °C leads to the recurrence of the chemical long-range order that is destroyed during deformation. It is shown that the transition to the long-range-ordered phase evolves in the form of small ordered domains homogeneously distributed inside the nanosized grains. At temperatures between 220 and 370 °C recovery of dislocations and antiphase boundary faults cause a reduction in the grain size from 77 to 35 nm. Grain growth occurs at temperatures above 370 °C. The evolution of the strength monitored by microhardness is discussed in the framework of grainsize hardening and hardening by defect recovery.
Nanocrystalline materials containing a large volume fraction of grain boundaries are of great interest as they frequently exhibit improved mechanical and new physical properties [1,2]. One widely used approach to produce nanocrystalline (NC) structures is severe plastic deformation (SPD) of coarse-grained materials, as achieved, for example, by high-pressure torsion (HPT) of bulk materials [3]. To understand the properties of NC materials, their physics and thermodynamics have to be studied. A detailed knowledge of the processes occurring during the thermal treatment is of prime importance not only for applications but also for a deeper understanding of the stability of the deformation-induced metastable phases. It has been shown that multiple changes in structure occur during annealing of SPD-processed nanocrystalline metals and alloys since they contain, in addition to small grains, a high dislocation density and high internal strains [4][5][6][7].
For intermetallic alloys the formation of the nanocrystalline structure during ball milling is accompanied by loss of the long-range order (LRO) present in the initial coarse-grained material [8]. For B2-ordered FeAl, the destruction of LRO also induces a transition from the paramagnetic to the ferromagnetic state [9][10][11][12]. Therefore, modifications during annealing of nanocrystalline disordered FeAl are manifold. The modifications at low annealing temperatures (below 250 °C) have been studied by several authors using different integral methods, like differential scanning calorimetry (DSC), X-ray and neutron diffraction, as well as magnetometer measurements [13,14,9,15]. The corresponding processes causing structural modifications are very sensitive to impurities. Consequently, in the studies of ball-milled FeAl powders. different behavior during annealing was revealed that can be attributed to contamination occurring during milling. For instance, mechanically milled FeAl powders annealed for 1 h exhibit a continuous decrease in microhardness [15] or a peak at 500 °C [16]; the latter was attributed to the precipitation and growth of fine oxide particles. In order to eliminate the effect of contamination on the processes causing structural modification, SPD of bulk materials has to be applied. To date, there have been few studies of intermetallic FeAl alloys deformed severely in the bulk due to their usually inherent brittleness. In addition, the structural state of nanocrystalline FeAl as a function of temperature has not been studied in detail using transmission electron microscopy (TEM), nor has a correlation with the DSC signal over the whole interesting temperature range been established. A few TEM studies have been conducted of disordered FeAl, though only for milled powders after compaction (e.g. [17]).
Recently, we have successfully achieved the production of bulk nanocrystalline disordered FeAl of high purity by high-pressure torsion of B2-ordered FeAl. Therefore, it was the aim of this paper to investigate the temperature-dependent structural modifications of the SPDinduced metastable state using integral methods, like DSC measurements and microhardness testing, in correlation with local systematic TEM studies.
Fe-45 at.% Al single crystals were grown from high-purity Fe (99.99%) and Al (99.9997%) under argon in alumina crucibles using the Bridgman technique at a growth rate of about 10 mm h -1 followed by an annealing treatment for 1 week at 400 °C. This treatment was used to achieve a defined initial state of order and of vacancy concentration [18].
HPT samples (8 mm in diameter, 0.8 mm thick) were cut from a single crystal by spark erosion. Several samples were HPT deformed by up to three rotations under a pressure of 8 GPa to achieve deformation grades larger than 10,000%. The deformation was done at room temperature, which corresponds to a temperature of 0.18 T m (T m being the melting temperature). For DSC and subsequent TEM investigations, discs of 2.3 mm diameter were prepared from the outer rim of the HPT samples using spark erosion.
DSC studies of the nanocrystalline samples were carried out using a Netsch DSC 204 Phoenix device in aluminum crucibles under argon flux at a heating rate of 20 K min -1 and the samples were heated up to 500 °C. Each sample was subjected to two subsequent heating runs and the second one was used as baseline.
For a systematic study of the evolution of microhardness and grain size, as well as the state of order, additional samples were heated in the DSC device to 130, 170, 220, 370 and 500 °C (corresponding to homologous temperatures of 0.24, 0.26, 0.29, 0.38 and 0.46T m ) at a heating rate of 20 K min -1 followed by an immediate cooling process at a cooling rate of 20 K min -1 . Measurements of the microhardness were carried out at room temperature using the Vickers technique with a Paar MHT-4 indentor. Indentation was done at a gradient of 0.1 N s -1 , with a final force of 2 N being applied for 10 s. Subsequently, the imprints were measured by digital imaging techniques after recording with a Zeiss Axioplan Optical microscope equipped with a CCD camera.
TEM samples were prepared by twin-jet electropolishing in a solution of methanol with 33% nitric acid at -25 °C [19]. TEM studies were carried out using a Phillips CM200 operating at an acceleration voltage of 200 kV.
Fig. 1 shows the signal obtained from the DSC measurements containing three exothermic peaks. The onset of the pronounced first peak (I) was measured by putting a tangent at the slope of the peak. The first exothermic peak (I) has an onset at about 130 °C, the end is at about 220 °C and the center at about 170 °C; from the area an enthalpy change of about 54 J g -1 (4.5 kJ mol -1 ) was deduced. The other two exothermic peaks (II and III) are centered around 320 and 410 °C, respectively. They are strongly overlapping and too small for a proper analysis of their areas, onset-and endpoints.
To identify the processes causing these exothermic peaks, individual samples were annealed to selected temperatures between the peaks (cf. the temperatures marked by the crosses in Fig. 1). The samples were then studied by TEM, always taking bright-field and dark-field images in combination with diffraction patterns. Fig. 2 shows the TEM images and the corresponding selected area diffraction (SAD) patterns obtained from a sample of the as-deformed state and from samples heated to 170, 220, 370 and 500 °C. In all cases the same size of SAD aperture (1.2 μm) was used. The dark-field image of the as-deformed sample shows a bright area, which reveals a grain. To measure the grain size by TEM methods, special care is needed to identify the large-angle (>15°) grain boundaries since the contrast caused by dislocation networks and subgrain boundaries can be complex. In addition, in the case of nanograins, their overlap is frequent even in TEM foils, and this leads to the formation of moiré contrast fringes [20]. Therefore, to get an unambiguous identification of the grains, it is necessary to tilt the beam or the specimen slightly (a few degrees) in different directions. In Fig. 2a the grain boundary is marked by a dashed line, which means that the contrast variations inside the grain are caused by dislocations and subgrain boundaries. The SAD pattern of the as-deformed sample (cf. Fig. 2b) shows diffraction rings corresponding to the body-centered cubic structure only since the superlattice reflections of the B2 structure are missing. The intensity along the rings is rather homogeneous, indicating that there is no pronounced texture. It should be noted that the same results in real and reciprocal space are obtained for samples heated up to 130 °C. Even for samples heated to 170 and 220 °C (cf. Fig. 2c), the grains observed in the TEM images do not change significantly in size or morphology compared to the as-deformed sample, and the larger grains show similar substructures caused by subgrain boundaries. However, the diffraction patterns of the samples annealed to 170 and 220 °C (cf. Fig. 2d) show the appearance of additional rings that are caused by the B2 superstructure. Therefore, the process responsible for peak I in the DSC curve can be identified unambiguously as chemical reordering. The TEM images obtained from samples heated to 370 °C (cf. Fig. 2e) show that most of the grains are nearly defect free, with clear grain boundaries, and the corresponding SAD pattern (cf. Fig. 2f) reveals sharp diffraction rings. Therefore, the broad second peak (II) in the DSC signal can be correlated to the recovery of defects inside the grains and at the grain boundaries. The third exothermic peak (III) is related to grain growth, which can be deduced from a comparison of the grain sizes of the samples annealed to 370 and 500 °C (cf. Fig. 2e and g), respectively. The bright-field image shows large grains with a very low density of dislocations (cf. Fig. 2g) and the corresponding diffraction pattern reveals sharp rings (cf. Fig. 2h).
The results of the evolution of the SAD patterns are summarized in the intensity profiles shown in Fig. 3. The integration along the rings as well as an automatic background subtraction was performed using the PASAD software package [21]. The profiles confirm that both the intensity and the sharpness of the superlattice reflections ((1 0 0), (1 1 1), (2 1 0) and (3 0 0)) increase during heating. As shown by the intensity profiles, the superlattice reflections are emerging at 170 °C. It should be pointed out that the subgrain structure present in the form of scattering domain size increases with temperature (as can be concluded from the change in the intensity profile shown in Fig. 3), whereas the grain size measured in the TEM dark-field images decreases.
To analyze the thermally induced process of reordering, TEM dark-field images were taken with fundamental reflections and compared with those of superlattice reflections (cf.Fig. 4). In the case of the fundamental reflection (2 0 0) (cf. Fig. 4a), the variation in contrast inside the imaged grain is caused by structural defects leading to small changes in the orientation of the lattice with respect to the incoming beam. In the case of the dark-field image taken with the corresponding superlattice reflection (1 0 0) (cf. Fig. 4b), small nanosized domains show up that are not visible in Fig. 4a. Since the intensity of the superlattice reflections is sensitive to the chemical LRO, comparison of Fig. 4a and b shows that at 170 °C the grains contain small chemically ordered domains of nanometer size (<5 nm). Therefore, the recurrence of the B2 superstructure takes place inside the grains by small ordered domains that show a rather homogeneous distribution (cf. Fig. 4b). In Fig. 4c the corresponding diffraction pattern is shown to illustrate the positions of the objective aperture used to form the dark-field images of the fundamental reflection (cf. Fig. 4a) and of the superlattice reflection (Fig. 4b). In the aperture only a small sector of the diffraction ring is included (∼10°). (It should be mentioned that all the dark-field images were taken with a tilted incident beam to show the reflected beam forming the image on the optical axis.) The experimental results show that the ordered domains are completely out of contrast in fundamental reflections, indicating that in the B2 superlattice structure the antiphase boundary (APB) faults bounding the ordered domains do not contain non-chemical fault components. As explained in the discussion (cf. Section 4.2), the situation is different from that in other intermetallic structures (e.g. Fe 3 Al [22] and Ni 3 Al [23]).
An analysis of the grain-size distribution was carried out from samples with different heat treatments (cf. Fig. 5). In total, more than 500 grains were analyzed using TEM dark-field images (as in the examples given in Fig. 2). The area of each grain was determined by digital imaging processing; the grain size is defined as the diameter of a circle with the same area as the actual grain. The histograms of the grain-size distribution together with their log-normal fits of the samples heated up to 220, 370 and 500 °C are shown in Fig. 5a-c, respectively. Fig. 5a reveals that the median of the grain diameter is ∼77 nm. (It should be mentioned that samples heated to lower temperatures have a similar median value.) From Fig. 5b, which shows the grain-size distribution at 370 °C, a median grain-size diameter of ∼35 nm only is deduced, which is smaller by a factor of 2 than that obtained at 220 °C. At 500 °C the median grain-size is ∼158 nm (cf. Fig. 5b); since in this case the thickness of the TEM foil and the grain size are similar, the measured grain size has to be multiplied by a factor of about 1.3 [24] leading to an actual grain size of ∼204 nm. In Fig. 5d the results of the measurements of the median grain sizes of the samples with different heat treatments are summarized.
Fig. 6 shows the evolution of the microhardness as a function of the heat treatment. The values of the microhardness reveal an inverse temperature-dependent behavior compared to that of the grain size: a more or less constant value up to 220 °C, a maximum at 370 °C and a drastic drop at 500 °C. Based on the measured values of the grain size (cf. Fig. 5d), the microhardness can be calculated by using the Hall-Petch relationship [25,26]. From the calculated values a curve is drawn that fits the measured data (cf. Fig. 6).
Metastable disordered nanocrystalline FeAl modifies its structure upon heating (heating rate 20 K min -1 ). The structural changes of nanocrystalline FeAl produced by HPT were tracked by combining the methods of DSC and TEM to identify the different processes occurring at different temperatures. The first exothermic peak (I) in the DSC signal that shows a strong maximum at 170 °C is interpreted by the occurrence of chemical reordering (cf. Fig. 1). The peak is asymmetric as the initial level of the heat flow is not reached after peak I because of an overlap with peak II. The recurrence of the B2 superstructure is indicated by the DSC curve, showing that it already starts above 130 °C (0.24T m ). The reordering is confirmed by analyzing the TEM diffraction patterns (cf. Fig. 3) that show rings corresponding to the B2 superstructure at temperatures above 130 °C, whereas the observed nanostructures do not change significantly in the temperature range up to 220 °C (cf. Fig. 2). The recovery of the B2 superstructure occurs in the form of small, chemically ordered domains (cf. Fig. 4), that are rather homogeneously distributed and grow during further heating, as indicated by the sharpening of superlattice reflections (cf. Fig. 3). It can be assumed that small ordered volumes retained after SPD as reported for L1 2 LRO Ni 3 Al [27] grow during the process of reordering. They are retained since the deformation induced process of disordering facilitated both by the formation of a high density of dislocations with APB faults and APB tubes is fragmenting the grains [28,29] and destroys the LRO locally by APB faults and not randomly as thermal disordering. As a consequence, a high density of very small ordered domains are retained in the course of severe plastic deformation.
During heating, the ordered nuclei grow, and the growth process is facilitated by the large number of dislocations and vacancies produced during the HPT deformation. During reordering, the deformation-induced dislocations can also act as nucleation sites, and at the temperatures at which they become mobile superlattice partial dislocations can remove APB faults by moving backward. This leads to the conclusion that in FeAl disordered by SPD the process of reordering is a combination of both thermal ordering and defect-induced ordering. (It should be mentioned that in stoichiometric B2 FeAl and in the present composition it is in principle impossible to separate these two processes since B2 FeAl is ordered up to the melting point and therefore disordered FeAl can only be achieved by means of SPD.)
The reduction in the area of APB faults occurring during the growth of the ordered domains causes the exothermic peak I in the DSC curve. The temperature regime in which recovery of the B2 superstructure takes place is in good agreement with several studies carried out on ballmilled samples [30,13,14,9]. The measured enthalpy change of 4.5 kJ mol -1 given by the area of peak I caused by reordering should be treated as an estimation since peak II centered around 320 °C is so broad that it overlaps with peak I. It is interesting to note that the value of enthalpy change determined in this study compares with that obtained for ball-milled samples [13] but is about a factor two larger than the value reported from the mechanically alloyed material [14]. This can be explained in the following ways: (i) in the case of ball-milled samples [13], already alloyed starting material was subjected to severe plastic deformation by ball milling; therefore the situation is similar to the present one, where an alloyed starting material is severely deformed by HPT. (ii) In the case of mechanical alloying, powders of the pure metals are used as starting materials and alloyed by the method of ball milling. The fact that in this case a smaller value of the enthalpy change was reported [14] than in case (i) indicates that only a fraction of the material had been alloyed.
The structural evolution during heating up to 500 °C reveals the nature of the two other peaks (II and III) in the DSC curve. According to the TEM investigations (cf. Figs. 2 and 5), two further processes were identified: (i) defect recovery, leading to static recrystallization, and (ii) grain growth. Peak II centered around 320 °C (0.35T m ). This can be attributed to the process of recovery of dislocations starting at low temperatures of about 200 °C (0.28T m ), at which the process of reordering is not finished, as concluded from the TEM study. Therefore, reordering and defect recovery occur simultaneously between 220 and 370 °C. From the TEM dark-field images and the temperature-dependent grain-size distribution, it can be concluded that during these early stages of dislocation recovery the grain size decreases by a factor of 2 by heating (cf. Fig. 6).
In this context, it is interesting to note that in FeAl the TEM images used to determine the grain size are not disturbed by the residual contrast of APB faults and therefore confusion of the grain-size measurements can be excluded. This is based on the experimental observations (cf. Fig. 3). The results of theoretical considerations show that in B2 alloys the APB faults are pure chemical faults, as deduced from pair potential calculations [31]. Therefore the situation is different from other intermetallic alloys. For example, in Fe 3 Al (DO 3 structure) a residual contrast of the APB faults is visible in fundamental reflections that can be explained by an appropriate model [22]; in the case of Ni 3 Al (L1 2 structure) the observed residual contrast of the APB faults was analyzed [28] and agrees with a non-chemical-fault component as deduced by ab initio calculations [32].
The thermally induced reduction in grain size can be considered as continuous static recrystallization triggered by defect rearrangement since smaller grains with highly reduced defect density are formed. It is proposed that a possible mechanism to obtain smaller grains during this thermal treatment could be the transformation of subgrain boundaries into highangle grain boundaries by absorbing dislocations. This transformation seems to be closely related to the process of reordering by the growth of chemically ordered domains as, first, their boundaries interact with dislocations, and secondly, the reduction in the grain size is reached at 370 °C (cf. Fig. 6), at which temperature the LRO is already pronounced (cf. Fig. 3).
Peak III in the DSC curve (cf. Fig. 1) is interpreted as classical grain growth starting at about 370 °C, at which temperature the recurrence of order is finished and the majority of defects are already recovered. Thus, the driving force for grain growth is mainly given by the reduction in the grain-boundary area in the ordered alloy. In the ordered state the migration of grain boundaries slows down and grain growth is expected to be reduced [33], which might explain the broad peak III.
Microhardness (cf. Fig. 6) is nearly constant (within the error bars) in the temperature regime of pronounced reordering (up to 220 °C). Therefore, ordering and the formation of ordered domains within nanograins have only little impact on the microhardness. The strong increase in the microhardness in the temperature regime of defect recovery leading to grain size reduction can be explained either by a Hall-Petch mechanism or by hardening due to the reduction in dislocation sources. The variation of microhardness according to the Hall-Petch relationship as it is shown in Fig. 6 indicates grain-size hardening as the maximum correlates with the minimum of the grain size (cf. Fig. 6). On the other hand, the dislocation density decreases considerably in the corresponding temperature regime, yielding grains of small size (∼30 nm) that have sharp boundaries and a reduced dislocation density. This leads to the conclusion that the recovery process could effect a reduction in the number of dislocation sources both by the formation of equilibrium grain boundaries and by the reduction in the dislocation density in the grain interior. This could also cause the observed increase in the hardness since a higher stress is needed to activate new dislocation sources in the small grains [34]. As a consequence, other deformation processes, such as grain-boundary mediated processes, can become active at high stresses [35]. Finally, the drop in microhardness above 370 °C is interpreted as a consequence of grain growth.
The identification of different processes resulting from different heat treatments allows the materials properties to be tuned. Therefore, for example, the production of fully dense bulk B2 ordered nanocrystalline FeAl exhibiting a low dislocation density is possible by HPT deformation followed by heating to 370 °C.
Systematic TEM investigations combined with DSC measurements are used to identify the processes of structural modification taking place during heating of bulk nanocrystalline disordered FeAl produced by HPT plastic deformation of B2 long-range ordered samples:
The recurrence of the B2 LRO taking place mainly between 130 and 220 °C (0.24 and 0.29T m ) can be explained by the formation of nanosized chemically ordered domains that are homogeneously distributed within the nanosized grains. It is assumed that ordered volumes retained during severe plastic deformation act as heterogeneous nuclei for ordering. Dislocation density and grain size were shown to be unaffected below 0.29T m .
• A thermally induced reduction in the grain size by a factor of two occurs in FeAl disordered by SPD. It is concluded that the reduction is caused by the rearrangement of a high density of superlattice partial dislocations transforming subgrain boundaries into high-angle grain boundaries. After heating to achieve reordering, the dislocations in the interior of the grains are, for topological reasons, connected with APB faults. By further heating up to 370 °C (0.38T m , a temperature above the recovery peak), the density of the APB faults is reduced by dislocations moving to the subgrain boundaries. This leads to smaller grains, with a low dislocation density in their interior and large-angle grain boundaries.
• Classical grain growth of nanograins occurs predominantly at about 410 °C (0.41T m ) and is governed by the reduction in the grain boundary area only, since dislocation recovery leading to defect-free grains takes place at lower temperatures.
The variation in the strength monitored by microhardness can be interpreted either by a Hall-Petch mechanism or by the reduction in dislocation sources as a consequence of the grain size reduction. DSC signal obtained by heating at a constant rate (20 K min -1 ) of disordered nanocrystalline FeAl after HPT deformation. The curve shows three exothermic peaks, I, II and III, at about 170, 320 and 410 °C, respectively. The temperatures at which samples were studied by TEM are marked by crosses (⊗). FeAl intensity profiles (intensity vs. diffraction vector g) obtained by azimuthal integration of TEM diffraction patterns taken from the as-deformed state and from samples heated to 130, 170, 220, 370 and 500 °C. The fundamental reflections ((1 1 0), (2 0 0), (2 1 1) and (2 2 0)) are present at all temperatures, whereas the intensity of the superlattice reflections increases with increasing temperature, indicating the transition from a disordered structure to a B2 LRO structure.
Mangler et al. Page 11 Published as: Acta Mater. 2010 October ; 58(17): 5631-5638.
Sponsored Document Sponsored Document Mangler et al. Page 12 Published as: Acta Mater. 2010 October ; 58(17): 5631-5638. Sponsored Document Sponsored Document Sponsored Document
Published as: Acta Mater. 2010 October ; 58(17): 5631-5638.Sponsored DocumentSponsored Document Sponsored Document
Published as: Acta Mater. 2010 October ; 58(17): 5631-5638.
The authors thank
The most common cause of amyotrophic lateral sclerosis (ALS) is mutations in superoxide dismutase-1 (SOD1). Since there is evidence for the involvement of non-neuronal cells in ALS, we searched for signs of SOD1 abnormalities focusing on glia. Spinal cords from nine ALS patients carrying SOD1 mutations, 51 patients with sporadic or familial ALS who lacked such mutations, and 46 controls were examined by immunohistochemistry. A set of anti-peptide antibodies with specificity for misfolded SOD1 species was used. Misfolded SOD1 in the form of granular aggregates was regularly detected in the nuclei of ventral horn astrocytes, microglia, and oligodendrocytes in ALS patients carrying or lacking SOD1 mutations. There was negligible staining in neurodegenerative and nonneurological controls. Misfolded SOD1 appeared occasionally also in nuclei of motoneurons of ALS patients. The results suggest that misfolded SOD1 present in glial and motoneuron nuclei may generally be involved in ALS pathogenesis.
Amyotrophic lateral sclerosis (ALS) is a neurodegenerative disease mainly characterized by progressive loss of upper and lower motoneurons and death usually ensues within 3-5 years after diagnosis. More than 150 mutations in the ubiquitously expressed enzyme superoxide dismutase-1 (SOD1) have been associated with familial ALS (FALS) and such mutations are found in about 6% of all ALS patients [35,45] (http://alsod.iop.kcl.ac.uk). The mutations confer a cytotoxic property on the enzyme that is poorly understood [16]. Nine of the mutations cause long C-terminal truncations. These mutant SOD1s lack the stabilizing C57-C146 intrasubunit disulfide bond and also strand 8 in the b-barrel core of the protein, and thus cannot adopt native folding. The most extensively studied truncation mutant, Gly127insTGGG (G127X), has been found to be rapidly degraded and to be present only in minute amounts in the human spinal cord [22]. The mutant SOD1s associated with ALS should most likely cause ALS by essentially the same mechanism. This implies that any common cytotoxic conformational species of SOD1 should be misfolded and present in minute amounts in the tissue.
The wild-type human SOD1 can also have neurotoxic effects. Transgenic overexpression in mice leads to axonal abnormalities and a late moderate loss of ventral horn neurons [19,24], and it exacerbates disease caused by mutant SOD1s [11,19]. Compared to D90A, the most common of the ALS-associated mutant SOD1s, wild-type human SOD1 is between half and equally neurotoxic in mice [23]. Damage to this long-lived protein causes destabilization and results in interaction properties and toxicity similar to those found for mutant SOD1s [5,34].
There is considerable evidence for the involvement of non-neuronal cells in the pathogenesis of ALS. In murine SOD1 models, down-regulation of SOD1 expression in astrocytes [47] and microglia [2] slows progression after disease onset, whereas down-regulation in motoneurons delays onset [4,47]. In chimeric mice, motoneurons expressing mutant SOD1s are spared if surrounded by nontransgenic glial cells [10]. Neuron-restricted synthesis of G93A mutant SOD1 in mice is, however, sufficient to cause motoneuron pathology and a late ALS-like disease [20,43]. Mutant SOD1 expression in oligodendrocytes might protect against loss of motoneurons [46]. The exact role of non-neuronal cells in the pathogenesis of human ALS is unknown, but all cells implicated in the transgenic models show signs of activation or alterations.
We have developed a set of antibodies which allow specific detection of misfolded SOD1 against the abundant background of native enzyme in the human CNS. Using these antibodies, we have previously shown that inclusions containing misfolded SOD1 are regularly present in motoneuron somas of ALS patients, both with and without SOD1 mutations [12]. The finding of misfolded SOD1 in sporadic ALS (SALS) was recently confirmed by Bosco et al. [5], and they also found that wild-type SOD1 immunopurified from SALS tissues inhibited fast axonal transport in a manner similar to H46R mSOD1 indicating a pathological mechanism common to SALS and SOD1 FALS. In this study, we searched for signs of SOD1 involvement in ALS focusing on non-neuronal cells. The major novel finding is that misfolded SOD1 is regularly present in the nuclei of ventral horn astrocytes, microglia, and oligodendrocytes in ALS patients carrying SOD1 mutations as well as in sporadic and familial patients lacking such mutations. Only negligible staining was seen in control patients with neurodegenerative and non-neurological disease.
Material was collected at autopsy from patients enrolled at the Department of Neurology, Umea ˚University Hospital. The group consisted of 43 patients with SALS [mean age 68 ± 12 (17-88) years] and 8 patients with FALS without SOD1 mutations [mean age 59 ± 11 (39-71) years]. The non-SOD1 FALS patients were subjected to genetic screening for mutations in the following genes; transactivation-responsive DNA-binding protein of molecular weight 43 kDa (TDP-43), fused in sarcoma (FUS), vesicle associated membrane protein B (VAPB), angiogenin and optineurin, and none were found. Moreover, nine ALS patients carrying SOD1 mutations were included in the study: six patients homozygous for the D90A mutation, two patients heterozygous for the G127X mutation, and one patient heterozygous for the D101G mutation [mean age 62 ± 11 (43-75) years]. ALS patients were diagnosed in accordance with the revised criteria of El Escorial [6]. In addition, two patients with spinobulbar muscular atrophy (SBMA) (aged 52 and 76 years, respectively) were included in the study.
Tissues were also collected from 26 control patients with other neurodegenerative diseases [mean age 74 ± 19 (2-92) years; 12 with Alzheimer's disease; six with Parkinson's disease; three with multiple sclerosis, one with tuberous sclerosis, one with Huntington's disease, one with spinocerebellar ataxia 7, one with autosomal dominant progressive external ophthalmoplegia, and one with argyrophilic grain disease]. Tissues from 20 control patients without neurological disease [mean age 69 ± 17 (37-91) years; all patients died from heart conditions or pneumonia) were also examined. The post-mortem time for all patients was estimated to be between 0 and 3 days, and there was no difference in postmortem time between the groups. The study adhered to the tenets of the Helsinki Declaration, and was approved by the Ethical Committee of Umea ˚University. For detailed information on subject collection, see Electronic Supplementary Material.
For detection of misfolded SOD1, a set of polyclonal rabbit (Ra) antibodies (ab) raised against keyhole limpet hemocyanin-coupled peptides corresponding to amino acids 4-20 (Ra 4-20 ab), 57-72 (Ra 57-72 ab), and 131-153 (Ra 131-153 ab) in the human SOD1 sequence were used. A set of polyclonal chicken (Chi) antibodies corresponding to the same set of peptides was also used (Chi 4-20 ab, Chi 57-72 ab, Chi 131-153 ab, respectively). The antisera were affinity purified in two steps as described in previous papers [12,23]. G127X mutant SOD1 has following Gly-127 a 5 amino acid long neopeptide before the C-terminal truncation. CNS material from G127X mutant SOD1-carriers was also stained with a mutant-specific Ra SOD1 peptide antibody directed against amino acids 123-132 in the mutant: CADDLGGQRWK, (neopeptide shown in bold). This antibody shows no reaction with wild-type human SOD1 [22]. Other commercially available antibodies used in the study are presented in Online Resource, Table S1.
After blocking, sections were incubated with primary antibodies and subsequently with corresponding fluorescent secondary antibodies or with biotin-conjugated secondary antibodies coupled to an avidin-horseradish peroxidase conjugate, and were visualized using aminoethylcarbazole as the precipitating enzyme product. The primary antibodies used were either the anti-SOD1 peptide antibodies described above or the antibodies shown in Online Resource Table S1. The sections were examined by confocal laser microscopy or using an Olympus BX50 light microscope. A four-tiered semi-quantitative scale was used to estimate the number of glial cell nuclei in each cross-section of spinal cord showing misfolded SOD1 staining. The levels were: 0 = no glial cells with nuclear staining; 1 = \25% of the glial cells showing nuclear staining; 2 = 25-75% of the glial cells showing nuclear staining; and 3 = [ 75% of the glial cells showing nuclear staining. Sections from cervical, thoracic, and lumbar spinal cord were analysed for each patient and the result is presented in Table 1. Data for proportion of glial cell nuclei with staining were calculated as median (range). For detailed information on staining methods, see Electronic Supplementary Material.
We have developed a set of rabbit and chicken antibodies against peptides in the human SOD1 sequence. The antibodies show high specificities for SOD1 in western immunoblots of human ventral horn extracts and they react only with misfolded SOD1 species in immunocapture experiments [12] (see also Online Resource, Fig. S1a, b). To further ascertain their specificities for misfolded SOD1 in immunohistochemical applications, the Ra 57-72 and the Ra 131-153 antibodies, and the Chi 57-72 and 131-153 antibodies, respectively, were subjected to preincubations. When pre-incubated with the appropriate immunizing peptide or with denatured SOD1, complete blocking of the signal was seen (Online Resource, Fig. S2b, d). Native SOD1 had no blocking effect (Online Resource, Fig. S2c). Exclusion of the primary antibodies resulted in the disappearance of all immunostaining. These studies show that the antibodies can be used for the specific detection of misfolded SOD1 species against the abundant background of natively folded SOD1 in tissues. In the following, positive immunohistochemical staining with these antibodies will be referred to as misSOD1 staining.
Spinal cord sections from the nine ALS patients carrying SOD1 mutations were stained with the anti-SOD1 peptide antibodies. Aggregates/inclusions of misfolded SOD1 in the soma, axons, and dendrites of affected motoneurons were seen as previously described (Fig. 1d) [33,36]. There was, however, also prominent staining of the nuclei of glial cells (Figs. 1a, d, 2a). The staining was composed of granules measuring 0.5-2 lm, which were dispersed throughout the nuclei. The immunohistochemical findings
Table of findings from sections stained with the Ra 131-153 ab. Data for number of patients with nuclear staining in glial cells show the total number of patients in parenthesis. Data for proportion of glial cell nuclei with staining are shown as median (range) referring to a four-tiered semi-quantitative scale (0 = no glial cells with staining; 1 = \ 25% of the glial cells with staining; 2 = 25-75% of the glial cells showing staining; 3 = [ 75% of the glial cells showing staining). The total number of patients in each group was as follows: SOD1 FALS, 9; non-SOD1 FALS, 8; SALS, 43; neurodegenerative controls, 26; non-neurological controls, 20. Sections from all levels were not available from all patients a Since thoracic spinal cord only was available from two patients and lumbar spinal cord only from three patients from the neurodegenerative control group, each result is presented individually were identical using the different rabbit and chicken anti-SOD1 peptide antibodies, and the intranuclear SOD1 positive granules in glial cells were seen at all levels of the spinal cord and in both grey and adjacent white matter. The staining was detected with both bright-field and immunofluorescence microscopy (Figs 1a, 2a). Using a four-tiered semi-quantitative scale to estimate the proportion of glial nuclei with misSOD1 staining in one section, all patients homozygous for the D90A mutation showed misSOD1 staining in glial cell nuclei, ranging from less than 25% in two patients and up to almost 100% in four patients (Table 1). The D101G patient showed misSOD1 staining in approximately 25% of the glial cell nuclei (Table 1). Both patients with the G127X mutation showed nuclear staining of glial cells, but they differed in the proportion of cells that stained. Using the G127X-specific Ra 123-132 ab, misSOD1 staining was seen in \25% of the glial cell nuclei in both patients. One of the two patients also had some nuclear staining in motoneurons. Interestingly, the Chi 131-153 ab, raised against a peptide sequence that is absent in the G127X mutant SOD1, gave rise to misSOD1 staining in glial cell nuclei in both G127X patients, and in one of them it was seen in more than 75% of the glial cell nuclei. Double staining with the Chi 131-153 ab and the G127X mutant-specific ab mostly showed separate aggregates of misSOD1 staining, without co-localization (Online Resource, Fig. S3a-c). Performing double-staining with an antibody that detect both misfolded wildtype and G127X mutant SOD1 protein, the Chi 53-72 ab, with the mutantspecific G127X ab yielded sometimes co-localization as in large aggregates in motoneurons (Online Resource, Fig. S3d-f) but also single staining with the Chi 57-72 ab as in the arrow-head marked glial nucleus. This indicates that To establish that the misSOD1 staining observed was present inside neuronal and glial nuclei, double immunofluorescence staining with an antibody to the structural nucleoporin protein NUP 62, which localizes to the nucleoplasmic region of the nuclear envelope, was performed [29]. Double staining with the NUP 62 antibody revealed the nuclear envelope, and inside it misSOD1 staining was seen (Online Resource, Fig. S4). MisSOD1 staining can thus be intranuclear in motoneurons and is principally intranuclear in glial cells.
Misfolded SOD1 is regularly found in the nuclei of glial cells and occasionally in motoneuron nuclei of ALS patients lacking SOD1 mutations Sections of spinal cord from 43 patients with SALS and from 8 patients with FALS, all of whom lacked SOD1 mutations, as well as 2 patients with SBMA, were stained with the antibodies to SOD1 peptides. Intranuclear staining in glial cells was seen both in the SALS and the FALS patients. The staining patterns were virtually identical to those seen in the patients carrying SOD1 mutations. All sporadic ALS patients showed misSOD1 staining in their glial nuclei (Figs. 1c, e, 2b, 3a, d, g, 4a and Online Resource, Fig. S4f), as did seven of the eight FALS patients who lacked SOD1 mutations (Figs. 1b, 2c, 4d). We were unable to find nuclear misSOD1 staining in one FALS patient. On average, the proportion of glial cells in each section with misSOD1 staining was close to 75%, and ranged from \25% in some patients up to 100% in others (Table 1). The intranuclear SOD1 positive granules in glial cells were primarily seen in the ventral horn and adjacent white matter but some could also be observed in the dorsal horn and the dorso-medial white matter. Overall, the nuclear misSOD1 staining in glial cells was at least as prominent in patients lacking SOD1 mutations as in patients carrying such mutations (Table 1). Intranuclear staining for SOD1 in glial cells was also present in the two SBMA patients included in the study.
Small and numerous aggregates of misfolded SOD1 were seen in the cytoplasm of motoneurons from ALS patients nucleolus (d). In motoneurons with nuclear staining, the cytoplasm appears to contain less of the small misSOD1 inclusions. In panels b and e, misSOD1 aggregates can be seen scattered throughout the cytoplasm of motoneurons, and are particularly abundant in the somal area, while not penetrating the nuclear envelope (arrows). f MisSOD1 staining in a sample from a neurodegenerative control patient with Alzheimer's disease. Glial cell nuclei and motoneurons lack mis-SOD1 aggregates/inclusions. Scale bars are 40 lm (a, b), 50 lm (c), 20 lm (d, e) or 30 lm (f)
lacking mutations in the SOD1 gene, as we have previously reported [12]. These misfolded SOD1 aggregates were often scattered throughout the cytoplasm and were particularly abundant in the perinuclear area, but did not penetrate the nuclear envelope (Figs. 1f, 2b, e, 4d) [12]. However, occasionally motoneurons also showed nuclear misSOD1 staining. The staining then appeared as nuclear aggregates approximately 0.5-3 lm in size, often in association with the nucleolus (Figs. 1e, g, 2d and Online Resource, Fig. S4a). In motoneurons with nuclear staining, the cytoplasm often appeared to contain fewer aggregates of misfolded SOD1 (Fig. 1e). The motoneurons thus showed two types of misSOD1 staining: either intranuclear staining in motoneurons or a perinuclear misSOD1 staining with numerous aggregates of misfolded SOD1 in the soma but no sign of intranuclear staining (Fig. 2e). In the latter case, the nucleus appeared to be spared, leaving the impression that the nuclear envelope was acting as a barrier. The two types of motoneuron staining appeared to co-exist in the ventral horns and they could be seen in the same section, although containing misSOD1 are astrocytes. d-f Double immunofluorescence labeling showing misSOD1 (green) and the microglial marker Iba1 (red). In panel f, some of the glial cells with nuclear misSOD1 staining are microglia. g-i Double immunofluorescence labeling showing misSOD1 (green) and the oligodendroglial marker Olig 2 (red). In panel i, misSOD1 is present in the nuclei of oligodendrocytes. Scale bars are 20 lm (a-c) and 10 lm (d-i)
somal and perinuclear staining was more common and was seen in all ALS patients investigated. On the other hand, the intranuclear staining in motoneurons was seen in 30 of 43 SALS patients and in six of eight FALS patients, all of whom lacked SOD1 mutations. Four out of the six patients homozygous for the D90A mutation and one of the two SBMA patients showed inclusions of misfolded SOD1 in the nuclei of motoneurons. The nuclear staining was only seen in a few motoneurons per section, and not in all sections investigated.
Nuclear staining of SOD1 in glial cells was specific for tissues affected by ALS and was not seen in controls with non-neurological or other neurodegenerative diseases Spinal cord sections from 20 patients who died of nonneurological diseases were stained with the anti-SOD1 peptide antibodies. Only one patient showed misSOD1 staining in the nucleus of glial cells and motoneurons (Fig. 1h). The proportion of staining was \25% in that patient, and staining was not seen in all sections investigated (Table 1). Twenty-six controls with other neurodegenerative diseases were also examined. Twentythree of these controls did not stain for misfolded SOD1 in ventral horn glial cell nuclei; nor was misfolded SOD1 detected in the nuclei of motoneurons (Figs. 1i, 2f). As the spinal cord is not the primary site of pathology in other neurodegenerative diseases, we also examined the brain areas mainly affected in Alzheimer's disease, Huntington's disease, and Parkinson's disease. We could not detect any misfolded SOD1 immunoreactivity in areas of the hippocampus, temporal, and frontal cortex, putamen, striatum, and substantia nigra (data not shown). The 4 of 46 controls that stained positive for misSOD1 showed increased GFAP staining in the spinal cord. All four also showed cytoplasmic inclusions in motoneurons when stained with ubiquitin, and in one neurodegenerative control, a patient with Alzheimer's disease, a few skein inclusions could be seen in the motoneuronal cytoplasm as well. (c) revealed some co-localization of ubiquitin and misSOD1 (arrows), although misSOD1 could also be seen separate from ubiquitin (arrowheads). Staining with anti-TDP-43 antibody revealed that some glial nuclei contained TDP-43 (e, arrows). Occasionally, misSOD1 and TDP-43 were present in the same glial cell nuclei (f) but no distinct co-localization of misSOD1 and TDP-43 was apparent. Scale bars are 20 lm (a-f)
Misfolded SOD1 is also found motoneurons and glial cells in brainstem and motor cortex Mid-olivary sections from the brainstem from 12 neurodegenerative control patients, from 11 SALS patients, 2 non-SOD1 FALS patients and 8 SOD1 mutated patients were investigated for misfolded SOD1 by immunohistochemistry. In brain stem motor nuclei, we found intranuclear misSOD1 staining in glial cells in all ALS patients, although the number varied between patients. Staining was generally found in fewer glial cells than seen in the spinal cord. The staining was mainly found in motor nuclei and in the inferior olivary nucleus. We also found a primarily cytoplasmic staining in brain stem motoneurons in all ALS patients, while nuclear staining was only occasionally observed. We found no staining in the control patients investigated.
Sections from motor cortex were investigated in eight SALS patients. All patients had misSOD1 staining in glial cell nuclei and some cytoplasmic staining in neurons. The staining was less than found in brainstem and spinal cord.
Identification of glial cell types with nuclear staining of misfolded SOD1 Next, we wanted to identify the glial cell types carrying nuclei that were positive for misfolded SOD1. We used double-labeling immunofluorescence with the anti-SOD1 peptide antibodies and antibody markers for astrocytes, microglia, and oligodendrocytes. Using confocal laser scanning microscopy, we could show that most of the nuclei carrying misSOD1 staining were present in cells positive for the astrocytic marker GFAP (Fig. 3a-c). Double-labeled astrocytes were seen both in white and grey matter, and they were seen in all cases where positive staining for misfolded SOD1 was seen. We also found that some of the nuclei that were positive for misfolded SOD1 were present in cells that expressed the microglial marker Iba1 (Fig. 3d-f), although not all microglia had misfolded SOD1 in the nucleus. In addition, a few of the glial cell nuclei that stained for misfolded SOD1 expressed the olig 2 marker (Fig. 3g-i). These double-labeled oligodendroglia nuclei were much smaller and occurred less frequently than the astrocyte nuclei in the sections investigated. SOD1-positive inclusions in glial cell nuclei occasionally co-localized with ubiquitin, but not with TDP-43 or p62
In the samples from SALS and FALS patients lacking SOD1 mutations stained with an antibody directed against TDP-43, cytoplasmic inclusions in motoneurons were seen in all sections investigated, as skein-like and/or fine granular punctuate inclusions, in accordance with the previous reports [1,31]. For a more detailed description, see Electronic Supplementary Material. When ALS patients lacking SOD1 mutations were double-stained with the anti-SOD1 peptide antibodies, TDP-43 was not found to colocalize with the misfolded SOD1 aggregates seen in the somas of motoneurons. Occasionally, misfolded SOD1 and TDP-43 were located in the same glial or motoneuron nucleus but no co-localization could be demonstrated (Fig. 4d-f). Interestingly, two control patients, one with Alzheimer's disease and another with cardiovascular disease, had TDP-43 positive neuronal cytoplasmic inclusions in the shape of skein inclusions in the cytoplasm of spinal cord motoneurons (data not shown). The staining was not different from that seen in SALS patients. Both of these controls were negative when staining for misfolded SOD1. The two SBMA patients also showed cytosolic TDP-43 staining typical of SALS patients (data not shown).
Since misfolded proteins present in the nuclei are degraded through the ubiquitin-proteasome pathway, we performed double immunofluorescence staining for misfolded SOD1 and ubiquitin. Occasionally, misfolded SOD1 co-localized with ubiquitin in glial cell nuclei, although most misSOD1 staining in glial cell nuclei was ubiquitinnegative (Fig. 4a-c).
The polyubiquitin-binding protein p62 (sequestosome 1) has been proposed to act as a factor for shuttling of polyubiquitinated proteins to the proteasome for degradation [38], and it has also been reported to interact with mutant SOD1 in cultured cells [14]. Using two anti-p62 antibodies, staining of sections from ALS patients and controls failed to show any nuclear immunoreactivity in glial cells and motoneurons (data not shown). However, all ALS patients and 8 of the 47 controls had p62 staining in the cytoplasm, in the shape of punctuate and skein-like inclusions, as has been described before [32]. The cytoplasmic p62 aggregates showed co-localization when double-labeling was performed with ubiquitin, but no co-localization of misfolded SOD1 and p62 was seen (Online Resource, Fig. S5).
The main new finding in this study is that misfolded SOD1 is present as small aggregates in the nuclei of glial cells of spinal cord tissue from ALS patients. These aggregates were found in ALS patients carrying SOD1 mutations as well as in sporadic and familial ALS patients lacking such mutations. Nuclear staining of misfolded SOD1 in glial cells was found in 59 of 60 ALS patients investigated. Only 4 of the 46 non-neurological and neurodegenerative controls had misSOD1 staining, but it was sparse and not seen at all levels of the spinal cord (Table 1). Since the vast majority of control cases lacked staining, the misSOD1 staining seen in ALS cases is unlikely to represent postmortem changes, which strongly suggests that the finding was related to ALS pathology. The glial cells containing misfolded SOD1 in the nucleus were mostly astrocytes but some microglial and oligodendroglial nuclei also stained positive for misfolded SOD1 (Fig. 3). Misfolded SOD1 was occasionally found in the nuclei of motoneurons.
Many histopathological studies which included staining for SOD1, have been carried out before on ventral horns from ALS patients with and without SOD1 mutations. It could appear curious that the glial nuclear staining not has been observed before. The likely explanation is that in most cases antibodies raised against whole human SOD1 have been used. Such antibodies react avidly with both native and denatured SOD1 [12], and the minute amounts of misfolded SOD1 will thus be masked by the staining of the abundant native SOD1 in the cells. Other studies using monoclonal conformation-specific antibodies directed towards peptides in the SOD1dimer interface or in the unfolded beta-barrel have failed to show misSOD1 staining in patients with sporadic disease or non-SOD1 mediated FALS [28,30]. A possible explanation for this could be that SOD1 in sporadic disease might have different misfolded conformation/-s compared to SOD1-mediated disease. Monoclonal peptide antibodies with only one specific antigen determinant might fail to detect other misfolded conformations, whereas polyclonal antibodies with several specific antigenic determinants could detect different misfolded confirmations. Finally, the antigen retrieval could have differed. If it is too strongly denaturing, extensive misfolding of the background native SOD1 might occur, hampering detection of the misfolded SOD1 present in the tissue.
Our finding that misfolded SOD1 is present in glial cell nuclei of FALS patients carrying SOD1 mutations as well as in ALS patients lacking such mutations support the notion of a pathological interplay between neurons and glial cells. It raises the question of how these non-motoneuron abnormalities might contribute to motoneuron degeneration and disease progression. Studies using transgenic ALS mice have shown aggregates of mutant SOD1 in astrocytes, microglia, and oligodendrocytes [7,17,20]. Regarding studies on human post-mortem tissue, cytosolic inclusions in spinal cord astrocytes, intensely stained with an antibody to SOD1, have been observed in two long-term surviving FALS patients carrying SOD1 mutations [27]. In the present study, all nine FALS patients carrying SOD1 mutations had staining of mutant SOD1 in glial cells, although the staining of misfolded SOD1 was primarily seen in the nuclei.
When using the mutant-specific G127X Ra 123-132 SOD1 ab, approximately 25% of the glial cell nuclei in both FALS patients with the G127X SOD1 mutant were stained, whereas using the Chi 131-153 ab, (raised to a peptide sequence that is absent in the G127X-mutated SOD1 protein) also gave staining of misfolded SOD1 in glial cell nuclei in the G127X patients. Performing doublestaining with the two antibodies yielded mostly separate aggregates of misSOD1 staining and no co-localization was seen (Online Resource, Fig. S3a-c). This indicates that misfolded wild-type SOD1 protein may participate in the disease process even in patients carrying mutations.
Wild-type native SOD1 is normally located in the cytosol, the intermembrane space of mitochondria, and in the nucleus [44]. Bidirectional movement of small proteins occurs via passive diffusion through the nuclear pore complex, whereas proteins larger than about 40 kDa generally require specific transportation [41]. The size of the native SOD1 dimer is 32 kDa, and it should, therefore, be able to move freely between the nucleus and the cytosol. Accordingly, Chang et al. [8] found the native SOD1 concentration in the nucleus to be at least half that in the cytoplasm. In transgenic ALS models carrying mutant human SOD1s, misfolded SOD1 lacking the stabilizing intrasubunit C57-C146 disulfide bond is enriched in the susceptible spinal cord [48]. Such disulfide-reduced SOD1 readily forms aggregates in vitro [9,13] and is the major component of SOD1 aggregates present in symptomatic transgenic mice [3,26]. Misfolded disulfide-reduced SOD1 is consequently a likely source of the aggregates/inclusions of SOD1 observed in the present study. In glial cells, the aggregation mainly appears to take place in the nuclei whereas in neurons it is more prominent in the soma. Unlike dimeric and monomeric SOD1, aggregated SOD1 will be trapped and potentially accumulated in the nuclei. It is possible that both the nucleus and the cytoplasm contain targets that are vulnerable to misfolded SOD1, but that the cytoplasm of glial cells may be better equipped to neutralize or degrade the proteotoxic agent. Autophagy is the major mechanism for degradation of aggregates [18,37] and the absence of this process in nuclei might contribute to the staining pattern of misfolded SOD1 observed. There is no evidence for extensive glial cell death in ALS. This suggests that any harmful effects of the misfolded SOD1 would not lead to glial death, but rather to activation causing damage to motoneurons. Whether the nuclear aggregates are cytotoxic, or merely terminal markers of toxic soluble misfolded disulfide-reduced SOD1 (or both species essentially innocent) can only be speculated on at present. Misfolded SOD1 may be toxic by exposing internal structures that interact with essential nuclear factors, or it may aggregate with such factors, or (as aggregates) it may physically block processes in the nuclei. Mutant SOD1 has been reported to have both RNA [15] and DNA-binding capacity [21], indicating the possibility of direct interaction with polynucleotides. In this context, it is interesting to note that several proteins found to be mutated in motor neuron diseases, such as TDP-43, FUS, angiogenin, and senataxin, can influence gene transcription and RNA turnover [40]. More details on TDP-43 and SOD1 are reported in Electronic Supplementary Material.
In neurodegenerative conditions, such as Alzheimer's and Parkinson's diseases, proteins found to be mutated in some familial patients (amyloid precursor protein/b-amyloid, a-synuclein) are generally assumed to be key players in the pathogenesis in sporadic patients. Immunohistochemical results constitute the strongest evidence for these suppositions. Here we have shown that inclusions of the misfolded SOD1 are regularly present in glial nuclei, both in ALS patients carrying SOD1 mutations and in sporadic and familial cases lacking such mutations. The finding expands on the previous demonstration of SOD1 inclusions in the soma of motoneurons in ALS patients lacking SOD1 mutations [5,12]. Together, these studies suggest that misfolded SOD1 is generally involved in the pathogenesis of ALS.
As with SOD1, mutations in two other proteins, TDP-43 and FUS, have been found to cause ALS with a phenotype spectrum similar to that seen in sporadic disease. Inclusions containing TDP-43 are found in the soma of motoneurons both in carriers of TDP-43 mutations and in sporadic patients [25,39] but not in carriers of SOD1 mutations [31] and FUS mutations [42]. Based on this discrepancy, it has been suggested that the pathogenesis of ALS caused by mutant SOD1s is different from that in sporadic ALS and ALS provoked by TDP-43 mutations. Our present findings and the findings of Bosco et al. [5] contradict this proposition even though no details on TDP-43 or FUS were provided in the latter study. Regarding involvement of the proteins found mutated in FALS pedigrees in sporadic ALS, these proteins possibly occupy different steps or occur in parallel routes in pathogenic chains of events. The fact that intranuclear glial and motoneuron inclusions were seen in the two SBMA patients as well as in the familial non-SOD1 patients suggests that SOD1 might be involved in downstream events even in motoneuron disease induced by mutations in other genes [12].
In summary, our observation that misfolded SOD1 is present in the nucleus of motoneurons and astrocytes also implicates the nucleus as a potential site of SOD1 toxicity. More studies are needed to clarify the distribution of misfolded SOD1 in non-neuronal cells and to explore in more detail the modes of SOD1 toxicity at different sites.
Acknowledgments We thank
The authors declare that they have no conflict of interest.
We evaluated 35 cases of a mechanical approach to abdominal wall lifting, used in office-based gasless laparoscopic sterilization under local anesthesia. Lifting of the abdominal wall, using the camera trocar as an anchoring device and complemented by suprapubic lifting by means of a towel clamp, led to passive intra-abdominal air filling, giving sufficient space to identify, anesthetize, coagulate and cut the Fallopian tubes. Only mild sedation was necessary. All women walked to and from the operating room. All had successful tubal ligation. The overall satisfaction rate was 97%. The mechanical lifting moment was not painful. With the exception of one woman with failed tubal anesthesia, all women had a low mean pain score of 2.6 (VAS 0-10). No complications occurred except one wound infection. The costs were £ 1 / 4 of those of traditional laparoscopic sterilization and office hysteroscopic sterilization. This approach is effective for office-based laparoscopic sterilization. Room air, two strings and a needle replace active gas insufflation and narcosis.
Modern healthcare is expensive, especially surgery performed in a regular operating theatre. General anesthesia involves life-threatening risks and undesirable side-effects. Local anesthesia minimizes these drawbacks and can be administered by the surgeon. This promotes simplification.
Female sterilization in an office situation using laparoscopy with a low-pressure pneumoperitoneum administered under local anesthesia has been performed for more than 30 years (1), but has not gained widespread popularity. A survey of the American Association of Gynecologic Laparoscopists (2) showed that only 5.1% of its members performed officebased laparoscopy under local anesthesia. This is a low incidence for doctors with a special interest in laparoscopy. One explanation can be that active insufflation of gas gives too much pain and necessitates heavy conscious sedation and surveillance by specially trained personnel (3). Even a low-pressure pneumoperitoneum is painful. By contrast, a zero-pressure (gasless) pneumoperitoneum is painless. Officebased laparoscopy and gasless laparoscopy are not new procedures, but gasless laparoscopy under local anesthesia is a new procedure. Two articles have been published on gasless laparoscopic female sterilization but these operations were done under general anesthesia (4,5).
Carbon dioxide gas is cold, very dry, dissolves in water and provokes hypothermia, desiccation, tissue irritation, and acid-base and blood gas changes. By comparison, room air is warmer, humid, insoluble, not irritating and available everywhere. Gasless laparoscopy depends on mechanical lifting of the abdominal wall with a passive inflow of room air. An open access technique for trocar insertion is safer than a closed access technique and is therefore a better choice for office laparoscopy (1). The open access technique results in a 1.5-2 cm wide hole in the abdominal wall, so using a micro-laparoscope is meaningless. An ordinary laparoscope has advantages in focal distance and field of vision.
This was a prospective pilot study to evaluate a mechanical approach to lifting, used for officebased gasless laparoscopic sterilization. Women with a body mass index (BMI) < 30 kg/m 2 , no serious illnesses and no known abdominal adhesions were informed about general and local anesthesia during presterilization counseling sessions. All women without contraindications for office surgery chose local anesthesia and were given written information about the procedure. Between September 2003 and December 2005, 35 women were sterilized in a low resource setting situated one floor below a regular operating theatre. These women had a mean age of 39 years (range 27-48), BMI 24.1 kg/m 2 (19-30). Fourteen women had had a previous abdominal operation. All women but one had a normal-sized uterus. The procedure room had resuscitation equipment including oxygen and suction, an emergency tray with diazepam (5 mg/ml), atropine (1 mg/ml), catastrophic adrenaline (0.1 mg/ml) and a narcotic antidote, electrocardiography equipment, pulse oximeter and an alarm button.
Women had fasted for at least 6 hours and were premedicated orally with 200 mg ibuprofen and 1 g paracetamol/30 mg codeine phosphate about 1 hour before the start of surgery. After voiding, each patient walked to the procedure room where she was given intravenous fluid. Personnel included the surgeon, an assistant to arrange the instruments and a midwife to monitor the patient's blood pressure, pulse rate, respiratory rate, blood oxygen saturation, electrocardiogram, level of sedation and to give medications on request from the surgeon. Only the surgeon was dressed in sterile clothing. The patient was cleaned by the assistant and draped by the surgeon and mildly sedated with 5 mg diazepam and 25 mg meperidine: doses that could be repeated once. The muscle relaxant effect of diazepam was of value in the mechanical stretching of the abdominal wall. A mixture of 40 ml 1% lidocaine hydrochloride/ adrenaline and 60-80 ml 0.9% sterile NaCl was used for local anesthesia. The anterior region of the cervical portio was anesthetized and a Hulka forceps attached. The abdominal wall just beneath the umbilicus was anesthetized and a ‡ 12.5-mm trocar was inserted into the abdominal wall using an open access technique. The lifting technique used the camera trocar as an anchoring device in the abdominal wall. The open trocar gas inlet allowed a free inflow of room air. An operative (0 , 10 mm) laparoscope with a 6-mm working channel was used. The trocar/abdominal wall were lifted with a loop of polydioxanone suture (PDS # 1) snared around the shaft of the trocar with a hang knot and needledriven through the fascia, cutaneous tissue and skin in the lower end of the abdominal wall incision. The loop suture was attached to a horizontal metal arm mounted on the operating table and placed above the woman.
Padded shoulder supports were an important prerequisite, because if a woman were to slide downwards on the table she would be likely to get scared and tense her abdominal muscles to stay put. This would tend to press the intestines into the pelvis. However, despite the shoulder supports it was necessary for the women to lie with their pelvic region 10 cm outside the table. This position gave necessary sliding distance to prevent the Hulka forceps from being blocked by the table when it was placed in steep Trendelenburg position.
Mechanical lifting of the abdominal wall with the camera trocar as an anchoring device and with the trocar gas inlet open, led to passive filling of the abdomen with air. Lifting the skin and subcutaneous tissue 6-8 cm above the symphysis pubis in the same way using a towel clamp allowed sufficient air into the abdomen for laparoscopic sterilization. The combination of mechanical lifting, the Trendelenburg position and the forward rotation of the uterus created adequate intra-peritoneal space to identify, anesthetize, coagulate and cut the tubes. If the mechanical lifting procedure did not create sufficient space to identify the tubes, a small amount of additional room air (1) was insufflated actively using a rubber bulb. A towel clamp closing the upper end of the abdominal wall incision then secured air tightness. The mechanical lifting supported most of the weight of the abdominal wall and only a small volume of actively insufflated air and a low intra-abdominal pressure increase were needed to create additional space. After the tubes had been anesthetized, coagulated, divided with hook-scissors and the intra-peritoneal air been reduced to a minimum, the abdominal wall opening was closed. The women were able to sit up for a minute and then walk back to the recovery room.
All women received postoperative information from the surgeon or the midwife. They were given oral medication for 2 days (200 mg ibuprofen  3 and 1 g paracetamol/30 mg codeine phosphate  3) and a questionnaire to be sent back in 1 week (pain score, worst painful moment, satisfaction rate, complications, validity of the presterilization counseling session).
All women had a successful tubal ligation. The overall satisfaction rate was 97%. Two of the first 10 women in the study had a small amount of filtered room air insufflated actively, but when the camera trocar lifting procedure was complemented by suprapubic lifting, there was no need for active air filling because the intra-abdominal laparoscopic view was always good. No woman reported the mechanical lifting moment to be painful: one stating, 'It was like being lifted in the pants'. Thirty-four of 35 women were satisfied in that they answered 'yes' to the question of whether they would recommend the same operation to their best friend. One woman was not satisfied; the reason was much pain when one of the tubes was electro-coagulated. This was the only woman to report that the interprocedural pain was worse than was expected from the information given in the presterilization counseling. Per/ postoperative pain was measured using a visual 0-10 analog scale (VAS). The 34 satisfied women reported an average score of 2.6 (range 0-7.5) and three expressed no pain at all. The potentially painful moments during the operation were the different needle pinpricks, the Hulka forceps manipulations, unintentional rough touching of pelvic organs and anesthetic failure. One woman reported VAS 7.5 for the steep Trendelenburg position and one woman VAS 7.0 for the Hulka forceps application. Minor sedation was sufficient. All women walked to and from the operating room. No operation was converted to general anesthesia and there were no complications except for one wound infection. All women left for home after 1-5 hours and all submitted a completed outcome questionnaire.
In the first 10 operations, a 1.2-mm puncture needle was used for tubal anesthesia. This needle had a tendency to push the tube in front of itself rather than to penetrate. Later, a 0.4-mm needle was used. When using a 10-mm laparoscope, a ‡12.5 mm camera trocar with an open high-flow gas inlet is necessary for the free passive inflow of room air. A smaller trocar can reduce the inflow of air, and result in a negative intra-abdominal pressure, that limits the space for inspection and instrumentation and causes pain. The used mechanical lifting procedure has a short setup time (< 1-2 min), introduces no extra devices into the peritoneal cavity, causes no trauma to the peritoneal surface and does not interfere with surgical movements.
To further simplify the process, the surgeon can use an amnioscope (20 Â 200 Â 25 mm), which permits direct visual inspection and instrumentation of the tubal areas (Figure 1). Short instruments are then used for an optimal visual distance. Since 1993, the author has performed more than 200 female sterilizations using an amnioscope. The only drawback compared to using a laparoscope are a less comfortable working position for the surgeon and the fact that the personnel and patient cannot watch the procedure on a monitor. An amnioscope with or without an operative laparoscope is a good choice, that supports a high inflow of room air and to insert a second working instrument beside the laparoscope to displace obscuring loops of distended bowel, if any.
Using an amnioscope and operate under local anesthesia and mild sedation in an office setting conforms well to the statement of the WHO Task Force on Female Sterilization (6). 'The ideal female sterilization would involve a simple, easily learned, onetime procedure that could be accomplished under local anesthesia and involve a tubal occlusion technique that caused minimum damage. The procedure would be safe, have high efficacy, be readily accessible, and be personally and culturally acceptable. The cost for each procedure would be low and there would be minimal costs for the maintenance of equipment'. In developing countries, where minilaparotomy is a common approach, women would probably benefit from this amnioscope lifting procedure. Compared with minilaparotomy, it gives a much better view of the pelvic organs, an aesthetically more acceptable scar, a shorter recovery period and less complications/ complaints (6). Using an ordinary anesthesia frame in the lifting process, with the horizontal arm draped in a sterile sleeve, makes the procedure progress even more smoothly. In developing countries, where resources are limited for the purchase and maintenance of more sophisticated laparoscopic equipment, this amnioscope lifting procedure is truly a cheaper and safer option for female sterilization than the traditional laparoscopic procedure presently used in the developed world.
Office laparoscopic sterilization requires, in contrast to office hysteroscopic sterilization, no scheduling of surgery according to the woman's menstrual cycle and has an almost 100% first-attempt success rate, an immediate effect on fertility, no need for tubal patency control, no material costs, a reversal success rate of 55-75% (7), an unchanged possibility of IVF and can be performed within 48 hours of delivery.
Office laparoscopy is claimed to cost much less than traditional laparoscopy. In one study, there was an almost 80% reduction in costs (8), which is in agreement with the present study. Our total costs in 2006 were calculated to be NOK 2,895 (US $446), which represents a 75% reduction. Gasless laparoscopic sterilization and hysteroscopic sterilization (Essure Ò device) can both be done in an officebased setting with the same OR-team and operation time. The Essure procedure is however approximately US $1,575 more expensive due to additional costs for the Essure Ò devices (ESS305, Conceptus, Inc., USA, Retail Price: $1,299) and for tubal occlusion control (HSG US $275) and is contraindicated within 6 weeks after delivery and in women with hypersensitivity to nickel or allergy to contrast media.
A comparative study/literature review of hysteroscopic sterilization versus laparoscopic tubal sterilization (9) showed overall standard complication rates for laparoscopic sterilization of 0.8-0.9% (6,9) and a major complication rate for hysteroscopic sterilization of 3.2% (9). Moreover, correct use of local anesthesia removes the single greatest source of risk in conventional laparoscopic sterilization procedures, general anesthesia.
A gastight mechanical lifting procedure, that creates adequate intra-peritoneal space for female sterilization under local anesthesia, is also useful in surgery under general anesthesia. Therefore, in conventional gas laparoscopy, the author also uses the described lifting technique. Such lift-assisted laparoscopy makes surgery in low gas pressure (1-6 mm Hg) possible with sustained optimal or adequate view (10) and, in case of need, to take temporary measures in a 'gasless' condition. An immediate shift between low, standard and zero gas pressure is possible. With no gas pressure, conventional open surgery instruments such as clamps, scissors and powerful suction devices can be used and with standard gas pressure, the complementary abdominal wall lifting means a 'double' outcome in forming the intraabdominal space. A special slit-trocar facilitates the shifting maneuver between the gas-based and gasless technique (Figure 2). The finding that the centrally positioned camera trocar, except for its normal function as a gastight sleeve for the laparoscope, is also a perfect anchoring device for mechanical lifting, is an enhancement for laparoscopic surgery. 'Yesterdays' gasless laparoscopy must not be confused with liftassisted laparoscopy. The European Association for Endoscopic Surgery states that 'gasless laparoscopy has no clinically relevant advantages compared to low-pressure (5-7 mm Hg) pneumoperitoneum' (11). There may be two exceptions to this: laparoscopy under local anesthesia and lift-assisted laparoscopy under general anesthesia if it is required to use a conventional open surgery instrument. All patients benefit from low-pressure gas laparoscopy and especially high-risk patients, including pregnant women.
The evaluated mechanical lifting technique can be applied effectively in office laparoscopic sterilization, even to overweight women. Risks of general anesthesia and active gas insufflation are eliminated and room air, two strings and a needle, replace a CO 2 gas filling machine and narcosis.
The author has invented the slit-trocar.
Background. Cutaneous metastases may cause considerable discomfort as a consequence of ulceration, oozing, bleeding and pain. Electrochemotherapy has proven to be highly effective in the treatment of cutaneous metastases. Electrochemotherapy utilises pulses of electricity to increase the permeability of the cell membrane and thereby augment the effect of chemotherapy. For the drug bleomycin, the effect is enhanced several hundred-fold, enabling once-only treatment. The primary endpoint of this study is to evaluate the effi cacy of electrochemotherapy as a palliative treatment. Methods. This phase II study is a collaboration between two centres, one in Denmark and the other in the UK. Patients with cutaneous metastases of any histology were included. Bleomycin was administered intratumourally or intravenously followed by application of electric pulses to the tumour site. Results. Fifty-two patients were included. Complete and partial response rate was 68% and 18%, respectively, for cutaneous metastases Ͻ 3 cm and 8% and 23%, respectively, for cutaneous metastases Ͼ 3 cm. Treatment was well-tolerated by patients, including the elderly, and no serious adverse events were observed. Conclusions. ECT is an effi cient and safe treatment and clinicians should not hesitate to use it even in the elderly.
A cutaneous metastasis can be defi ned as " a neoplastic lesion arising from another neoplasm with which there is no longer continuity " [1]. Cutaneous metastases account for 0.7% to 9% of all metastases [2]. Breast cancer accounts for 51% of the total cases of cutaneous metastases, while malignant melanoma accounts for 18% [3].
The management of cutaneous metastases often presents a challenge for the clinician as they may be widespread and may recur after radiotherapy or chemotherapy. In some cases, patients may have stable disease in sites other than the skin, and clinicians may be reluctant to use systemic chemotherapy for the skin metastases alone.
Electrochemotherapy (ECT) is a rapidly emerging and effective treatment option for cutaneous metastases from malignant tumours [4 -6]. ECT uses local application of short duration, electric pulses directly to the tumour cells via an electrode, causing destabilisation of the cell membrane and thereby making it transiently permeable (electropermeabilisation -Figure 1) [7]. Bleomycin, a chemotherapeutic agent used in the treatment of cancer, is under normal conditions unable to freely diffuse through the plasma membrane. However, electropermeabilisation allows this otherwise poorly permeating agent to enter the cell cytosol, thereby greatly increasing its concentration within the tumour cell (Figure 2). In high concentrations, such as those achieved with ECT, bleomycin can cause cell death within a few minutes [8]. Preclinical studies have demonstrated a 300-to 700-fold increase in bleomycin cytotoxicity using this method of drug delivery [9 -11].
ECT was originally employed for treatment of metastatic head and neck cancer [12] and has since been used in the treatment of cutaneous metastases from tumours independent of histology [5,13,14]. ECT can be used where surgery is not an option and is also effi cient in chemotherapy-resistant and radiotherapy-resistant lesions [5,15]. Treatment may provide palliation particularly where there is pain or bleeding from cutaneous metastases [16]. It is well tolerated with few side-effects, allows for immediate recovery and can be repeated [4,5]. In 2006 a European study was published (the ESOPE study) [5] and with that the standard operating procedures [17] which describe the ECT procedures in detail.
The present study aims at continuing the exploration of ECT as a highly effective treatment in order to improve and evaluate its benefi ts. To this end, we created the International Network for Sharing Practice in Elec-troChemoTherapy (INSPECT) database, the purpose of which is to gather, share and publish clinical data and experience. This is the fi rst report from this network.
Patients were recruited consecutively at two institutions: Copenhagen University Hospital Herlev, Denmark and James Cook University Hospital, Middlesbrough, UK. The primary endpoint was response rate. Secondary endpoints included safety and response rate according to size.
Patients at Herlev Hospital with cutaneous metastases, for whom no further surgery or conventional treatment was feasible, could be offered treatment with ECT within the framework of a non-randomised phase II study. Approval was granted by the local ethics committee and the Danish Medicines Agency.
Approval in the UK was granted in 2007 by the clinical effectiveness subcommittee for treatment of metastatic skin and subcutaneous lesions, palliation of bleeding or painful lesions and primary treatment of cancers not amenable to surgical excision or conventional treatments.
Patients fulfi lling the inclusion criteria were sequentially enrolled and all patients signed informed consent before inclusion.
Patients eligible for inclusion had histologically proven malignant cancer with measurable cutaneous or subcutaneous tumour nodules suitable for application of electric pulses. Patients had been offered standard treatment options, were Ն 18 years old, Figure 1. The electroporation procedure: A. Electroporation occurs when an applied external fi eld exceeds the capacity of the cell membrane. The formation of permeable areas happens in the frame of less than a second whereas resealing happens over minutes. As the resting transmembrane potential is negative on the inside respective to the outside, the fi rst part of the membrane that will be permeabilised is the pole facing the positive electrode. The positive electrode should be imagined in the left of the picture and the negative electrode on the right. Pulses were delivered to a cell suspended in medium containing propidium iodide and after the pulses propidium iodide is trapped within the cells [9]. B. The application of pulses to skin tumours must be preceded by local or general anaesthesia, in local anaesthesia the lidocain is injected around the metastasis. C. The cliniporator equipment allows monitoring of voltage and current during the pulse. D. A treatment situation is shown where a patient is receiving local injection of bleomycin followed by application of pulses under local anaesthesia. The application of pulses lasts only a few minutes in total.
had ECOG performance status Յ 2, had an expected life expectancy of at least three months and, where appropriate, were using adequate contraception. A platelet count Ն 50 mia/l was required, with a prothrombin time Յ 40 s and an activated partial thromboplastin time in the normal range.
Patients were ineligible if they had previously had allergic reactions to bleomycin or to any of the components required for anaesthesia, if the cumulative dose of 250 mg bleomycin/m 2 (400.000 IU bleomycin/m 2 ) had previously been exceeded, had chronic renal dysfunction (serum creatinine Ͼ 150 μ mol/l) or acute lung infection.
Follow-up was planned for up to six months. Patients who had been started on systemic antineoplastic treatment after ECT were excluded from the study at that time.
The ECT sessions were performed based on the standard operating procedures for electrochemotherapy [17]. Bleomycin was administrated either intratumourally (i.t.) or intravenously (i.v). The decision to treat either i.t. or i.v. was based on the number of cutaneous metastases to be treated and the size of the metastases.
General anaesthesia, was preferred for multiple metastases, large metastases ( Ͼ 3 cm), metastases adhering to the periosteum or situated in sensitive regions (e.g. face and scalp), and in accordance with patient preference.
Intratumoural treatment (for small or few cutaneous metastases). Bleomycin was injected into the cutaneous metastases according to size. Pulses were delivered after administration of the drug (all pulses must be administrated within 10 minutes of bleomycin injection).
Intravenous treatment (for large or many cutaneous metastases). Bleomycin was injected intravenously (15000 IU/m 2 ϭ 15 U/m 2 which is approximately equal to 8.5 mg/m 2 bleomycin depending on the activity of the drug and the manufacturer). Pulses were delivered 8 -28 minutes following injection when bleomycin is known to be present in high concentration in the tumour [18,19].
Local anaesthesia (for small or few cutaneous metastases). Lidocaine with epinephrine was injected around the metastasis (Figure 1). The electrode was placed in and around the metastasis and the pulses administered. In the middle panel the electric pulses are subsequently applied, cells are permeabilised and the drug enters. In the left panel the cells reseal after a few minutes and the extracellular drug is washed out while the internalised molecules remain trapped intracellularly.
Under local anaesthesia, patients do not feel the insertion of the electrode needles but do feel a brief local muscle contraction upon administration of the electrical impulse.
Anaesthesia for electrochemotherapy (ECT) was tailored to the patient ' s condition, the position of the lesions, the extent of treatment and the special considerations pertaining to general anaesthesia and the use of bleomycin [20].
Depending on the clinician ' s choice, one of the following electrodes was used: 1) Type I electrodes: two plates with a 6 mm gap between the plates; 2) Type II electrodes: two parallel rows of needles with 4 mm between the rows; 3) Type III electrodes: a hexagonal array of electrodes with 7.9 mm between the needles.
Electric pulses (eight pulses of 100 μ s duration) were delivered using a square wave electroporator (IGEA, Carpi, Italy). The applied voltage was 1.3 kV/ cm for plate electrodes and 1.0 kV/cm for needle electrodes, i.e. for the type II needle electrode with a 4 mm gap between the needles the applied voltage was 400 V. For type I and II electrodes, the pulses are applied with 1 Hz or 5 kHz, whereas for type III electrodes, pulses can only be applied with 5 kHz. Electrodes were single use. The duration of the procedure was recorded from the start of the bleomycin injection to the completed delivery of the last pulse. After ECT, the treated metastases were covered with standard dressings where necessary.
Evaluation of the tumour response was by measurement of the extension or regression of the treated metastases. This was documented using digital photography. A maximum of seven cutaneous metastases per patient were registered as target lesions in order not to skew data by inclusion of patients with very large numbers of cutaneous metastases. The response was registered for each target lesion and new cutaneous metastases were not considered in response evaluation but could be treated in a second ECT session. The response rate was evaluated similarly to the Response Evaluation Criteria in Solid Tumours (RECIST version 1.0) [21]: Complete response (CR) was defi ned as disappearance of the target lesion; partial response (PR) with at least 30% decrease in the diameter of the target lesion; progressive disease (PD) with at least 20% increase in the diameter of the target lesion and stable disease (SD) with neither suffi cient shrinkage to qualify for PR or suffi cient increase to qualify for PD. In some cases with exophytic ulcerated tumours evaluation was not possible due to crust formation (Figure 3).
Safety was reported in the form of adverse events using Common Toxicity Criteria version 3.0. Patients were asked if they would potentially agree for another session as a measure of how patients felt about the treatment procedure.
Descriptive methods were employed for statistical analysis using SPSS 13.0.
Patients were followed for six months, but excluded from further evaluation if new antineoplastic treatment was started within the six months.
All patients treated with ECT were included for evaluation of effi cacy and safety.
A total of 52 patients with cutaneous metastases were enrolled between June 2007 and April 2010. Table I presents patient characteristics at baseline. Fifty-one patients underwent electrochemotherapy for 196 cutaneous metastases from primarily malignant melanoma or breast cancer (Figure 3). In one patient with malignant melanoma in the head region, treatment was not given due to poor lung function.
Forty-fi ve patients were evaluable for safety and toxicity, and 24 patients with 97 cutaneous metastases had a follow-up of 60 days or more. Eleven patients received a second treatment with ECT.
The median diameter of the cutaneous metastases was 12 mm ranging from 1 mm to 200 mm. Locations of the cutaneous metastases are presented in Table I.
Treatment data are listed in Table II.
Patients treated with i.v. or i.t. administration of bleomycin had a median number of three treated cutaneous metastases. The median size of the cutaneous metastases treated with i.v. bleomycin was 10 mm (range 1 -200 mm) and for i.t. 9 mm (range 1 -50 mm).
Of the 51 treatments, 28 (55%) were performed under general anaesthesia and 23 (45%) were performed under local anaesthesia (see Table III). There was no statistical difference between number of nodules per patient and choice of anaesthesia.
The electrodes used for treatment were as follows: type II electrodes for 119 (61%) of the cutaneous metastases; type III electrodes for 47 (24%); type I electrodes for 21 (11%); both type I and II electrodes for two (1%) and both type II and III electrodes for seven (4%).
Data on duration were available for 42 procedures. The median duration of a treatment session from start of bleomycin administration to last pulse delivered was 20 min (range 5 min to 1 hour and 9 min) for local anaesthesia and 25 min (range 11 min to 1 hour 27 min) for general anaesthesia (see Table II).
No serious adverse events (SAE) were observed. Reported adverse events were fl u-like symptoms one to two days after treatment (fi ve patients, 10%), pain in the treated area one to two days after treatment (fi ve patients, 10%), ulceration of treated area (two patients, 4%), cough (one patient, 2%), allergic skin reaction (one patient, 2%) and anxiety (one patient, 2%). There was no CTC grade 3 or 4 toxicity. Most side-effects were seen when treated under general anaesthesia with systemic administration of bleomycin.
Six patients were lost to follow-up before evaluation due to systemic disease progression. Forty fi ve patients with 162 treated cutaneous metastases had a median follow-up of 79 days (range 8 -180), and 24 patients with 97 nodules had a follow-up Ͼ 60 days. Responses are presented in Table II.
For patients with a follow-up Ͼ 60 days, CR was observed in 58 (60%) metastases, PR was observed in 18 (19%) metastases, SD was observed in 11 (11%) metastases and PD in seven (7%) metastases. Response was not evaluable in three (3%) metastases.
Table I. Patients ' characteristics at baseline. Patients Total (N ϭ 52) Patients (%) Patients Herlev, N ϭ 30 Patients Middlesbrough N ϭ 22 Median age in years (range) 69.6 (38.9 -94.7) 72.1 (53 -89.8) 68.3 (38.9 -94.7) Age distribution 1 80 ϩ 11 11% 70 ϩ 25 48% 60 ϩ 44 85% 50 ϩ 48 92% Sex Female 35 67% 24 13 Male 17 33% 6 8 ECOG 2 performance status 0 35 67% 58% 84% 1 12 23% 33% 5% 2 5 10% 9% 11% Previous Treatment Surgery 42 81% 71% 95% Radiotherapy 20 38% 42% 21% Chemotherapy 21 40% 42% 37% No previous treatment 8 15% 21% 0% Number of metastases treated pr. patient 3 Median (range) 3 (1 -7) 3 (1 -7) 4 (1 -7) Diagnosis 4 Malignant Melanoma 21 40% 36% 47% Breast Cancer 15 29% 33% 21% Adenocarcinoma (other than breast) 5 10% 15% 0% Basocellular Carcinoma 5 10% 9% 11% Squamous Cell Carcinoma 3 6% 6% 5% Other 3 5% 0% 16% Location of metastasis Chest 79 40 40 41 Lower limbs 54 28 22 36 Head and Neck 30 15 22 6 Scalp 21 11 12 9 Upper limbs 6 3 4 1 Abdomen 5 3 1 5 Back 1 1 0 1 Size of metastases Median diameter in mm (range) 12 (1 -200) 15 (2 -200) 5 (1 -140) Յ 3 cm 138 Ͼ 3 cm 24 1 Number of patients at or above a given age. 2 ECOG ϭ Eastern Cooperative Oncology Group. 3 Maximum seven metastases per patient registered, for 51 patients. 4 No signifi cant differences among distribution of diagnosis of primary tumour among centres could be observed (p ϭ 0).
Of the 51 patients treated, 46 (90%) would agree to another treatment, four (8%) would not agree to another treatment and one patient is not accounted for.
Cutaneous metastases or recurrent malignant disease in the skin, particularly after treatment of malignant melanoma, head and neck carcinoma or breast cancer, is often diffi cult to manage. Patients have often received multimodal treatment with surgery, radiotherapy and chemotherapy and are faced with obviously progressing disease. The uncontrolled cutaneous metastases can, in many ways, adversely affect self-esteem and body image. The cutaneous metastases and the treatment of the cutaneous metastases will seldom affect life expectancy, but may be very important for the patient ' s quality of life.
Electrochemotherapy is a method where the combination of electric pulses and bleomycin increases the cytotoxicity of bleomycin 300 -700 times [9 -11]. When electric pulses are delivered to tissue in the presence of bleomycin, the cell membrane becomes permeable and bleomycin enters the cell where it is trapped. The large increase in bleomycin cytotoxicity makes it possible to do " once-only " treatment suitable for the palliative patient. The clinical effectiveness of ECT was fi rst
Table II. Treatment data and response. TREATMENT DATA All Patients (n ϭ 51) All Patients (%) Herlev Middlesbrough Chemotherapy Bleomycin I.T. 21 41% 41% 42% Bleomycin I.V. 30 59% 59% 58% Anaesthesia Local 23 45% 50% 37% General 28 55% 50% 63% ECT session duration 1 , hours:minutes Median (range) (hh-mm) 00:16 (00:05 -01:27) 00:29 (00:08 -01:27) 00:18 (00:05 -00:35) Would agree for another session yes 46 90% 87% 90% no 4 8% 13% 5% no answer 1 2% 0 5% RESPONSE All Metastases (n ϭ 97) 2 Metastases (%) Herlev Middlesbrough Response for registered metastases 3 CR 58 60% 54% 68% PR 18 19% 20% 17% SD 11 11% 18% 2% PD 7 7% 4% 12% Not evaluable 3 3% 5% 0% Time from treatment to CR (days) Median (range) (days 47 (16 -110) 41 (16 -110) 63 (38 -100) Size of metastases Յ 30 mm (n ϭ 84) CR 57 68% PR 15 18% SD 5 6% PD 5 6% Not evaluable 2 2% Size of metastases Ͼ 30 mm (n ϭ 13) CR 1 8% PR 3 23% SD 6 46% PD 2 15% Not evaluable 1 8%
1 Data available for 42 patients, the time is from start of chemotherapy administration till the last pulse was given. This means it does not include time anaesthesia. One patient with the procedure lasting 1 hour and 9 min was treated in local anaesthesia with i.t. injection of bleomycin had three nodules where treatment of the fi rst nodule was fi nished before the anaesthesia of the next nodule began. One patient with the procedure lasting 1 hour and 27 min was treated in general anaesthesia with i.t. injection of bleomycin had seven nodules where treatment of the fi rst nodule was fi nished before injection of bleomycin in the next nodule. This explains why some procedures lasted longer than one would expect. 2 24 patients with 97 metastases with a follow-up Ͼ 60 days. 3 Maximum seven metastases per patient registered.
demonstrated in head and neck squamous cell tumour nodules in 1991 [22]. Subsequent clinical investigation has shown that ECT using bleomycin is also a feasible and effective treatment for cutaneous and subcutaneous metastases of other malignancies [14]. The ESOPE study in 2006 [5] produced standard operating procedures for ECT treatment (including dosage, pulse parameters, electric pulse generators and electrodes), pain control and indications for treatment. The ESOPE study demonstrated that electrochemotherapy is an easy, highly effective and safe treatment for small ( Յ 3 cm) cutaneous or subcutaneous metastases of various malignancies. The objective response rate after one treatment was 85%. Similar results have been demonstrated by Campana et al. [13] and additionally by many case-reports and smaller studies [23 -25].
In the present study we have tried to manage cutaneous metastases with ECT as a routine procedure in two cancer centres. The primary endpoint of this study was response rate.
Fifty one patients from Denmark and the UK were treated for 192 cutaneous metastases with ECT -the majority with either malignant melanoma or breast cancer. This is in agreement with breast cancer being the most common malignancy with cutaneous metastases in women and malignant melanoma being the most common in men [3].
ECT treatment in this study was provided as a palliative procedure to patients with performance status Յ 2. A broad spectrum of patients was included, which is refl ected in six patients lost to follow-up before any evaluation and only 24 patients having a follow-up Ͼ 60 days. Some patients travelled long distances to reach the centre offering ECT which hindered follow-up, and some patients had systemic progression during follow-up and were offered other antineoplastic treatment. These factors may explain the relative high rate of patients lost to follow-up before any evaluation and only 47% of patients having a follow-up period Ͼ 60 days.
The classic RECIST criteria [21] were unsuitable as tumour assessment in evaluation of ECT treatment as RECIST includes measurable lesions in other organs if present, a minimum size of 1 cm and a maximum fi ve lesions per organ. Also the aim of ECT is local and not systemic control. Instead, the defi nitions of CR, PR, PD and SD from RECIST were adapted and seven cutaneous metastases were registered as target lesions. This seems a feasible way to evaluate ECT.
In this study, objective response rate (OR) for patients with a follow-up period of Ͼ 60 days was 86% for cutaneous metastases Յ 3 cm and 31% for cutaneous metastases Ͼ 3 cm. The metastases were divided into smaller or larger than 3 cm to enable comparison with previous studies. The response rate for the cutaneous metastases Յ 3 cm is similar to the ESOPE [5] and other studies [13,23,24], whereas for larger metastases, the OR is considerably lower, which is in agreement with previous observations [13]. In patients with large volume disease, the purpose is not necessarily to eradicate the cutaneous metastases, but to obtain palliative relief in terms of decreased odour, exudate and bleeding. Therefore, SD can be the aim of ECT treatment for large volume disease. However, in the management of small cutaneous metastases, control and disappearance of cutaneous metastases can be the aim of treatment. Due to the low incidence of complications, treatment can be repeated several times in order maintain local control or obtain control if not achieved by the fi rst treatment. In this study 11 of the 51 patients were resubmitted for treatment, either due to new metastases or progression of previously treated metastases, with no SAE ' s observed.
For patients with cutaneous metastases, local control during their remaining life period is the goal of treatment. Regional and local techniques such as palliative surgery, re-irradiation, hyperthermia, isolated limb perfusion and isolated limb infusion can be offered to patients with cutaneous metastases in order to provide local symptom control. When offering treatment to patients, the risk of complications and toxicity should always be carefully addressed and the likely benefi t should always be compared with the risks to the patient. Electrochemotherapy offers a minimally invasive local treatment with swift symptomatic relief and few sideeffects. ECT can also, as shown in this study with 48% of patients being Ͼ 70 years, be offered to elderly patients for whom other treatments may not be a possibility.
In this study, treatment was performed at two different centres -a department of plastic surgery and a department of oncology -with similar results.
Table III. Choice of anaesthesia according to location of metastases and size. Local anaesthesia General anaesthesia Location of metastases Chest 23 32% 56 46% Lower limbs 31 42% 23 19% Head and Neck 11 15% 19 15% Scalp 4 5% 17 14% Upper limbs 3 4% 3 2% Abdomen 1 1% 4 3% Back 0 0 1 1% Size of metastases Median (range) (mm) 7.5 (1 -60) 10 (1 -200) Number of metastases per patient Median (range) 3 4
This demonstrates that the treatment functions well in different types of units. ECT may easily be implemented as limited training is needed.
In conclusion, our results in concordance with previous studies suggest ECT is an effi cient treatment that may improve quality of life in patients with metastatic disease and clinicians should not hesitate to use it even for elderly patients. ECT is simple to administer, and can therefore be implemented by smaller hospital units with resultant benefi ts for patients.
In our two centres, we concurrently found ECT to be an excellent treatment choice for the patient suffering from cutaneous metastasis where other treatments have failed.
We would recommend that more centres offer ECT and that referral for this once-only and simple treatment should be considered.
Julie Gehl is a research fellow of the
The authors report no confl icts of interest. The authors alone are responsible for the content and writing of the paper.
In 2001, we saw a 54-year-old woman with destructive seronegative rheumatoid arthritis (RA). Because of secondary osteoarthritis, she had received a replacement of the right hip and of both knees, and a triple arthrodesis of the right foot. Furthermore, she had chronic obstructive pulmonary disease (COPD) and bronchiectasia.
She complained of severe pain and a sensation of heavy pressure at the site of the manubrio-sternal joint (MSJ), which had developed at the beginning of 2000, making coughing very difficult. She had a tender swelling at the same location on the sternum and we saw a dorsal dislocation of the manubrium. A severe thoracic kyphosis was also observed. A radiograph and CT showed a luxation of the MSJ (Figure ). The patient also required a hemiarthroplasty of the left shoulder, so we planned to perform an arthrodesis of the MSJ in the same session.
Hemiarthroplasty of the left shoulder was first performed and then the arthrodesis of the sternum was started. However, during this procedure, on trying to place the manubrium back into position, we observed that there would still be a bone gap of about 1 cm and too much tension on the manubrium. We decided, therefore, to do a resection arthroplasty. Approximately 1 cm of the manubrium and 1 cm of the sternum was resected and the gap was closed by the soft tissue layers, which had been displaced. The skin was closed and a deep drain was left in.
During the first 2 years after surgery, she had no pain or sensation of pressure on the sternum, and the swelling had almost disappeared. She could sit in a more relaxed position, which made coughing easier. Later on, the swelling slowly progressed but the patient still did not experience any pain. At 7 years, physical examination revealed an eminent swelling of the sternum but without any tenderness. The patient was very satisfied with the result. Radiography showed that the sternum was positioned 4 cm anteriorly to the manubrium.
In 1950, Bogdan and Clark first described 5 cases of painful rheumatoid involvement of the MSJ. In 2 cases, the symptoms disappeared spontaneously after just 1 month. In 1 case, the pain was relieved by aspirin and in the other 2 cases no treatment was mentioned. In 1979, Rapoport et al. (1979) described an RA patient with a painful anterior subluxation of the sternal body on the manubrium. No treatment was discussed, however. In 1981, Wiseman described a patient with rheumatoid arthritis with posterior dislocation of the manubrium on the sternum, and pointed out that complaints in this area could be the result of a dislocation. Khong and Rooney (1982) reviewed approximately 400 patients with classical RA. 10 of them had a subluxation of the MSJ, all of whom had substantial swelling over the MSJ with crepitus and deformity. Erosions and subluxation were found on lateral radiographs of the sternum in all 10 cases. The authors referred to Laitinen et al. (1970) and Kormano (1970) who had also reported that the MSJ is often involved in RA, but seldom leads to significant clinical problems. They did not mention the exact symptoms of the patients with subluxation of the MSJ, and nothing was written about the patients' desire for treatment of their symptoms, if this should be available. Rapoport et al. (1979), Holt andRooney (1980), Wiseman (1981), Khong and Rooney (1982), and Kelly et al. (1986) described 12 RA patients in total with thoracic kyphosis and MSJ dislocation. In the patients they described, 11 had severe kyphosisas in our patient. Kelly et al. (1986) described the anatomical development of MSJ as the reason that it can be involved in RA, and described the role of kyphosis in transmitting force via the first rib to the manubrium, which can lead to dislocation.
Conclusion: The saccular duct and endolymphatic sinus run in the bony groove, before reaching the orifice of the vestibular aqueduct. We first clinically visualized this sulciform groove using three-dimensional (3D) cone beam CT images. This strategy can be useful to assess the condition of the saccular duct and endolymphatic sinus concerning the longitudinal flow system of endolymph. Objective: To assess the saccular duct and endolymphatic sinus in the endolymphatic system in order to advance clinical studies on inner ear dysfunction. Methods: The sulciform groove of the saccular duct and endolymphatic sinus of human subjects was analyzed by cone beam CT and compared with that of a cadaver. Results: We could obtain reconstructed 3D CT images of the sulciform groove of the saccular duct and endolymphatic sinus using several CT window levels.
The saccule connects with two ducts, the reuniting and saccular ducts, which are respectively the paths of endolymph. These paths are important in the longitudinal flow theory [1], whereby the endolymph goes directly from the cochlea to the endolymphatic sac. This theory explains the cause of Meniere's disease, which exhibits endolymphatic hydrops [2,3]. From this viewpoint, the vestibular aqueduct that follows the saccular duct was initially clinically visualized and used as a diagnostic tool for Meniere's disease [4,5].
We also clinically analyzed the reuniting duct by visualizing it with three-dimensional computed tomography (3D CT) to diagnose Meniere's disease [6][7][8]. However, the saccular duct has not been clinically visualized. The reason is that the saccular duct is membranous, facing endolymph inside and perilymph outside, and so CT and MRI have limitations for visualizing it. The saccular duct becomes endolymphatic sinus running in the bony sulciform groove before ending in the vestibular aqueduct, which may be visualized in 3D CT images as a bony substance, like the bony groove of the reuniting duct (YT groove) in a previous study [6,7].
In the present study, we investigated whether the saccular duct and endolymphatic sinus could be demonstrated on clinical images using cone beam 3D CT.
We investigated the saccular duct and endolymphatic sinus and their sulciform groove in the temporal bones of bilaterally healthy ears of 12 controls (6 males and 6 females; mean age 58.4 years, range 35-79 years), which had been included in a previous study [8] employing cone beam CT (3D Accuitomo; J. Morita Mfg Corp., Kyoto, Japan) and the temporal bone of a cadaver donated to our medical university of anatomy, with the consent of the deceased and our university, to compare the findings in human subjects with those in the cadaver.
CT images were investigated using cone beam CT to obtain images under the following conditions: 80 kV, 6 mA, voxel 0.125 Â 0.125 Â 0.125 mm, slice thickness 0.5 mm. CT images of this region were taken as a column with a diameter and height of 6 cm. Reconstructed 3D images of the inner ear were obtained using rendering software (IVIEW; J. Morita Mfg Corp.) using a perspective view with a viewing angle of 15 and a 0.25 mm voxel (0.25 Â 0.25 Â 0.25 mm).
To obtain 3D CT images of the saccular duct and endolymphatic sinus in accordance with cadaver specimens, we adopted the landmarks reported in a previous study to reduce discrepancies due to rendering effect [7]. 3D CT images are manipulated so that the saccule, YT groove, and vestibular portion of the posterior semicircular canal and ridge of the cochlea just sloping down to the vestibule in front are all viewed in one frame. With this view, the bony sulciform groove of the saccular duct and endolymphatic sinus is visible. Then, this image can be manipulated through pitch or yaw to yield a view from directly above the sulciform groove to decrease the artifact by rendering effects. Also, several CT window levels were used in one view to visualize the groove from the surrounding architecture. Approval was obtained from the ethics committee of Osaka City University Graduate School of Medicine.
A sulciform groove from the saccular fossa to the orifice of the vestibular aqueduct in which the saccular duct and endolymphatic sinus run could be localized in the cadaver specimen (Figure 1A-D). The sulciform groove of the endolymphatic sinus was distinct but that of the saccular duct was unclear (Figures 1C and 2A). Employing a direction in which the saccule and saccular duct and endolymphatic sinus can be seen from directly above, the groove-like 3D CT image of the saccular duct and endolymphatic sinus were consistent with macroscopic findings (Figure 2A-C). This groove-like image was made up of two colors when CT was imaged with these bone and soft tissue CT window levels (Figure 2C). Further, this image was consistent with macroscopic cadaveric findings because there was a torn membranous labyrinth around the groove that could be imaged as soft tissue by CT.
Groove-like 3D CT images of the saccular duct and endolymphatic sinus were obtained in all 12 human subjects. The aspects of the groove were similar among the 12 healthy subjects (Figures 3 and 4).
The images of the groove reconstructed using several CT window levels such as a high or low density of bone revealed more accurate images (Figure 3B and C), especially the lumen of the groove, which was more manifested three-dimensionally compared with single CT window levels (Figure 3A-C).
The vestibular aqueduct has been assessed in relation to the etiology of Meniere's disease, where the endolymphatic duct runs and reaches the endolymphatic sac. This clinical strategy is based on the fact that idiopathic endolymphatic hydrops, a pathological entity of Meniere's disease, is caused by the blockage of endolymph in the endolymphatic duct or sac [9,10].
However, there are few reports on the saccular duct and endolymphatic sinus despite these sites and the vestibular aqueduct being directly connected.
Most saccular duct and endolymphatic sinus studies are limited to animal experiments or specimens of human temporal bone [3,10], because the saccular duct and endolymphatic sinus are too small and membranous to assess clinically. The saccular duct and endolymphatic sinus are not a bony groove themselves but run in the bony groove, which could be visualized clinically using 3D CT. When the saccular duct or endolymphatic sinus is involved as a lesion in inner ear disease, their bony grooves may reflect the lesional effect. A similar strategy was adopted to assess the reuniting duct to visualize the YT groove in a previous study [6,7].
The rendering strategy for 3D CT analysis sometimes leads to a misunderstanding [7]. The CT image of the cadaver with a double CT window level such as bone and soft tissue showed the groove and its circumscribing torn membranous labyrinth, in which the image was consistent with the macroscopic cadaverbased findings. Therefore, we judged that our strategy could markedly reduce the negative effect of rendering. As precise differentiation of the saccular duct from the endolymphatic sinus was difficult when employing 3D CT images, we assessed both portions as one unit in the present study.
In human cases, the CT images of the sulciform groove of the saccular duct and endolymphatic sinus from directly above using different bone CT window levels on one image were consistent with cadaver specimen findings, and much more informative than the images employing a single bone CT level.
We previously reported the importance of assessing the reuniting duct, which is frequently involved in Meniere's lesions, and dislodged otoconia from the saccule into the reuniting duct [8]. As the saccular duct directly connects with the saccule, we cannot deny the possibility that dislodged otoconia from the saccule may enter the saccular duct or endolymphatic sinus in a similar manner to that in the reuniting duct.
Investigating the saccular duct and endolymphatic sinus will provide novel information regarding inner ear diseases caused by longitudinal flow disturbances of endolymph, especially Meniere's disease. We will also further investigate the role of these architectures.
We greatly appreciate the valuable technical assistance of
The authors report no conflicts of interest. The authors alone are responsible for the content and writing of the paper.
Aim: To determine whether the size and shape of the placental surface predict blood pressure in childhood.
We studied blood pressure in 471 nine-year-old Indian children whose placental length, breadth and weight were measured in a prospective birth cohort study.
Results: In the daughters of short mothers (<median height), systolic blood pressure (SBP) rose as placental breadth increased (b = 0.69 mmHg ⁄ cm, p = 0.05) and as the ratio of placental surface area to birthweight increased (p = 0.0003). In the daughters of tall mothers, SBP rose as the difference between placental length and breadth increased (b = 1.40 mmHg ⁄ cm, p = 0.007), that is as the surface became more oval. Among boys, associations with placental size were only statistically significant after adjusting for current BMI and height. After adjustment, SBP rose as placental breadth, area and weight decreased (for breadth b = )0.68 mmHg ⁄ cm, p < 0.05 for all three measurements).
The size and shape of the placental surface predict childhood blood pressure.
Blood pressure may be programmed by variation in the normal processes of placentation: these include implantation, expansion of the chorionic surface in mid-gestation and compensatory expansion of the chorionic surface in late gestation.
People whose birthweights were towards the lower end of the normal range have higher blood pressures and are at increased risk of developing hypertension in later life (1)(2)(3). This is thought to reflect foetal programming, the process by which malnutrition and other adverse influences during development alter gene expression and programme the body's structures and function for life (4,5). These adverse influences may also slow foetal growth, leading to low birthweight (6). A baby's birthweight reflects its success in obtaining nutrients from its mother. The source of these nutrients is not only the mother's current diet but her metabolism, which is a product of her lifetime's nutrition (7). Height is one indicator of a woman's nutrition in early life (8), and women who are short have lower rates of protein synthesis during pregnancy than do women who are tall (9).
A baby's birthweight also depends on the placenta's ability to transport nutrients to it from its mother (6). Small babies generally have small placentas (10), which suggests that placental size is an indicator of placental function. In some circumstances, however, an undernourished baby can expand its placental surface to extract more nutrients from the mother (11). This leads to high placental weight in relation to birthweight. Both low placental weight and high placental weight in relation to birthweight have been shown to predict later hypertension (12)(13)(14). Placental weight has inconsistent associations with blood pressure levels in children. There are reported associations between increased levels of blood pressure and low placental weight and a high ratio of placental weight to birthweight (15,16). Other studies have found no associations (17). While the size of the placenta is linked to foetal nutrition, which nutrients are delivered to the foetus is conditional on their availability in the maternal circulation. This will be related to the mother's body size. The effects of placental size on later hypertension are conditioned by the mother's body size (11) because her body is the source of nutrients.
The weight of the placenta does not distinguish its thickness from its surface area. To increase the surface for nutrient and oxygen exchange, the placenta can expand its invasion across the surface of the uterine lining or invade the maternal spiral arteries more deeply. The long-term consequences of these may be different.
The surface of the placenta is generally described as oval or round. To measure the ovality, two so-called diameters of the surface were routinely recorded in some hospitals (11). The maximal diameter described the length of the surface, while the lesser one bisecting it at right angles described the breadth. Studies in Finland and Holland have shown that hypertension in the offspring in later life is related to the size of the placental surface (11). This relation depends more on the breadth of the surface than on its length.
We measured the length and breadth of the placental surface in a study of newborn babies in Mysore, South India (18). The children's blood pressures were measured at the age of 9 years. We hypothesized that blood pressure levels would be related to the size of the placental surface and would relate more to the breadth than the length. We also hypothesized that the associations would depend on the mother's height. The relation between placental size and later hypertension differs in the two sexes (19). We therefore examined boys and girls separately.
In 1997, the Mysore Parthenon study recruited pregnant women attending the antenatal clinic at the Holdsworth Memorial Hospital (HMH) in Mysore, South India (18,20). The hospital ethical committee approved the study, and informed written consent was obtained from the parents and children. All women who had a singleton pregnancy and who were not diabetic before pregnancy were eligible. Out of 1233 eligible women, 830 (67%) took part in the study. At 30 ± 2 weeks of gestation, their body size was measured, including their height, using standardized methods. One hundred and fifty-six women delivered elsewhere and were not followed up any further: the remaining 674 women delivered live-born babies at HMH. Seven babies were stillborn and four had major congenital abnormalities. Neonatal anthropometric measurements were made on the remaining 663 within 72 h of birth by one of the four trained measurers, again using standardized methods. Weight was measured using digital weighing scales (Seca, Germany); crown-heel length was measured using a Harpenden neonatal stadiometer (CMS instruments, London, UK); head circumference was measured with blank anthropometric tape, marked and measured against a fixed ruler.
Placental dimensions were measured in 653 of the babies. The placentae were initially checked for completeness, and the cord clamp was released to allow the blood to drain. The amnion was stripped off, and the chorion was trimmed close to the placental edge. The cord was cut flush with the placenta, and any obvious clots were removed. The placenta was weighed using an electronic weighing machine. It was then placed on a flat surface, with the cotyledons facing upwards. The longest diameter (length) was identified by eye and measured using a graduated transparent plastic ruler placed on the surface. The diameter perpendicular to the length was defined as the breadth and was measured in the same way.
The children were followed up annually from birth until 5 years of age, and every six months thereafter. Eight children were excluded because of major medical conditions, while a further 25 children died. Ninety-one children did not attend the nine and a half year follow-up (nine untraceable, 26 moved away and 56 declined to participate) so that 539 children were examined (Figure 1). Further anthropometry was carried out including weight (Salter, Tonbridge, Kent, UK), height (Microtoise; CMS instruments) and mid-upper arm circumference. Systolic and diastolic blood pressures were measured in the left arm using an automated blood pressure monitor (Dinamap8100, Criticon, FL, USA). Appropriate-sized cuffs were selected for use based on the mid-upper arm circumference. Two recordings were made after five minutes seated at rest, and the average taken. The observers were unaware of the neonatal measurements of the children. The techniques used by different observers were standardized by regular intra-and inter-observer variation studies.
Blood pressure was recorded for all 539 children who attended the 9.5-year follow-up. Because placental size is increased in gestational diabetes, we excluded from the analysis 35 children whose mothers developed diabetes during gestation and 25 children for whom maternal glucose tolerance test data were missing. We included seven children born to mothers who developed hypertension and 29 babies born preterm. Placental measurements were missing for eight children, and hence our study sample comprises the remaining 471 children with complete measurements of the placenta.
The data included the weight of the placenta and the length and breadth of its surface. Placental area was calculated assuming an elliptical surface, using the maximal diameter (length) • lesser diameter (breadth) • p ⁄ 4 (11). We calculated the difference between the length and breadth of the surface and the ratio of placental area to birthweight. We examined associations between placental variables and blood pressure using linear regression, unadjusted and then adjusted for the child's body mass index (BMI) and height at 9.5 years; where p values from adjusted models are used, this is stated. We used linear regression to examine interactions between the effects on blood pressure of placental size and maternal height. Among girls, there were interactions between the effects of placental size and maternal height (see Results); and, as in the analyses of the Helsinki Birth Cohort (11,19), we examined separate models for girls born to mothers above and below the median height (154.5 cm). There were no similar interactions between placental size and maternal height in boys.
Table 1 shows the mean measurements of the 471 mothers, placentas, babies and children. None of the mothers smoked tobacco. At birth, boys were larger than girls, but their placental size was similar. Twelve boys and eight girls had systolic hypertension, using standard criteria (21).
There were no significant differences in birthweight or any of the placental measurements between the 471 children included in the study sample and the 118 children who were eligible but who were not included, either because of incomplete placental measurements or because of loss to follow-up (birthweight: 2850 g vs. 2781 g, p = 0.2; placental length: 19.5 cm vs. 19.3 cm p = 0.4; placental breadth: 17.0 cm vs. 16.7 cm p = 0.1; placental area: 261.6 cm vs. 255.1 cm p = 0.2; placental weight: 0.408 kg vs. 0.411 kg p = 0.8). However, there was a significant difference in maternal height between these two groups: 154.3 cm vs. 156.4 cm, p = 0.0002). Among boys, neither systolic nor diastolic pressure was related to birthweight, head circumference, birth length or the length of gestation. Table 2 shows the trends in blood pressure with placental size. Systolic pressure tended to rise as placental breadth, weight and area decreased. These trends became statistically significant after adjusting for the child's current BMI and height (for breadth b = )0.68 mmHg ⁄ cm). The trends with breadth and weight also became statistically significant after adjusting for birth weight (p = 0.04 for breadth and p = 0.03 for weight). Systolic pressure tended to rise as the ratio of placental area to birth weight decreased; however, this trend was not statistically significant. Diastolic pressure was unrelated to placental size.
Among girls, systolic and diastolic pressure fell as birthweight increased (systolic pressure fell by 4.2 mmHg per kg increase in birthweight, 95% CI 1.4-6.9, p = 0.003; diastolic pressure fell by 2.6 mmHg per kg increase, 95% CI 0.3-4.9, p = 0.03, after adjusting for current BMI). Systolic pressure in girls also fell as birth length increased (p = 0.04, adjusted for current BMI). It was not related to head circumference or to the length of gestation. Table 2 shows the trends in blood pressure with placental size. Systolic blood pressure rose as the ratio of area to birthweight increased. Diastolic blood pressure fell as placental weight increased. In Table 3, the trends in blood pressure with placental size among girls are divided according to the mother's median height. Among girls whose mothers' height was below the median, systolic pressure rose as placental breadth increased (b = 0.69 mmHg ⁄ cm) and as the ratio of placental area to birthweight increased. In a simultaneous regression, both larger placental area and lower birthweight were associated with higher systolic pressure (p = 0.001 and 0.002, respectively). Diastolic pressure also rose as the ratio of placental area to birthweight increased; however, this was a weaker trend than that for systolic pressure.
Among girls whose mothers' height was above the median, there were no trends in systolic blood pressure with either placental length, breadth or area (Table 3), but systolic pressure rose as the difference between the diameters increased (b = 1.40 mmHg ⁄ cm, p = 0.007), that is as the placental surface became more oval. In contrast to systolic pressure, diastolic pressure rose as placental breadth, length and area decreased, but was unrelated to the difference between diameters. The differing trends in systolic pressure with placental size in the two maternal height groups were statistically significant (p for interaction = 0.03 for breadth, 0.02 for area ⁄ birthweight, and 0.03 for the difference between length and breadth).
We found that the blood pressures of 9-year-old children in south India were related to the size and shape of the placental surface at birth. We suggest that these associations reflect the role of placental function in programming blood pressure (22). Consistent with findings in the Helsinki birth Cohort, we found different associations in boys and girls (19). Boys grow faster than girls from an early stage of gestation, even from before implantation, and this makes them more vulnerable if their nutrition is compromised (8,23).
More newborn boys than girls have retarded growth and placental abnormalities, and most of them die during the perinatal period (24). The associations between blood pressure and placental size were modified by adjustment for the child's current body size. This reflects the known association between raised blood pressure and rapid postnatal growth (3). We found that three different placental phenotypes predicted raised blood pressure depending on the child's sex and the mother's height. Common to each phenotype was that blood pressure was related to the shape or size of the placental surface and to the breadth rather than to the length. Many studies have shown that raised blood pressure in children and adults is related to lower birthweight within the normal range (1-3,15-17). This suggests that blood pressure levels are linked to foetal malnutrition (6). We now examine each of the three placental phenotypes and discuss their possible relation to foetal malnutrition.
In boys, higher systolic blood pressure was associated with smaller placental area. This association depended on reduction in placental breadth rather than length. This is consistent with findings among men in Holland, in whom hypertension was related to a short placental breadth but not to length (25). One suggestion is that tissue along the breadth is more closely related to nutrient delivery to the foetus than tissue along the length of the surface, and shorter breadth results in lesser delivery (26). The length and breadth of the placental surface are established by growth of the chorionic surface in mid-gestation. We suggest that reduced growth in placental breadth is associated with foetal malnutrition and raised blood pressure in boys. Among boys, the effects of placental size on blood pressure were not conditioned by the mothers' height. In the Helsinki Birth Cohort (19), the effects of placental surface area on hypertension were not related to the mother's height in men, but were stronger among women with short mothers. This led to the suggestion that, compared to boys, the nutrition of girls in utero depends more on the mothers' lifetime nutrition and metabolism, reflected in her height than on the mothers' diet in pregnancy (19).
In girls whose mother's height was below the median, raised systolic pressure was associated with a large placental area in relation to birthweight. This association is therefore opposite to that seen in the Helsinki cohort, in which hypertension was associated with smaller placental area, especially in women with shorter mothers (19). Again, the association in our study depended more on the breadth of the placental surface than on its length. A study of men and women born in a maternity hospital in Preston, UK, showed that high placental weight in relation to birthweight was associated with later hypertension in the offspring (13). Measurements of the placental surface were not available. This observation has been replicated, and high placental weight in relation to birthweight has also been shown to predict coronary heart disease (14,27). In response to maternal undernutrition in mid-gestation, foetal lambs are able to extend the area of the placenta by expanding the individual cotyledons (28). This increases the area available for nutrient and oxygen exchange and, if normal nutrition is restored in late gestation, there is a larger lamb than there would otherwise have been. This is profitable for the farmer, and manipulation of placental size by changing the pasture of pregnant ewes is standard practice in sheep farming.
There is evidence that a similar process occurs in humans and involves broadening of the placental surface rather than elongation (11,19). Compensatory placental expansion may be beneficial in some circumstances, but if the compensation is inadequate, and the foetus continues to be undernourished, the need to share its nutrients with an enlarged placenta may become an added metabolic burden that has long-term costs, which includes raised blood pressure.
In girls whose mothers' height was above the median higher, systolic pressure was predicted by a larger difference between the breadth and length of the placental surface that is by a more oval-shaped surface. In contrast, diastolic pressure was not predicted by the shape of the surface but by its size, pressures rising with decreasing placental area. In pregnancies complicated by pre-eclampsia, a disorder that is initiated by impaired implantation, the placental surface is oval (26). We therefore suggest that the association between systolic pressure and an oval placenta reflects disruption of the processes of implantation, which includes spiral artery invasion and recruitment, with consequent foetal malnutrition. The association between small placental area and diastolic pressure may reflect reduced expansion of the chorionic surface in mid-gestation. We can offer no explanation as to why impaired implantation would affect systolic pressure while reduced expansion of the chorionic surface would affect diastolic pressure.
In the Parthenon study, placental and newborn size and childhood blood pressure were measured by trained research staff according to standard protocols. None of the mothers smoked, and only seven had pregnancy-induced hypertension. The study has achieved high follow-up rates in the children; 80% of the original live-born babies of nondiabetic mothers and 84% of surviving children were studied at 9 years. The study is based on births in one hospital in Mysore, and the participants may therefore be unrepresentative of the whole Mysore population. At the time of the study, the Holdsworth Memorial Hospital was one of three large maternity units in Mysore. It is situated in a relatively poor area of the city, and most of the patients come from 'lower middle-class and middle-class' families. However, it is not a specialist referral hospital, and most women choose to deliver there because of its proximity to home. We do not think loss to follow-up would have introduced significant bias.
Although the mothers of the children included in our study sample were significantly shorter than those lost to follow-up, differences in birthweight and placental measurements between these groups were small.
We suggest that variations in three normal processes of placentation lead to foetal undernutrition and consequent raised blood pressure. The processes are those that accompany implantation and expansion of the chorionic surface in mid-gestation and compensatory expansion of the chorionic surface in late gestation. In girls, the effects of these disruptions on blood pressure are conditioned by the mother's early nutrition as indicated by her height.
p values for the differences between boys and girls using unpaired t-tests.
*p values adjusted for the child's current body mass index and height.
ª2011 The Author(s)/Acta Paediatrica ª2011 Foundation Acta Paediatrica 2011 100, pp. 653-660
This study was funded by the
This is an informal personal review of the development over time of my ideas about the concentrating mechanism of the mammalian renal papilla. It had been observed that animals with a need to produce a concentrated urine have a long renal papilla. I saw the function of the long papilla in desert rodents as an elongation of the counter-current concentrating mechanism of the inner medulla. This model led me to overlook contrary evidence. For example, in many experiments, the final urine has a higher osmolality than that of the tissue at the tip of the papilla. In addition, we had observations of the peristalsis of the renal pelvis surrounding the papilla. The urine concentration falls if the peristalsis is stopped. I was wrong; together, these lines of evidence show that the renal papilla is not just an elongation of the inner medulla. We are left without a full explanation of the concentrating mechanism of the mammalian renal papilla. It is hoped that other researchers will tackle this interesting problem.
My interest in how some mammalian kidneys can produce highly concentrated urine goes back to 1947-1948 when I first worked with desert rodents in Arizona. As young research associates in Dr Laurence Irving's laboratory, my husband Knut and I were given the exciting problem of finding out how desert rodents manage to exist entirely without drinking water.
Working in a small desert laboratory in the Santa Rita Mountains in Arizona, we soon found that Kangaroo rats and several other desert rodents live and thrive on a diet of dry seeds only without access to drinking water. As I collected and analysed urine from the animals, I marvelled at these high concentrations of urea and salt in the samples. At that time I made up my mind to learn kidney physiology. Little was known about kidney function; in fact, so little was known that a specialist was quoted as saying: 'All we really know about the kidney is that it makes urine.' It struck me as very interesting that the kangaroo rats and several other desert rodents have very long renal papillae. At that time, my husband, Knut, got hold of a dissertation by Ivar Sperber in Sweden. Examining all of Sperber's drawings of various mammalian kidneys, I found that all animals with a long renal papilla were from arid environments (see Figure 1). Conversely, animals from moist habitats have a very short or no papilla at all. A few years later, the counter-current hypothesis was proposed and I assumed that the renal papilla could be looked at as an elongation of the inner medulla and therefore as a part of the counter-current system. According to the counter-current hypothesis, the urine concentration would follow the tissue concentration gradient. This is clearly not the case in the papilla.
In 1964, Bruno Truniger and I found that urea concentration in the tissue stopped increasing at the border between the renal papilla and the upper part of the inner medulla. In the papilla, the concentration remained lower than in the urine (see Figure 2). The upper points in the figure are from animals on a highprotein diet and are the ones of interest here. A graduate student of mine, Susan Zell, did careful studies on a number of different desert rodents. In all cases, the urine osmolality far exceeded the tissue osmolality in any part of the papilla. Figure 3 shows an example of her results for the desert rodent Acomys. The concentration in the urine was more than 1000 mOsm higher than that in the tissue.
My interest in studying the peristalsis of the renal papilla stems from a chance observation. I had invited Karl Ullrich from Berlin to the US to collaborate with me. As we were preparing an animal for micro-puncture after we had injected lissamine green to colour the urine, we noticed that the peristalsis moved the bluegreen urine in waves through the collecting ducts in the renal papilla. This intrigued me and I decided to study the peristalsis further. Another observation showed the importance of the peristalsis. In a collaborative study with Carl Gottschalk on the counter-current mechanism, his assistant would prepare antidiuretic animals for micro-puncture studies. For a better access, she would remove the pelvic wall surrounding the papilla. To our surprise, this procedure resulted in increased urine production with a lower osmolality in the Figure 1 Contrasting kidney anatomy from dry and wet environments. Adapted from Sperber (1944) p. 317. The kidney from the desert rodent Psammomys has a long papilla, while the water dwelling Hydromys has no apparent papilla (Sperber 1944).
Figure 2 Distribution of urea in urine and renal tissue in rats on high-protein (solid triangles) and low-protein, high salt diet (open circles). U/P, urine-to-plasma ratio; t/P, tissueto-plasma ratio (Truniger & Schmidt-Nielsen 1964).
Figure 3 Average osmolality of bladder urine compared with average osmolality of papillary tissue from antidiuretic Acomys cahirinus killed by decapitation. Bladder urine was collected 30 min after spontaneous micturation (Zell 1973).
prepared animals. At that point, we did not understand the meaning of the observation. Later, together with Bruce Graves, I started studies of the peristalsis.
There is a film clip of peristaltic waves in the supplemental materials.
To track the peristalsis, we trans-illuminated the papilla with a fibre-optic light and recorded the light coming through the papilla on a chart recorder (see Figure 4). You can see the peristaltic contraction and relaxation phases clearly on the figure. To study the effect of the peristalsis on the renal papilla, we wanted to stop the action. Using the snare shown in Figure 5, we stopped the peristalsis at relaxed and contracted phases, as indicated in Figure 6, and rapidly fixed the tissue by pouring fixative directly on the preparation. Figure 7 is a schematic diagram of the papilla. As you can see, the entire papilla including the papillary epithelium is surrounded by the contractile pelvic wall. On the left of Figure 8 are cross sections of a contracted papilla and on the right are sections of a relaxed papilla (cf. Fig. 6). You can clearly see the contracted pelvic wall as a dark outline on the left. Figure 9 shows higher magnification cross sections 300 lm from the tip. We counted the number of collecting ducts in each cross section to define the segments of the papilla. We measured cell volumes by cutting out the relevant portions of micrographs of serial sections and weighing them. Figure 10 shows our results for cell volumes along the length of relaxed and contracted papilla. As you can see, epithelial cell volume is actually greater in the contracted papilla. And for the collecting duct cells, the effect is stronger and increases along the length of the papilla.
To study the physiological effect of the peristalsis, Bruce Graves and I performed experiments on pelvic paralysis or removal of the pelvis. Similar experiments were also performed in Mark Knepper's laboratory (Pruitt et al. 2006). One kidney in the live rat was treated and the other was sham operated as a control. Either paralysis with Xylocaine or surgical removal of the pelvis significantly reduces the papillary osmolality.
In 2002, together with an artist in Gainesville, I created an animated cartoon of my proposed model of the effect of the peristalsis on the function of the renal papilla.
A film clip narrated by myself at that time is available in the supplemental materials.
Hypothesis on how the renal pelvic peristaltic pumping of the papilla might contribute to the concentrating mechanism [This section is a mildly edited transcript of the video. It should be noted that the hypothesis shown in the video and in the following transcript is incomplete. In particular, an accounting of the forces moving water is far from complete. I suggest that this is a fertile area for future research].
On the basis of experimental findings by a number of investigators including us, I shall now present a hypothesis and an animation on how the renal pelvic peristaltic pumping of the papilla might contribute to the concentrating mechanism. The model deals only with mammals with a relatively long papilla. The highest degree of urinary concentration is found in mammals with a long renal papilla. In these, the peristalsis has the strongest effect on the papilla. I suggest that the kidney papilla works as a pump through alternating positive and negative pressures generated by the peristaltic contractions of the pelvic wall.
(1) Water moves into the collecting duct cells as a result of the small positive hydrostatic pressure on the walls of the cells, generated by the peristaltic wave pushing the fluid through the collecting ducts. (2) Fluid moves out of the cells as a result of the negative pressure generated by the elastic forces, which expand the papilla during rebound. (3) Fluid is removed from the interstitium by the vasa recta, which contain no blood at the time the fluid enters.
(4) The animated model shows a hamster papilla in cross section and then in a longitudinal section. (5) To concentrate the urine in the collecting ducts, water must be removed from the collecting duct fluid. (6) Partly, this water removal is caused by the accumulation of solutes in the papillary interstitium.
The model presented here tries to explain how the hydrostatic pressure generated by the pelvic wall peristalsis could contribute to the removal of water from the collecting duct urine. It does not deal with the solute. Water movements through a membrane result from the difference in water potential in the two compartments separated by the membrane. Water potential is decreased by solutes in solution and increased by hydrostatic pressure. Water moves through water channels (aquaporins). As shown by Mark Knepper and his colleagues (Nielsen et al. 2002), water leaves the collecting ducts lumen through the aquaporin water channels in the plasma membranes of the collecting ducts cells. Aquaporin-2, the antidiuretic hormone-sensitive water channel, is present in the apical membrane of the collecting duct cells.
Water molecules move through the aquaporins by single-file diffusion without entrainment of solutes.
Urea moves by diffusion through urea transporters. Water can leave the collecting duct cells through the water channels, aquaporins three and four, which also permit solutes to pass through.
In the animated model, the sequence of the events occurring in the papilla is shown at low urine flow. velocity with which the fluid is formed, thus creating an increment in the pressure on the wall of the ducts. longer and narrower. The vasa recta close and blood flow stops. (5) Some blood moves retrograde, some down towards the tip of the papilla. The loops of Henle close as fluid is pushed both retrograde and towards the tip. (6) During rebound, the papilla becomes shorter and broader. Collecting ducts are still closed, but water moves out of the cells into the interstitium due to negative hydrostatic pressure generated by the elastic properties at the interstitial matrix. (7) Ascending vasa recta are tethered to other structures and will open as tissue expands during rebound permitting water to enter the vasa recta. At this point in time, there is no blood in either the ascending or descending vasa recta. Loops of Henle, descending as well as ascending, are still empty. (8) Early relaxation lasts about 1 s. The papilla resumes its original shape. Collecting ducts are still empty and closed because the urine has not reached them yet. Vasa recta, first the descending then the ascending, are filled with blood pushing the column of water that had entered the ascending vasa recta towards the cortex. Loops of Henle are also filled with fluid.
The peristaltic mechanism proposed above is currently incomplete. Our knowledge of the papillary concentrating mechanism is insufficient to explain its action. However, the data clearly show that the papilla is an important part of the overall concentrating mechanism.
The mammalian kidney has two concentrating mechanisms. One is the counter-current mechanism that establishes a concentration gradient in the renal tissue increasing from cortex to the inner medulla by the addition of solutes to the tissue. The other less well understood concentrating mechanism is based on removal of water from the collecting duct fluid to conserve water for the organism. The data have shown that the increase in urine osmolality in the renal papilla is due to water removal. Fluid of a much lower osmolality than that of the final urine is removed from the collecting ducts. The forces responsible for this water movement are not yet understood. For example, Zell's observations (Fig. 3) of urine osmolalities more than 1000 mOsm kg )1 greater than that of the tissue would suggest osmotic pressure differences greater than 25 bar. Such pressures are far greater than can reasonably be generated by the peristalsis. The peristalsis of the pelvic wall enclosing the papilla is nevertheless an essential component of the concentrating mechanism. If peristalsis is disabled by removal of the pelvic wall or by paralysis with Xylocaine, the concentrating ability is significantly reduced (Pruitt et al. 2006).
At the Arizona Symposium (Dantzler 2006), it was suggested anecdotally by Bill Dantzler that the renal papilla functions differently from the inner medulla. It is not just an elongation of the inner medulla. This comment triggered a rethinking of my ideas. Many of my previous findings strongly support such a proposal. Knepper et al. (2003) have written an excellent review of the concentrating mechanism of the renal inner medulla. They have also shown that aquaporins are present in the papilla. These presumably play an important role in the concentrating mechanism.
Some insects also have a remarkable concentrating ability in their excretory organs. Such insects might use a related mechanism. When I gave the August Krogh lecture at the American Physiological Society (Schmidt-Nielsen 1995), I compared the concentrating mechanism in insects with that of desert rodents. Insects have amazing powers of conserving water. There again, we do not fully understand the mechanism. However, the insects have a mechanically strong chitinous tissue that is muscular and packed with mitochondria.
We are left with many problems yet to be solved. My father would quote Piet Hein saying, 'Problems worthy of attack prove their worth by hitting back.'
Acta Physiol 2011, 202, 379-385 Ó 2011 The Authors Acta Physiologica Ó 2011 Scandinavian Physiological Society, doi: 10.1111/j.1748-1716.2011.02261.x
Acta Physiologica Ó 2011 Scandinavian Physiological Society, doi: 10.1111/j.1748-1716.2011.02261.xFunction of mammalian renal papilla AE B Schmidt-Nielsenand B Schmidt-Nielsen Acta Physiol 2011, 202, 379-385
Ó 2011 The Authors Acta Physiologica Ó 2011 Scandinavian Physiological Society, doi: 10.1111/j.1748-1716.2011.02261.x Function of mammalian renal papilla AE B Schmidt-Nielsen and B Schmidt-Nielsen Acta Physiol 2011, 202, 379-385
I thank
There are no conflicts of interest.
Additional Supporting Information may be found in the online version of this article.
Video Clip S1. Syrian hamster kidney papilla. Infused with lissamine green. Showing waves of green colored urine pushed by the peristalsis.
A cartoon animation showing a proposed model of the concentrating mechanism of the renal papilla under peristalsis. The animated model shows a hamster papilla in cross section and then in a longitudinal section.
Please note: Wiley-Blackwell are not responsible for the content or functionality of any supporting material supplied by the authors. Any Queries (other than missing material) should be directed to the corresponding author for the article.
Objective: We report a patient who experienced delusional symptoms during gradual discontinuation of low-dose venlafaxine and required antipsychotic treatment. Method: Case report. Results: A 31-year-old woman with major depression had been treated abroad with venlafaxine before returning to Japan. Since venlafaxine is unavailable here, we supplemented her regular venlafaxine dosage of 37.5 mg ⁄ day with clomipramine 20 mg ⁄ day. After 5 weeks we reduced venlafaxine to 18.75 mg ⁄ day and uptitrated clomipramine to 40 mg ⁄ day. Four days later she developed delusions of reference, palpitations and nausea. Clomipramine was increased to 60 mg ⁄ day, and her symptoms subsided. Eight weeks later her supply of venlafaxine ran out, and within 4 days her condition deteriorated into more severe symptoms that required 4 monthsÕ antipsychotic treatment. Conclusion: We speculate that her symptoms were discontinuation syndrome, including psychotic symptoms and physical symptoms, caused by (i) venlafaxine-clomipramine interaction and ⁄ or (ii) the serotonin reuptake inhibitor-like effects of low-dose venlafaxine.
Venlafaxine, a serotonin-norepinephrine reuptake inhibitor, is known to carry a risk of discontinuation syndrome due to its relatively short half-life. The frequency of discontinuation syndrome was reported to increase with higher doses and abrupt discontinuation. The common symptoms are physical, including dizziness, headaches and nausea, as well as psychiatric symptoms such as agitation and anxiety (1,2). Conversely, reports of psychotic symptoms have been few (2). We report a patient who experienced psychotic and physical symptoms during gradual discontinuation of low-dose venlafaxine while being administered clomipramine. That combination of symptoms fits the diagnostic criteria for discontinuation syndrome and required long-term antipsychotic treatment until resolution.
A 31-year-old woman with major depression came to our out-patient clinic after having lived abroad. She had no prior history of physical illness, delusional symptoms, manic episodes or substance abuse. Before moving abroad, she had suffered from mild depressive mood, slight anxiety and low self-confidence for the first time, and had been treated for about 1 year with clomipramine (20-40 mg ⁄ day).She fully recovered and safely stopped medication (i.e. she experienced no discontinuation syndrome) 2 months before leaving Japan. Within 6 months after going abroad, she had a recurrence, was treated with venlafaxine for 6 months and almost recovered on a dosage of 37.5 mg ⁄ day before returning to Japan. When we first saw her she exhibited a slight depressive mood. Since venlafaxine is not available in Japan, we decided to start her on another antidepressant and gradually discontinue venlafaxine. We began coadministration of 20 mg ⁄ day of clomipramine, which she had previously taken and had been safe and effective for her condition, and her regular venlafaxine dosage of 37.5 mg ⁄ day. Her symptoms subsided after 5 weeks, so we reduced venlafaxine from 37.5 to 18.75 mg ⁄ day and uptitrated clomipramine from 20 to 40 mg ⁄ day. Since the imipramine equivalent doses of venlafaxine and clomipramine are the same (3), the decrease in the venlafaxine dosage was almost exactly compensated by the increase in the clomipramine dosage. Unfortunately, 4 days later, she began to suffer from depressive mood and delusions of reference along with palpitations and nausea. Therefore, we increased the clomipramine to 60 mg ⁄ day. Two weeks later her depressive mood and physical symptoms had recovered, but she still exhibited some slight indications of delusional feelings.
The patientÕs condition remained stable until her supply of venlafaxine ran out 8 weeks later, and deterioration was seen within 4 days. She lapsed back into delusions of persecution and fear of death along with palpitations, dizziness, nausea and stomachache that rendered her unable to work. At this stage, because we were unable to resume venlafaxine, we decided to administer an antipsychotic agent, perospirone, which is a type of serotonin dopamine antagonist available only in Japan (4). The dosage was 8 mg ⁄ day at the beginning. After 5 weeks her symptoms began to show visible improvement. It took a further 3 months of this antipsychotic treatment for her to gradually recover from her symptoms and be able to work in the same capacity as before. Meanwhile, perospirone was tapered down to 4 mg ⁄ day and finally 2 mg ⁄ day in the last month and discontinued. Twelve months have now passed since then, and the patient has completely recovered and shows no signs of any depressive mood or delusional symptoms on the current clomipramine dosage of 60 mg ⁄ day.
Our case is characterized by several significant findings. First, the patient exhibited prolonged delusional symptoms and fear of death as well as palpitations and nausea after discontinuation of venlafaxine. Her condition was diagnosed as venlafaxine discontinuation syndrome because she experienced not only psychotic symptoms but also physical symptoms (5). Second, she experienced these symptoms during gradual discontinuation of an extremely low dosage of venlafaxine combined with clomipramine, and they finally required long-term antipsychotic treatment. We speculate that the following factors contributed to the manifestation of the symptoms in this patient.
First, the symptoms may have resulted from an interaction between venlafaxine remaining in her system and clomipramine. In vitro studies indicate that venlafaxine is a relatively weak inhibitor of cytochrome P450 (CYP) 2D6 (1), and clomipramine is metabolized by CYP 1A2, 2C, 2D6 and 3A4 (6). Therefore, some changes in clomipramine metabolism may have been caused by adding or stopping venlafaxine, leading to some adverse reactions. A study of eight patients by Gomez and Perramon showed that, when venlafaxine was added to clomipramine, none of the patients experienced any adverse reactions (one showed an increase in the serum clomipramine level, but the others did not) (7). Meanwhile, Benazzi reported a patient who developed severe anticholinergic adverse reactions and hand tremors when clomipramine was augmented with venlafaxine (data on the serum clomipramine level were not reported) (8). Another patient experienced no adverse reactions and no change in the serum clomipramine level upon stopping venlafaxine (8). Although the data in those case studies were somewhat limited and we were unable to determine our patientÕs serum clomipramine level, we speculate that there might have been an interaction between the two drugs in our patient. Furthermore, the sensitivity to adverse reactions to drugs differs among individuals.
Another possibility may be that our patient had intolerance to the discontinuation of low-dose venlafaxine. She experienced psychotic and physical symptoms not only when the dosage of venlafaxine was decreased from 37.5 mg ⁄ day to 18.75 mg ⁄ day, but also when 18.75 mg ⁄ day of venlafaxine was discontinued. Clomipramine was being coadministered with venlafaxine at both times, but the clomipramine dose remained unchanged in the second case. Given the stable dosage of clomipramine when venlafaxine was discontinued, it would seem that the symptoms that occurred were a result of the gradual discontinuation of the low-dose venlafaxine, although, as noted above, there may have been interaction between venlafaxine remaining in her system and clomipramine. It was reported that patients generally experience discontinuation syndrome upon stopping higher dosages of venlafaxine (2). However, there have been reports that in discontinuation syndrome associated with low-dose venlafaxine certain patients developed severe psychotic symptoms, as was the case with our patient. Louie et al. (9) reported a 46-year-old woman who experienced auditory hallucinations 3 days after reducing the venlafaxine dosage (from 37.5 mg b.i.d. for 10 days to 18.75 mg b.i.d.). The symptoms persisted until her venlafaxine dosage was increased to 56.25 mg ⁄ day. Regarding combination therapy, Parker and Blennerhassett (10) reported a 42-yearold man who had auditory hallucinations for a period of 5 days after discontinuing treatment with 75-mg ⁄ day venlafaxine combined with 1000 mg ⁄ day lithium. Fava (11) described a 59year-old man with bipolar disorder (type 1) who manifested manic symptoms after venlafaxine (37.5 mg ⁄ day) was discontinued and replaced with lithium. The symptoms resolved within a few days after venlafaxine was restarted at 37.5 mg ⁄ day. Venlafaxine was discontinued after 3 months, and the patient again experienced discontinuation symptoms, which subsided within 2 weeks. All of the above patients also experienced physical symptoms.
We believe that discontinuation of the low dosage of venlafaxine played a significant role in these symptoms. One possible explanation for these symptoms is the serotonin reuptake inhibitor (SRI)-like effects of venlafaxine at low dosages. It was reported that, like other SRIs, venlafaxine selectively inhibited 5-HT uptake at low dosages (12), while others reported that discontinuation of SRIs resulted in a rapid decrease in serotonin availability in the brain (13). This accounts for the manifestation not only of physical symptoms such as dizziness and intestinal symptoms, but also of psychiatric symptoms. It is likely that the same mechanism as in SRI discontinuation occurs in discontinuation of low-dose venlafaxine and induces these unusual psychotic and physical symptoms.
Drug manufacturers recommend gradual reduction of the dose of venlafaxine to prevent discontinuation syndrome. However, our patient shows that unusual psychotic and physical symptoms (i.e. discontinuation syndrome) can be associated even with discontinuation of a low dose of venlafaxine, even when the discontinuation is gradual, and even when combinedwithotherdrugs. It isimportant, therefore, to consider that some patients may show intolerance to even low-dose venlafaxine discontinuation.
None.
None.
Depression has been associated with impaired recollection of episodic details in tests of recognition memory that use verbal material. In two experiments, the remember/know procedure was employed to investigate the effects of dysphoric mood on recognition memory for pictorial materials that may not be subject to the same processing limitations found for verbal materials in depression. In Experiment 1, where the recognition test took place two weeks after encoding, subclinically depressed participants reported fewer know judgements which were likely to be at least partly due to a remember-to-know shift. Although pictures were accompanied by negative or neutral captions at encoding, no effect of captions on recognition memory was observed. In Experiment 2, where the recognition test occurred soon after viewing the pictures, subclinically depressed participants reported fewer remember judgements. All participants reported more remember judgements for pictures of emotionally negative content than pictures of neutral content. Together, these findings demonstrate that recognition memory for pictorial stimuli is compromised in dysphoric individuals in a way that is consistent with a recollection deficit for episodic detail and also reminiscent of that previously reported for verbal materials. These findings contribute to our developing understanding of how mood and memory interact.
Depression is associated with impaired memory function. Previous findings have indicated that dysphoric mood specifically impairs recognition memory accompanied by the recollection of the encoding context in which material is first seen, and not recognition memory that is based in familiarity (Drakeford et al., 2010;Hertel & Milan, 1994;Jermann, Van der Linden, Adam, Ceschi, & Perroud, 2005;MacQueen et al., 2003;MacQueen, Galway, Hay, Young, & Joffe, 2002;Ramponi, Barnard, & Nimmo-Smith, 2004; see also Lemogne et al., 2006;Raes et al., 2006). Studies such as these have narrowed down the memory impairment associated with depression to the recollection component of recognition memory. This finding is of great significance as therapeutic interventions could exploit the spare, more automatic, capacity of familiarity-based recognition. Nevertheless, the recollection deficit remains poorly understood and replication has not been universal (see, Jermann, Van Der Linden, Laurençon, & Schmitt, 2009). In the current research, we investigated the effects of subclinical depression on the familiarity and recollection components of recognition memory for pictorial material as this allowed us to explore the important role that visual cognition plays in depressive ideation.
Visual processing, compared to verbal processing, is differentially influenced by depression, thus memory for pictorial material may differ substantially from memory for verbal material. For example, Baker and Jessup (1980) argue that verbal processing is more characteristic of depressive ideation than visual processing; when dysphoric participants were directed to process information visually (with mental images), rather than verbally, visual processing was rated as more euphoric than verbal processing. Bywaters, Andrade, and Turpin (2004) reported a positive correlation between sad mood and vivid imagery when recalling pictures, suggesting that visual processing was less susceptible than verbal processing to disruption in depression. The processing of images is likely to make use of resources that do not completely overlap with those required for verbal ideation as proposed in different models of working memory (Baddeley, 2007;Barnard, 1999). Image processing may thus distract from, rather than prompt, classic abstract depressogenic themes such as "failure." On the basis of these kinds of evidence, it has been suggested (see, Holmes, Arntz, & Smucker, 2007) that mental imagery could be used in cognitive behaviour therapy in order to alleviate negative emotional symptoms. Yet, basic knowledge of the effects of sad mood on recognition memory for pictorial material remains sparse.
Pictures are informationally richer than words: whereas the word "telephone" covers a multitude of physical forms, a picture of a phone depicts a specific model of phone with some clearly delineated attributes and hence offers greater potential for recollection. In fact, memory for pictures is usually better than memory for words -the picture superiority effect. Dewhurst and Conway (1994) argued that pictures give rise more readily to an enriched automatic recollective experience because the encoding experience can be richer both perceptually and semantically. Dewhurst and Conway (1994) and Rajaram (1996) have demonstrated that in recognition memory, the picture superiority effect is confined to the recollection component by using the remember/know procedure (Gardiner, 1988;Tulving, 1985). After recognising a picture or a word as seen before, participants were asked to indicate whether they could recollect the encoding context, in which case they assigned a "remember" judgement, or whether they simply knew that they had come across the item before but could not recollect the context in which the item was first seen, (i.e. the item was familiar) in which case they assigned a "know" judgement. More remember judgements were reported for pictures than for words, whereas an equal or larger number of know judgements were reported for words relative to pictures. It has also been shown that different brain regions are involved in the recollection of pictures and words (Kensinger & Schacter, 2006;Woodruff, Johnson, Uncapher, & Rugg, 2005).
There is increasing evidence that the memory difficulties experienced in depression are better characterised as an impairment of recollection and not familiarity (Drakeford et al., 2010;Hertel & Milan, 1994;Jermann et al., 2005;MacQueen et al., 2002MacQueen et al., , 2003;;Ramponi et al., 2004). Depressed individuals when recognising items presented previously, have little difficulty with realising that the item was seen before, but they have considerable difficulty remembering the item spatial or temporal context. More specifically, with the remember/know procedure, it was found that recognition responses assigned with remember judgements were reliably fewer in depressed (Drakeford et al., 2010) and subclinically depressed participants (Ramponi et al., 2004) than in controls, whereas the number of recognition responses assigned with know judgements did not differ reliably. The first goal of the present investigation was to determine whether the same pattern of impaired recollection and intact familiarity would be observed for pictorial stimuli. Given the arguments just exposed that dysphoric mood affects verbal, but not visual, processing, it was possible that spared visual processing capacity could offset any recollection deficit.
The second goal was to study the modulatory effect of the emotional content of memoranda on the recollection deficit. Emotional salience is known to have a very strong effect on memory for events (for reviews, see Buchanan, 2007;Hamann, 2001;Kensinger, 2004;LaBar & Cabeza, 2006). As life experiences worth remembering must be discriminated from experiences that can be forgotten without cost, emotional salience is a clear candidate for marking out particular experiences for privileged mnemonic access. The memory enhancing effect of emotion has been observed across a number of paradigms and for different events or stimulus types. There is also substantial evidence that emotional stimuli are better recollected than neutral ones (Dewhurst & Parry, 2000;Kensinger & Corkin, 2003) and that this holds for pictorial stimuli as well (Comblain, D'Argembeau, Van der Linden, & Aldenhoff, 2004;Dahl, Johansson, & Allwood, 2006;Dolcos, LaBar, & Cabeza, 2005;Ochsner, 2000;Sharot, Delgado, & Phelps, 2004;Sharot & Yonelinas, 2008; but see Aupee, 2007) and also that the emotional context in which to-be remembered items are appraised influences memory. Pre-experimentally neutral objects or words that are embedded in emotional pictures are better recalled than those embedded in neutral pictures (Erk et al., 2003) and show different neural activation patterns at retrieval (Smith, Henson, Rugg, & Dolan, 2005). Similar effects have been reported for the retrieval of neutral words (Maratos, Dolan, Morris, Henson, & Rugg, 2001) and pictures (Cahill & McGaugh, 1995) when the emotional context is determined by accompanying prose.
In depression, the memory enhancement effect for negative emotional material is even more pronounced. Depression has been linked with an abnormal appraisal of emotional stimuli (for a review see Leppanen, 2006) that contributes to the onset and maintenance of depression. This distorted emotional information processing is reflected in enhanced memory for negative emotional material, i.e. material congruent with sad mood (Blaney, 1986;Elliott, Rubinsztein, Sahakian, & Dolan, 2002;Murphy et al., 1999). Cognitively, this effect is thought to be mediated by increased allocation of processing resources to negative material as more elaborate associations are generated to information consistent with an individual's current concerns (see Blaney, 1986;Williams, Watts, MacLeod, & Mathews, 1997). Neuroimaging investigations suggest that the abnormal amygdala activation typically associated with sad mood (for a review see Drevets, 2003) may mediate the memory enhancement of negative material (Moritz, Gläscher, & Brassen, 2005).
The memory enhancing effect for negative material in depressed individuals has been observed for faces that vary in emotional expression (Gilboa-Schechtman, Erhard-Weiss, & Jeczemien, 2002;Ridout, Astell, Reid, Glen, & O'Carroll, 2003;Ridout, Noreen, & Johal, 2009) and Jermann, Van der Linden, and D'Argembeau (2008) observed that this recognition memory enhancement was captured specifically by the recollection component, as is the case for words (Jermann et al., 2009;Lewis, Critchley, Smith, & Dolan, 2005). Whether pictures of scenes that have negative emotional connotations are also better retained by participants experiencing dysphoric mood, potentially offsetting the general recollection impairment remains unclear (for example see Jermann et al., 2009). This question was addressed in the current research by comparing dysphoric and control participants on recognition memory for pictures varying on an emotional dimension. In the present investigation, we tested, first, whether subclinically depressed participants demonstrated impaired recollection and intact familiarity for pictures, and second, whether emotion can modulate picture memory.
Experiment 1 evaluated the effect of mood on recollection for neutral pictures that did not convey inherent affective information. It was possible that a recollection deficit parallel to that found for verbal material (Drakeford et al., 2010;Ramponi et al., 2004) would be observed, although any deficit could be offset by a relative preserved visual processing capacity.
The remember/know procedure was employed to explore which component of picture recognition, if any, is affected by dysphoric mood. Contrary to previous protocol for verbal materials, we followed the precedent set by Ochsner (2000) for pictorial materials so that ceiling effect could be avoided and did not test memory immediately, but after a two-week delay. With this significant delay between study and test, a recollection deficit can also be expected to be evident as a reduction in the number of know judgements. Recognition responses that are initially assigned a remember judgement are later on experienced with a know-type awareness because of the loss of episodic detail, i.e. a "remember-to-know" shift occurs (Conway, Gardiner, Perfect, Anderson, & Cohen, 1997;Dudukovic & Knowlton, 2006;Herbert & Burt, 2003, 2004;Knowlton & Squire, 1995). If at shorter delays depressed participants initially report fewer 'remember' judgements, at longer delays when the episodic details are lost and remembering becomes knowing, they would be expected to report fewer 'know' judgements. Consequently, at long intervals the recollection deficit associated with dysphoric mood can be expected to be reflected in know judgements.
We also explored whether biasing the processing of pictures with an emotional or neutral interpretation affected memory. Picture processing was biased with the use of picture captions that participants were asked to read and evaluate when viewing the picture (Teasdale et al., 1999). For example, an image of a staircase could be paired with the caption 'where Johnny had a bad fall' (negative context), whereas the same image of a staircase could be paired with the caption 'where the telephone was usually kept' (neutral context). This procedure has the considerable advantage that basic picture processing is common to the neutral and negative conditions: the same image could be biased across participants with either a neutral or negative caption. An enhancement of picture memory embedded in negative captions was expected due to the memory enhancing effect of emotion.
Research Council Cognition and Brain Sciences Unit. Mood was assessed using the Beck Depression Inventory (BDI), a 21-item self-report measure of depression (Beck, Ward, Mendelson, Mock, & Erbaugh, 1961). The recommended cut-off point for mild depression in the BDI is 9/10 (Kendall, Hollon, Beck, Hammen, & Ingram, 1987). Fourteen subclinically depressed participants with a BDI score of 10 or above and 14 control participants with a BDI score of 9 or below participated; the two groups were matched for age and IQ (Scale 2, Form A of the Cattell Culture Fair Intelligence Test, Cattell & Cattell, 1973). Table 1 summarises the profiles of the 2 groups.
Affective Picture System (IAPS: Lang, Bradley, & Cuthberg, 1997) and other sources where necessary. All pictures were of neutral valence and medium levels of arousal. Using 9-point scales, pilot ratings resulted in an average valence of 5.19 (SD = 0.76) and average arousal of 4.58 (SD = 0.66). The pictures were divided into 4 lists of 20 pictures (A, B, C, and D) that were matched for image content, valence and arousal.
At encoding, pictures from two of the four matched lists were presented in a random order. Negative captions were paired with the pictures in one list, and neutral captions were paired with the pictures in the other. Using the example above, the image of a staircase in one list was paired with a negative caption, whereas a similar but not identical image of a staircase in second list was paired with a neutral caption. All captions were matched for numbers of words. Using 9-point scales, pilot ratings resulted in an average valence of 2.56 (SD = 0.63) and average arousal of 6.44 (SD = 0.84) for the negative captions and in an average valence of 4.97 (SD = 0.48) and average arousal of 3.16 (SD = 0.75) for the neutral captions. At test, the pictures from the remaining two unviewed lists were used as foils. Thus, eighty pictures from the four lists (two viewed lists and two unviewed lists) were presented in a random order. Encoding and test list assignment was randomly determined.
To ensure that attention was paid to the meaning given to the picture by the caption at encoding, participants were required to make an incompatibility judgment, indicating whether the caption was incompatible with the picture it was paired with. An additional 10 neutral filler pictures with incompatible captions (e.g., a picture of some cows in a field would be paired with the incompatible caption: "She washed up after the evening meal") were interspersed throughout the 40 encoding pictures. Ten filler pictures/caption pairs were also placed at both the beginning and end of the encoding list as buffers; for two of these, the captions were incompatible.
Procedure-Each participant was tested individually. In the encoding phase, neutral pictures were presented one at a time with either a negative or neutral caption presented below each image. Participants indicated when they considered the image-caption pair to be incompatible by pressing the space bar. They were not told that their memory for the pictures would be tested later. The image-caption pair appeared on the screen for 6000 ms with an interstimulus interval of 1500 ms. Participants were required to respond within this time frame.
The test phase took place approximately two weeks later (14.7 days +/-2.4 days). Before presentation of the picture sequence, participants were told that some of the pictures were from the image set presented during the encoding phase two weeks earlier, whereas others were new ones. For each image, participants were asked to indicate whether they recognised the image as one of those presented in the encoding phase. They were given extensive instructions (see Gardiner, Ramponi, & Richardson-Klavehn, 1998) on how to make remember and know judgments following a positive recognition. Participants were instructed to give a remember or know judgment only when they were certain that they recognised the image as one of those belonging to the image set presented at encoding. They were strongly discouraged from guessing: if they were uncertain that an image was an 'old' one, they were instructed to say that the image was new. Participants indicated their response by pressing labelled keys on the computer keyboard. Prior to presentation of the test pictures, participants' accuracy in assigning remember and know judgments was assessed with viewed and not-viewed filler pictures in a short practice session in which they explained their reason for a given recognition judgment and the experimenter verified that this response was appropriate.
Mean proportions of picture recognition and of pictures assigned remember and know judgments, with corresponding judgments to unviewed pictures (i.e. false alarms), are shown in Table 2. An alpha level of .05 was used for all statistical tests in this paper.
Recognition memory-In a 2 (group: subclinically depressed vs. controls) × 2 (image caption: negative vs. neutral) mixed ANOVA on recognition memory corrected for false alarms (i.e. overall recognition proportions minus overall false alarm proportions, for each participant), there was a significant main effect of group, F (1,26) = 4.87, MSE = .033, p < .05, η p 2 = .16, indicating that subclinically depressed participants recognised fewer pictures than controls; thus dysphoric mood appears to also affect memory for pictorial material. Neither Published as: Acta Psychol (Amst). 2010 November ; 135(3): 293-301.
Sponsored Document Sponsored Document the image caption effect, F (1,26) = .23, MSE = .009, η p 2 = .01, nor the interaction of group with caption, F (1,26) = .04, MSE = .009, η p 2 = .00 was significant. A measure of discriminability (d′; d′ = Z(hits) -Z(false alarm), see Table 3) between the old and new pictures showed exactly the same pattern: d′ measures for the subclinically depressed group were significantly lower than those for the control group, F (1,26) = 7.25, MSE = .356, p < .05, η p 2 = .22 and neither the image caption effect nor the group by caption interaction were significant (all Fs < 1).
A measure of criterion (c) or of the propensity of participants to produce a positive recognition response was also computed (c = (Z(hits) + Z(false alarms)) ⁎ -0.5, see Table 3). In the corresponding ANOVA a group effect was observed on this measure, F (1,26) = 18.07, MSE = .126, p < .05, η p 2 = .41. The image caption effect and the interaction were not significant (all Fs < 1). Thus, relative to controls subclinically depressed participants demonstrated a greater propensity, or bias, to report recognising a picture, as reflected in the larger number of false alarms (see Table 2) reported by the subclinically depressed participants t (26) = 4.55, SEM = .040, p < .05.
Recollection and familiarity-Remember and know judgments, corrected respectively for remember and know false alarms, were also analysed. In the corresponding 2 (group) × 2 (image caption) mixed ANOVA for the remember judgments, neither the main effects nor the interaction were significant (all Fs < 1). For the know judgments, in the corresponding ANOVA, the group effect approached significance, F (1,26) = 3.32, MSE = . 038, p = .08, η p 2 = .11, but neither the effect of image caption nor the interaction was significant (all Fs < 1).
The underlying assumption of the remember/know procedure is that the relationship of recollection and familiarity is a redundant one (Joordens & Merikle, 1993): all items are familiar and a subset of these is recollected. An alternative approach to the separation of recollection and familiarity is based on the assumption that the recollection and familiarity processes are independent, i.e. some items can only be recollected and some items can only be familiar but some items can be both recollected and familiar (Jacoby, Toth, & Yonelinas, 1993). We derived recollection and familiarity estimates (Table 3) from the remember and know judgements within a Dual-Process Signal-Detection Model (Yonelinas, Kroll, Dobbins, Lazzara, & Knight, 1998) based on the assumption of independence. This model specifically takes into account response bias, when there is a difference in the number of false alarms between two groups. The formula used to calculate recollection estimates from remember and know judgements is reported in Yonelinas et al. (1998). In this model, familiarity, unlike recollection, is assumed to reflect a signal-detection process, so familiarity is measured with d′.
For the recollection estimates, our results paralleled those reported for the remember judgments, with no effects of caption, group or their interaction (all Fs < 1). The same analysis on familiarity d′ showed a significant group effect, F (1,26) = 4.81, MSE = .44, p < .05, η p 2 = .16, indicating that relative to controls, subclinically depressed participants are impaired at discriminating old from new pictures on the basis of familiarity. With this measure no other effects were significant (all Fs < 1).
In Experiment 1 the subclinically depressed participants showed an impairment in recognition memory for pictorial material. They were less able to discriminate between older and newer pictures than control participants implying that dysphoric mood can have a considerable effect on the processing of pictures despite these conveying a richer set of information. Thus, there is no evidence that visual processing, in contrast to verbal processing, is differently affected in depression in the extent to which it can support mnemonic processes. The elevated number of false alarms in subclinically depressed participants relative to controls was also notable; subclinically depressed participants had an increased tendency to judge a picture as seen before at the expense of recognition accuracy.
The recognition deficit in this sample of subclinically depressed participants was captured by a reduction in the number of the know judgements. This result can be nevertheless consistent with the recollection deficit found for verbal material (Drakeford et al., 2010;Jermann et al., 2005;Hertel & Milan, 1994;MacQueen et al., 2002MacQueen et al., , 2003;;Ramponi et al., 2004) when considered in the context of the remember-to-know shift (Conway et al., 1997;Dudukovic & Knowlton, 2006;Herbert & Burt, 2003, 2004;Knowlton & Squire, 1995). In the current study recognition memory was not tested immediately as in the previous studies, but two weeks later to avoid ceiling effects, and what would have been fewer remember judgements at short delays become fewer know judgements at longer delays due to episodic details having been lost.
Finally, the processing of neutral pictures appeared unaffected by negative captions designed to bias picture elaboration. A memory enhancement for the pictures associated with negative captions was not observed. In Experiment 2, negative and neutral emotionality was manipulated by varying inherent picture content.
Emotional experience is often described as having two orthogonal dimensions: a valence dimension that characterises the pleasantness or unpleasantness of the experience, and an arousal dimension that characterises how calming or exciting an experience is (Russell, 1980). There is extensive evidence (e.g. Bradley, Greenwald, Petry, & Lang, 1992) that arousal plays a key role in enhancing memory for emotional material. The general finding has been that arousal levels associated with experimental stimuli modulate amygdala activation, which in turn has a role in memory consolidation (see Kensinger, 2004;Kensinger & Corkin, 2004;LaBar & Cabeza, 2006). Recent research in our laboratory (Croucher, 2006;Ewbank, Barnard, Croucher, Ramponi & Calder, 2009;Murphy, Hill, Ramponi, Calder, & Barnard, in press) has studied the "impact" of pictures on an observer; this factor is believed to have a critical role in the enhancement of memory. The term "impact" is used in visual media to describe particularly eye-catching pictures. The degree of affective impact reflects the extent to which a picture personally affects the viewer and how quickly the viewer grasps a picture's content. Distinct stimuli can be appraised impersonally and "propositionally" as having comparable negative emotional content even though they may not give rise to comparable levels of genuinely felt affect (Teasdale & Barnard, 1993). Croucher (2006) found that the degree of affective impact modulated recollection of pictures in a recognition task: negative pictures of high impact were better remembered than negative pictures of low impact, even though these image sets were matched for negative valence and arousal. Murphy et al. (in press) found that high impact negative pictures compared to matched low impact negative and neutral pictures receive priority processing in visual attention, and Ewbank et al. (2009) found increases in amygdala activation when viewing high impact relative to low impact negative pictures and yet no difference when viewing low impact negative pictures relative to neutral ones despite a difference in valence.
The second experiment examined the effect of dysphoric mood on recollection and familiarity for pictures of inherent negative emotional content. Subclinically depressed and control participants reported remember and know judgements after viewing neutral pictures and two sets of pictures that varied in affective impact but were matched on other key dimensions that could influence retention (arousal, valence, distinctiveness, approach/avoidance and visual complexity). In order to achieve a retention interval more comparable to that employed in studies of the influence of mood on recollection and familiarity for words whilst avoiding ceiling effects, we reduced exposure time and made use of an incidental monitoring task that did not require aspects of picture content to be actively processed.
A replication of the memory deficit observed in Experiment 1 in subclinically depressed participants would confirm that dysphoric mood influences memory for pictorial material and verbal material similarly. We predicted that the recollection deficit would be captured by the remember judgements as previously found for verbal material due to the similarly short gap between encoding and test. We expected that the negative pictures of high affective impact would be better remembered than both the low affective impact and neutral pictures.
3.1.1 Participants-Participants were other participants recruited from the volunteer panel of the Cognition and Brain Sciences Unit. Twenty subclinically depressed participants with BDI scores of 10 or above were matched for age and IQ with twenty control participants with BDI scores of 9 or below. Table 4 summarises the profiles of the 2 groups.
The negatively-valenced pictures were selected from a set of IAPS pictures rated for valence, arousal, distinctiveness, approach/avoidance and visual complexity by Croucher (2006). The impact scale varied from 9 (indicating maximum affective impact) to 0 (indicating minimum affective impact). Sixty negatively-valenced pictures were used, with half of the pictures of high impact (M = 7.13, SD = 0.93) and the other half of low impact (M = 3.67, SD = 0.53). The two sets were matched for valence (M = 2.01, SD = 0.73; t (58) = 1.16, SEM = .187, d = .30), arousal (M = 4.76, SD = 1.17; t (58) = .38, SEM = .304, d = .10), distinctiveness (M = 5.80, SD = 1.19; t (58) = 1.09, SEM = .306, d = .29), approach/ avoidance (M = 7.30, SD = 1.03; t(58) = .12 , SEM = .268, d = .03) and visual complexity (M = 4.12, SD = 1.09; t(58) = 1.05, SEM = .282, d = .29). There were also 30 pictures of neutral valence (M = 5.23, SD = 0.64) and arousal (M = 4.79, SD = 0.73): half of the neutral pictures were from the IAPS and the other half were drawn from other sources to allow them to be matched on arousal with the negatively-valenced pictures. The 30 pictures in each of the high impact, low impact and neutral sets were divided into two lists matched for each of the criteria listed above. At encoding, participants were shown one of the lists for each set such that 45 pictures were presented overall. At test, pictures from the second (unviewed) list from each set acted as foils. Encoding list assignment was randomly determined.
In order to ensure attention was directed at the pictures at encoding without requiring detailed attention to picture content, participants were asked to press the spacebar to pictures that appeared upside down. To this end, 15 filler pictures were presented in an inverted orientation. Fifteen filler pictures were also placed at the beginning and end of the testing pictures as buffers and two of these were also presented upside down. Fifteen additional fillers were presented interspersed with the 45 critical pictures to increase the numbers of pictures presented. At study and test pictures were presented in a pseudorandom order so that no more than 2 pictures of the same type were presented in sequence.
Procedure-Each participant was tested individually. Participants were not told that their memory for the pictures would be tested later and were asked to judge whether an image was presented upside down. A fixation cross appeared for 500 ms and after an additional 500 ms, the image appeared for 1000 ms. Forty-five minutes later participants' recognition for the pictures was tested. The same procedure as Experiment 1 was followed for the recognition decision and remember/know judgments task.
Mean picture recognition scores and the proportions of pictures assigned remember and know judgments, with respective false alarms, are shown in Table 5. Signal detection measures, (d′ and c), and recollection and familiarity estimates are reported in Table 6.
In a 2 (group: subclinically depressed and controls) × 3 (picture-type: high impact, low impact and neutral) repeated-measures ANOVA of recognition scores corrected for false alarms, the effect of picture-type was significant, F(2,76) = 24.215, MSE = .017, p < .001, η p 2 = .39; neither the effect of group nor its interaction with picturetype was significant (all Fs < 1). Planned comparisons showed that participants recognised more high impact than low impact pictures, t(39) = 3.77, SEM = .021, p < .001, d = .60, and more low impact than neutral pictures, t(39) = 3.56, SEM = .035, p < .001, d = .56.
The corresponding ANOVA on d′ (see Table 6) mirrors these results with a main effect of picture-type, F(2,76) = 12.306, MSE = .540, p < .001, η p 2 = .24 , but no main effect of group, F(1,38) = 2.579, MSE = 1.454, η p 2 = .06, or interaction, F(2,76) = .824, MSE = .540, η p 2 = . 02. In the corresponding ANOVA on the criterion measures there was a significant effect of picture-type, F(2,76) = 8.358, MSE = .089, p = .001, η p 2 = .18, and a picture-type by group interaction, F(2,76) = 3.356, MSE = .089, p < .05, η p 2 = .08. The interaction indicated that the subclinically depressed participants adopted a more liberal criterion as in Experiment 1, than the controls when recognising the neutral pictures (t(38) = 2.17, SEM = .135, p < .05, d = .80) but not when recognising the high impact (t(38) = .34, SEM = .154, d = .11) or low impact pictures (t(38) = .89, SEM = .150, d = .18).
Recollection and familiarity-For the remember judgments corrected for false alarms, the same 2 (group) × 3 (picture-type) ANOVA showed that the group effect was significant, F(1,38) = 4.874, MSE = .080, p < .05, η p 2 = .1, with the subclinically depressed participants remembering fewer pictures than controls. The picture-type effect was also significant, F(2,76) = 32.450, MSE = .017, p < .001, η p 2 = .46, whereas the group by picturetype interaction was not, F(2,76) = .895, MSE = .018, η p 2 = .02. The same results were observed in the corresponding ANOVA for the recollection estimates (reported in Table 6) where a significant group effect, F(1,38) = 4.132, MSE = .087, p = .05, η p 2 = .10, and picture-type effect, F(2,76) = 35.950, MSE = .017, p < .001, η p 2 = .49, were observed, but the interaction was not significant, F(2,76) = .947, MSE = .017, η p 2 = .02. In both groups, the high impact pictures were associated with more remember judgements than the low impact pictures, control: t(19) = 3.61, SEM = .041, p < .05, d = .81; subclinically depressed: t(19) = 2.11, SEM = .041, p < .05, d = .47, and were associated with higher recollection estimates, control: t(19) = 3.79, SEM = .040, p < .05, d = .85; subclinically depressed: t(19) = 2.42, SEM = .039, p < .05, d = .53; in planned comparisons the low impact pictures were associated with more remember judgements than the neutral pictures (Control: t(19) = 6.02, SEM = .044, p < .05, d = .68; subclinically depressed: t(19) = 2.41, SEM = .048, p < .05, d = .54) and with higher recollection estimates (Control: t(19) = 3.39, SEM = .040, p < .05, d = .76; subclinically depressed: t(19) = 2.35, SEM = .049, p < .05, d = .53). The recollection deficit observed in subclinically depressed participants was not attenuated for the negative pictures, as confirmed by the absence of a group by picture-type interaction For the corrected know judgments, in the corresponding ANOVA we did not find a significant effect of group, F(1,38) = 4.113, MSE = .062, η p 2 = .07, picture type, F(2,76) = 2.119, MSE = .009, η p 2 = .05, or an interaction F(2,76) = .683, MSE = .009, η p 2 = .01). With the familiarity d′ estimates we observed no effect of group, (1,38) = .336, MSE = 1.116, η p 2 = . 009, and no significant interaction, F(2,76) = .129, MSE = .300, η p 2 = .003), but in this case the effect of picture type, F(2,76) = 6.157, MSE = .300, p < .05, η p 2 = .14, was significant.
Published as: Acta Psychol (Amst). 2010 November ; 135(3): 293-301.
Sponsored Document
The neutral pictures were less familiar than the low impact, t(39) = 3.05, SEM = .124, p < . 05, d = .43, and high impact pictures, t(39) = 2.78, SEM = .131, p < .05, d = .48, but the high and low impact pictures did not differ in how familiar they were, t(39) = .134, SEM = .106, d = .02.
While a global difference in recognition memory between the subclinically depressed and control groups was not observed, a deficit was detected specifically in the recollection component. Subclinically depressed participants reported fewer recognition responses accompanied by remember judgements than control participants. These results replicate previous findings for recollection with verbal material (Drakeford et al., 2010;Hertel & Milan, 1994;Jermann et al., 2005;MacQueen et al., 2002;2003;Ramponi et al., 2004) and extend those findings to pictures. Dysphoric mood particularly affects memory for the episodic details accompanying events, rather than affecting memory for the event per se (Ramponi, Nayagam, & Barnard, 2009;Raes et al., 2006) and we can now conclude, on the basis of this current evidence, that this holds for pictures. Again, as in Experiment 1, the visual processes engaged to encode and retain (particularly context) information presented in pictorial form are clearly affected even in subclinical forms of depression.
Negatively-valenced pictures were better recognised and recollected than neutral pictures with negative pictures of high affective impact being better recollected, replicating Croucher (2006). With the familiarity estimates, but not with the know judgements, we found that for all participants negative pictures (whether of high or low impact) were judged to be more familiar than the neutral ones.
In this study we investigated how dysphoric mood affects recollection and familiarity for pictorial material that varies in emotionality as determined by verbal description or inherent content. We conclude that recognition memory for pictorial material is adversely affected by mood. Visual processing to the extent that it supports the encoding, elaboration and retention of information in pictorial form is also impaired by dysphoric mood.
In relation to recollection and familiarity, our results are largely consistent with the view that recollection of the episodic context is impaired by dysphoric mood, (Drakeford et al., 2010;Hertel & Milan, 1994;Jermann et al., 2005;MacQueen et al., 2002;2003;Ramponi et al., 2004, but see Jermann et al., 2009). In Experiment 1, recognition memory for the neutral pictures was significantly reduced in subclinically depressed participants compared to controls. This reduction was captured by the familiarity component of recognition memory, when memory is consolidated and a remember-to-know shift has occurred due to contextual details having been lost (Conway et al., 1997;Dudukovic & Knowlton, 2006;Knowlton & Squire, 1995), the original deficit in recollection can subsequently be reflected in recognition memory experienced as knowing. In Experiment 2, where recognition memory was tested soon after pictures were viewed, a marked deficit in recollective experience, but not in familiarity, was observed in the subclinically depressed participants.
A notable finding from both experiments involves criterion differences between the subclinically depressed and control groups. In Experiment 1, subclinically depressed participants had twice as many false alarms as control participants, showing that subclinically depressed participants adopted a very liberal criterion when deciding whether or not they had seen a neutral picture irrespective of the encoding context. As indicated by the discrimination measure (d′), subclinically depressed participants were less able to discriminate between old and new pictures and, as indicated by the criterion measure (c), they opted for a more inclusive strategy, thus risking more false positive judgements. A similar pattern was obtained in Experiment 2, but only with the pictures conveying neutral content. No reliable differences on criterion measures were obtained with the negative pictures.
It is possible that under considerable uncertainty, created by long retention intervals and/or indistinctive pictures of neutral content, subclinically depressed participants' negative appraisal of their own ability to remember events may contribute to a choice of an over-inclusive strategy. Previous findings reported in the recognition memory literature concerning criteria differences for dysphoric mood have shown very little consistency. Some studies have reported a more liberal criterion for negative words related to dysphoric mood (Deijen, Orlebeke, & Rijsdijk, 1993;Zuroff, Colussy, & Wieglus, 1983), while other studies have reported a more conservative criterion (Corwin, Peselow, Feenan, Rotrosen, & Fieve, 1990;Dunbar & Lishman, 1984) and yet others do not report any criterion shift (Brebion, Smith, & Widlocher, 1997;Channon, Baker, & Robertson, 1993;Watts, Morris, & Macleod, 1987). A number of methodological factors may be responsible for these differences, including differences in task demands, materials and severity of mood. The general conclusion is that, at a minimum, accurate interpretation of results must consider these biases (e.g., Zuroff et al., 1983).
The presence of a marked liberal criterion with neutral pictures at two different delays and for both low (Experiment 1) and high levels of recollection (Experiment 2) appears consistent with differences linked to dysphoric mood in the problem-solving literature (Slife & Weaver, 1992). One difference can be characterised as a pure cognitive deficit (represented by the discrimination measure d′) and the other difference can be characterised as a difference in metacognitive processes involved in strategy selection (as represented by the criterion measure c). This analogy could help to resolve earlier empirical discrepancies. For example, in discussing the theoretical basis for impaired initiative in depressed states, Hertel and Hardin (1990) identified motivational and metacognitive factors as plausible explanatory candidates. This hypothesis of differences in metacognitive processes related to mood when recognising neutral pictures could be further explored by manipulating the "expectancy" of the likelihood that old pictures would appear in the recognition test (see Gardiner, Richardson-Klavehn, & Ramponi, 1997;McCabe & Balota, 2007) and by measuring participants' self-perception of their mnemonic competence. Criterion differences, and by implication their metacognitive origins, have very notable effects on remember and know judgements (Donaldson, 1996;Gardiner, Ramponi, & Richardson-Klavehn, 2002;Parks & Yonelinas, 2007;Postma, 1999;Wixted, 2007) that could substantially alter the view presented thus far on the effect of mood on memory; hence, these warrant systematic future investigation.
While the neutral and negative captions of Experiment 1 did not have an observable effect on the recognition memory for either control or subclinically depressed groups, a mnemonic advantage for negatively valenced pictures was observed in Experiment 2. As anticipated, this effect was particularly marked for the recollection of pictures of high relative to low affective impact that were equal in rated arousal and valence. This result replicates the finding that the emotion effect is captured by the recollection component (Comblain et al., 2004;Croucher, 2006;Dahl et al., 2006;Dolcos et al., 2005;Ochsner, 2000;Ritchey, Dolcos, & Cabeza, 2008;Sharot et al., 2004;Sharot & Yonelinas, 2008).
The recollection deficit observed in subclinically depressed participants was not reduced for the negative pictures, and thus, a mood-congruent-memory effect was not observed. Moodcongruent-memory effects have been observed primarily with self-referent encoding or autobiographical elaboration (Bradley & Mathews, 1983;Lewis et al., 2005;Matt, Vazquez, & Campbell, 1992). Impact ratings reflected the extent to which an image had particular emotional impact and salience to the self, but not to which it was inherently autobiographical. It may be far easier for subclinically depressed individuals to relate words of negative valence, like "regret" or "sorrow", to their personal experience (Kensinger, 2004) than to relate pictures that depict specific events that are happening to someone else. Furthermore, in studies of recognition memory for emotional faces, mood-congruence effects have sometimes been observed (Gilboa-Schechtman et al., 2002;Jermann et al., 2008;Ridout et al., 2003;Ridout, Noreen, et al., 2009) and sometimes not (e.g. Ridout, Dritschel, et al., 2009). Ridout, Dritschel, et al. (2009) found that mood-congruent-memory was absent in participants that were not oriented towards the affective element of the faces at encoding. The task that participants were asked to carry out at encoding in the current experiments simply ensured attention to the pictures. In one case, they reported whether the caption was coherent with the picture and in the other case, whether the picture was inverted. At encoding, participants were not explicitly oriented to appraise the picture's affective component, a factor that may explain why a moodcongruent effect was not observed. Nevertheless, the intrinsically negative pictures were still better recollected, probably indicating that more semantic elaboration had occurred at encoding, but the semantic elaboration for negative pictures was not more extensive in the subclinically depressed group.
Finally, in relation to mood-congruent memory, one caveat of this study is that our results rely on normative valence and arousal ratings of the negative and neutral pictures. Subjective ratings were not collected, so it is possible that mood-congruent effects had occurred but these would only have been apparent if subjective ratings were considered. A second caveat of this study is that the effect of subclinical depression on recognition memory for the negative and neutral pictures was not measured at longer time intervals to test the patterns of the remember-to-know shift. This is also important in view of evidence that the effect of emotion on memory can differ with time (e.g. Sharot & Yonelinas, 2008). Ritchey et al. (2008) provide some evidence that, over a time interval of a week, recollection for emotionally negative pictures increased even though overall recognition memory remained similar. If this is the case, it is possible that remember-to-know shifts may have different time courses for pictures of negative content relative to neutral pictures.
In summary, the current experiments complement and extend prior work with verbal material by showing that dysphoric mood has adverse effects on the recognition of pictures. In Experiment 2 the deficit is confined to the recollection component of recognition memory and the effects observed after a two-week interval in Experiment 2 are consistent with controls remembering more at the outset leading to a greater opportunity than the subclinically depressed participants to report knowing pictures at the longer interval. Along with a core deficit in memory comparable to that observed with verbal material, the presence in subclinical depression of a more liberal criterion with neutral material is consistent with a second, qualitatively distinct difference in metacognition in the subclinically depressed group that was confined to pictures that were unremarkable in their content. Control and subclinically depressed participants showed comparable memory enhancement for negatively valenced pictures.
Published as: Acta Psychol (Amst). 2010 November ; 135(3): 293-301.Sponsored DocumentSponsored Document Sponsored Document
This work is funded by the
Table 1
Mean scores and SD of individual differences indices and of demographic characteristics for subclinically depressed and control participant.
Control participants (N = 14) Subclinically depressed participants (N = 14) M SD M SD BDI 2.79 2.32 14.43 3.84 IQ 118.93 9.69 118.36 24.41 AGE 33.71 11.81 33.57 13.69 Male:Female (N) 7 7 6 8 Published as: Acta Psychol (Amst). 2010 November ; 135(3): 293-301. Sponsored Document Sponsored Document Sponsored Document Ramponi et al. Page 18
Published as: Acta Psychol (Amst). 2010 November ; 135(3): 293-301.
Published as: Acta Psychol (Amst). 2010 November ; 135(3): 293-301.
Published as: Acta Psychol (Amst). 2010 November ; 135(3): 293-301.
Published as: Acta Psychol (Amst). 2010 November ; 135(3): 293-301.
Background: Effects of high iodine-concentration contrast material on the image quality of coronary CT angiography (CCTA) have not been well evaluated.
Purpose: To compare the image quality and attenuation values of CCTA between patients administered iopromide 370 and iomeprol 400 with the use of 64-slice multidetector CT. Material and Methods: Patients were prospectively enrolled and were randomized into two groups (group A, 151 patients received iopromide 370, iodine fl ux ϭ 1.48 g I/s; group B, 146 patients received iomeprol 400, iodine fl ux ϭ 1.60 g I/s). CT attenuation was measured in the coronary arteries and great arteries and measurements were standardized based on an iodine fl ux of 1.50 g I/s. The image quality of 15 coronary artery segments was graded by two radiologists in consensus with the use of a four-point scale (1 ϭ excellent to 4 ϭ poor enhancement). Non-parametric statistical approaches were used to compare the two groups.
Results: The median attenuation values in the coronary arteries were 454 HU and 464 HU for iopromide 370 and iomeprol 400, respectively, and they did not differ ( P ϭ 0.26). When standardizing for an iodine fl ux, signifi cantly higher attenuation values were found for iopromide 370 (median ϭ 460 HU, range ϭ 216 -791 HU) compared with iomeprol 400 (median ϭ 435 HU, range ϭ 195 -758 HU) ( P ϭ 0.006). The median image quality score of coronary arterial segments was 1 (range 1 -2) for both groups ( P ϭ 0.84).
The attenuation values in the coronary arteries after injection of the same amount of two high iodine-concentration contrast materials at the same fl ow rate with different iodine fl uxes are similar with no difference in image quality. With standardization for an iodine fl ux, the attenuation is signifi cantly higher when using iopromide 370.
The recent development of multidetector CT (MDCT) has enabled noninvasive imaging of the coronary arteries. Because of rapid cardiac motion, high temporal resolution is essential for cardiac CT techniques and spatial resolution should be suffi cient for depiction of the branches of coronary arteries. In addition, optimal enhancement is essential for the entire diagnostic process and all diagnostic images including curved multiplanar reformation and volume-rendering images for coronary arteries to ensure reliable results during image post-processing. Recent studies have described the use of contrast material with high iodine content (370 mg I/ml or 400 mg I/ml) for coronary CT angiography (CCTA) (1 -4). However, the effects of high iodineconcentration contrast material on the vessel visibility of CCTA using a 64-slice MDCT have not been well evaluated.
The purpose of this prospective study was to compare the image quality of CCTA using a 64-slice MDCT with two contrast agents with high iodine concentrations (iopromide 370 mg I/ml and iomeprol 400 mg I/ml). The attenuation obtained in the coronary arteries and the great arteries as well as the subjective degree of enhancement in the coronary arterial segments were evaluated.
From August 2007 to December 2008, 337 consecutive patients (206 men and 131 women; mean age, 54 Ϯ 10 years; age range, 22 -75 years) referred for CCTA for suspected coronary artery diseases were enrolled prospectively in the study. Patients who had arrhythmia, renal insuffi ciency (a serum creatinine level more than 1.5 mg/dl), a history of allergic reaction to contrast material, previous history of surgery or stenting for coronary artery diseases, heart failure, and women who were potentially pregnant or nursing were not eligible for study participation. Patients who were unable to cooperate with breath-holding for at least 10 s or had a body weight above 90 kg or below 40 kg were not enrolled, to limit the heterogeneity within the patient population. The institutional review board approved this study and all patients provided written informed consent to participate in this study. Images of the study were excluded from the analysis in the presence of poor image quality caused by a severe motion artifact, extensive calcifi cation or inappropriate scan coverage.
Patients were randomly assigned into two groups by the use of permuted block randomization (5), which differed with regard to the iodine concentration of the contrast agent that was administered. Group A patients received iopromide 370 (370 mg I/ml, Ultravist 370; Bayer Schering Pharma, Berlin, Germany) and group B patients received iomeprol 400 (400 mg I/ml, Iomeron 400; Bracco Imaging, Milan, Italy). For each patient, age, sex, height, and body weight were recorded.
Coronary CT angiography was performed on a 64-row detector system (Aquilion 64, Toshiba Medical Systems, Otawara, Japan) in the craniocaudal direction to cover from the aortic root to the caudal end of the heart. Before an examination, all patients were instructed to take a deep breath and to hold their breath. Patients with a prescanning heart rate of 65 beats per minute or higher were given 100 mg of metoprolol (Seloken; AstraZeneca, Zoetermeer, The Netherlands) orally 1 hour before CT scanning. Just before the injection of contrast material, 0.6 mg nitroglycerin was administered sublingually for vessel dilation.
The contrast agents were prepared at 37 ° C and were injected with an 18-guage needle through the right antecubital veins by the use of a dual-syringe power injector (Stellant-Dual Flow; Medrad, Pittsburgh, Pa., USA). Contrast agents were administered at a rate of 4 ml/s and were followed by a 40 ml saline fl ush at the same rate (6). The resulting contrast material volume and injection rate, respectively, were 70 ml and 4 ml/s (total injection time, 17.5 s) for all patients. This procedure refl ected a clinical routine that resulted in an iodine fl ux (iodine delivery rate, IDR) of 1.48 g I/s for iopromide 370 and 1.60 g I/s for iomeprol 400. The IDR was calculated as follows: IDR (g I/s) ϭ [iodine concentration (mg I/ml) ϫ contrast fl ow (ml/s)]/1000 mg/g (7). In the analysis, attenuation values in the vessels of the two groups were additionally standardized on a fl ux of 1.50 g I/s for the comparison of the attenuation values as described in the statistics section.
Synchronization between the passage of contrast material and data acquisition was achieved with the use of a real-time bolus tracking technique (SureStart; Toshiba Medical Systems, Tokyo, Japan) using a region of interest (ROI) positioned in the ascending aorta. The trigger threshold inside the ROI was set at 200 HU. The main scanning parameters were as follows: number of detectors, 64; individual detector width, 0.5 mm; gantry rotation time, 400 ms; tube voltage, 120 kVp; tube current, 400 mA; feed/rotation, 3.2 mm; feed/second, 8.0 mm. A phantom of the American Association of Physicists in Medicine (AAPM) was used for the calibration of the CT scanner to ensure reproducible measurement of CT ROI at a 6-month interval. An acceptable limit of the attenuation of water was 0 Ϯ 4 HU. For daily quality control of CT, air calibration was used.
Data collection and analysis were performed using previously reported methodology (2, 8) with some modifications. The data set was reconstructed with retrospective electrocardiography gating with time windows of 70%, 75%, and 80% of the R-R interval for one of the mid-diastolic phases of cardiac cycles to represent motion-free images. For patients with a heart rate of more than 70 beats per minute during scanning, additional reconstruction of images for systolic phases was performed. For patients who showed suboptimal image quality on predetermined phases due to variable or high heart rates, the best phase for image reconstruction was selected after a review of multiphase images of one slice at the mid-heart level by 10 ms or 1% interval throughout the cardiac cycle.
With the use of axial data, an experienced radiologist reconstructed the three-dimensional volume-rendered images and curved multiplanar reformation images of the coronary arteries using commercial software (Aquaris ver. 3.5.2.1; TeraRecon, San Mateo, Calif., USA).
Coronary CT angiography was analyzed by consensus of two experienced cardiac radiologists who were blinded to contrast material used. Oblique coronal or sagittal sections in the dataset were selected to measure the attenuation value objectively using an ROI placed at the proximal part of the four main coronary arteries (right coronary artery, RCA; left main artery, LM; left anterior descending artery, LAD; and left circumfl ex artery, LCX) (6). The ROIs were drawn as large as possible within the vessels with care taken to avoid motion artifacts as well as calcifi cation or soft plaques (Fig. 1). The overall visualized length without motion artifact of each coronary artery depicted on a curved planar reformation image was extracted by the use of the standard software and compared for both groups.
The readers subjectively evaluated the image quality based on a 15-segment American Heart Association (AHA) model ( 9). The proximal, middle, and distal segments of RCA (segments 1, 2, 3), posterior descending branch of RCA (segment 4), LM (segment 5), and the proximal, middle, and distal segments of the LAD (segments 6, 7, 8), the fi rst and second diagonal vessels (segments 9, 10), proximal and distal LCX (segments 11, 13), obtuse marginal branch (segment 12), and posterolateral and posterior descending branch of LCX (segments 14, 15). For each segment, image quality was graded with the use of a four-point scale. Scores were defi ned as grade 1 for excellent (strong homogeneous enhancement with sharply defi ned vessel edges), grade 2 for good (homogeneous enhancement with mildly blurred vessel edges), grade 3 for fair (inhomogeneous enhancement with moderately blurred vessel edges), and grade 4 for poor (inhomogeneous enhancement with markedly blurred vessel edges) image quality. Grades 1, 2, and 3 were assumed as scores of diagnostic image quality. If the segment was not delineated by the fi eld of view for its hypoplasia or aplasia, it was counted as " not applicable " .
The CT attenuation of both groups was measured objectively using the ROI technique at the proximal ascending aorta, main pulmonary artery, and distal descending thoracic aorta at the level of the inferior margin of the heart on axial images (Fig. 2). Consistency of contrast enhancement was also assessed by calculation of ROI differences as follows: consistency of contrast enhancement ϭ (attenuation of proximal ascending aorta -attenuation of distal thoracic aorta).
Artifacts due to insuffi cient mixing of contrast material in the right atrium were assessed using a four-point scale. Scores were defi ned as grade 1 for no streak artifact, grade 2 for a mild streak artifact without an obscured vessel segment, grade 3 for a moderate streak artifact with a mild but acceptable degree of vessel segment obscuration, and grade 4 for a severe streak artifact with markedly obscured vessel segments.
Vital signs including heart rate were monitored before the injection of contrast material and during CCTA. All types of adverse effects of contrast agents were also recorded and vital signs were closely monitored during CT examinations.
Sample size was calculated by assuming that the grades of image quality in each coronary segment in two groups would not be different. If one supposes that the proportion of grade 1 would be 90% in each group and that difference in proportions of grade 1 in the two groups less than 10% would be regarded as no difference in them, then one needs 220 patients with 80% power and 5% type I error.
Demographic data such as sex, age, body mass index (BMI), height, baseline heart rate, and patients ' characteristics were compared using the chi-squared test and Wilcoxon two-sample test as appropriate. The attenuation values of the coronary arteries and the great arteries and the visualized length of each coronary artery without motion artifact were averaged for all patients in each group and the overall average was used to compare the two groups using Wilcoxon rank sum tests for single vessel and t test on ranks based on a mixed linear model to account for the correlation of vessels within the same patient (coronary arteries only). All tests were based on ranks as deviances from normality were expected for the attenuation values. In addition, the attenuation values were standardized for an iodine fl ux of 1.50 g I/s by dividing the attenuation values by the iodine fl ux used and by multiplying by 1.50 (10). These were analyzed in the same way as for the raw attenuation values. The results of subjective analysis for the image quality of 15 coronary arterial segments graded with the use of the four-point scale and streak artifacts of the right atrium were assessed by the use of a multinomial regression based on generalized estimating equations (GEEs) taking into account multiple vessels within the same patient. Independence was used as working correlation matrix. The number of patients who showed adverse reactions to contrast agents was assessed using the chi-square test. All tests were performed two-sided; P values Ͻ 0.05 were regarded as statistically signifi cant. Data processing and analysis were performed with SPSS (version 10.0; SPSS, Chicago, IL, USA) and SAS (version 9.2, SAS Institute, Cary, NC, USA).
Among 337 patients initially recruited for this study, 297 patients (151 patients in group A and 146 patients in group B) were ultimately included in this investigation. Forty patients were excluded from the analysis because of poor image quality caused by a severe motion artifact ( n ϭ 11), extensive calcifi cation ( n ϭ 27) or inappropriate scan coverage ( n ϭ 2). There were 175 men and 122 women (age range, 22 -75 years; median age, 54 Ϯ 10 years). Patients ' demographics and characteristics were not signifi cantly different between the two groups in terms of sex, age, BMI, height, baseline heart rate, use of premedication, and symptoms except atypical chest pain (Table 1).
No differences were found between the median attenuation values (454 HU for iopromide 370 and 464 HU for iomeprol 400, P ϭ 0.26) in the coronary arteries in the two groups without standardization for an iodine refl ux (Table 2). After standardization for an iodine fl ux of 1.5 g I/s, the median attenuation value for iopromide 370 was higher than that for iomeprol 400 (460 HU and 435 HU, respectively, P ϭ 0.006) (Table 2, Fig. 3a). No signifi cant effect of the location of vessel was found on the attenuation values ( P ϭ 0.99 and P ϭ 0.99 for raw and standardized values, respectively). The median visualized length of the RCA, LAD, and LCX depicted on curved planar reformation images was not signifi cantly different between the two groups ( P Ͼ 0.05, each) (Table 3, Fig. 3b). A signifi cant effect of the location of vessel was found on the visualized lengths of the three coronary arteries in each group ( P Ͻ 0.0001).
A total of 3602 segments of 297 patients were evaluable for subjective analysis of contrast enhancement, excluding 853 segments for their hypoplasia or aplasia. In a subjective assessment with the use of a four-point scale, each segment of the coronary arteries showed excellent or good contrast enhancement. For iopromide 370, 91.8% ( n ϭ 1699) of 1851 coronary segments had a score of grade 1 and the remaining 8.2% ( n ϭ 152) had a score of grade 2. For iomeprol 400, 91.6% ( n ϭ 1604) of 1751 coronary segments had a score of grade 1 and the remaining 8.4% ( n ϭ 147) had a score of grade 2. There was no signifi cant difference for the image quality score of each coronary segment between the two groups ( P ϭ 0.84).
The median attenuation values of the proximal ascending aorta, main pulmonary artery, and descending thoracic aorta were not different between the two groups (Table 4). After standardization for an iodine fl ux of 1.5 g I/s, the median attenuation of proximal ascending aorta and descending thoracic aorta were signifi cantly higher for iopromide 370 than for iomeprol 400.
The consistency of contrast enhancement calculated by ROI differences between the proximal ascending aorta and distal thoracic aorta was not signifi cantly different between the two groups. The average consistency of contrast enhancement for iopromide 370 and iomeprol 400 was 49 Ϯ 47 HU and 55 Ϯ 56 HU, respectively ( P ϭ 0.12).
Nine of 151 patients (6%) in group A and 7 of 146 patients in group B (5%) showed a mild streak artifact.
No case showed moderate or severe streak artifacts that obscured the right coronary artery.
There was no moderate or severe adverse reaction to the intravenous contrast agent, but 13 patients had a mild adverse reaction such as nausea ( n ϭ 6), dizziness ( n ϭ 3), and urticaria ( n ϭ 4). Group A had eight cases (5.3%) with mild adverse reactions (three cases of nausea, one case of dizziness, four cases of urticaria) and group B had fi ve cases (3.4%) with mild adverse reactions (three cases of nausea, two cases of dizziness). There was no signifi cant difference in the frequencies of adverse reactions to the contrast agents between the two groups ( P ϭ 0.43). All of the patients ' symptoms were resolved soon after conservative treatment. The average change in the heart rate after injection of contrast material was 2.9 beats per minute (median, 2; range, 0 -20) for group A and 3.2 (median, 3; range, 0 -25) for group B, respectively. For group A, 105 patients showed a decrease or no change in heart rate after injection of contrast material, and 41 patients showed increased heart rates (mean, 3.1; median, 3; range, 1 -20 beats per minute). For group B, 111 patients showed a decrease or no change in heart rates after injection of contrast material and 35 patients showed an increased heart rate (mean, 3.1; median, 2; range, 1 -10 beats per minute).
An optimal contrast agent application protocol for CCTA is critical as the ability to diagnose coronary The raw attenuation values in the coronary arteries showed no significant difference between two groups with the median attenuation value of 454 HU (range 213 -780 HU) for iopromide 370 (group A, Gr.A) and that of 464 HU (range, 208 -809 HU) for iomeprol 400 (group B, Gr.B), respectively ( P ϭ 0.26). After standardization with an iodine fl ux of 1.5 g I/s, the attenuation using iopromide 370 was signifi cantly higher in the coronary arteries (except LAD, P ϭ 0.0539). (b) Boxplot of visualized length of coronary vessels along with P values for the comparison of groups. The measurements were similar in both groups. artery disease mainly depends on adequate visualization of the coronary arteries. Arterial enhancement is generally determined by the number of iodine molecules administered. The rate of iodine administration can be increased either by increasing the injection fl ow rate or by increasing the iodine concentration of the contrast agent (11). However, the degree of arterial enhancement following the intravenous injection of the same amount and type of contrast material is highly variable among individuals for physiological parameters such as cardiac output and central blood volume (11). Moreover,according to CADEMARTIRI et al. (12), contrast bolus geometry may not only depend on contrast density and fl ow but also on contrast volume, bolus chaser, and heart diseases. In our study, we strictly controlled the factors that could infl uence the contrast geometry, such as the total volume and injection rate of the contrast material as well as body weight.
Several investigators have reported a major impact of different iodine fl uxes on arterial enhancement (7,10,13,14). The arterial enhancement is proportional to the iodine fl ux; the higher iodine fl ux, the higher the arterial enhancement. The iodine fl ux can be increased by increasing the iodine concentration or injection rate (10). Slightly higher attenuation values found for iomeprol 400 can be attributed to the higher iodine delivery rate (1.60 g I/s) as compared with that of iopromide 370 (1.48 g I/s). However, the standardized attenuation values for an iodine fl ux of 1.5 g I/s were shown to be signifi cantly higher for the use of iopromide 370 as compared with iomeprol 400 in the proximal ascending aorta, thoracic descending aorta, and the average value across the great arteries. This advantage was also found in the RCA, LM, and LCX and the average across the coronary arteries. The higher values for the use of iopromide 370 after standardization of the iodine fl ux may refl ect the effects of the different viscosities of the two contrast agents (12.6 mPas for iomeprol 400 and 9.5 mPas for iopromide 370 at 37 ° C). The higher viscosity of iomeprol 400 might result in inhomogeneous mixing and therefore a non-superior attenuation value, which seemed to be compensated by the higher iodine fl ux. When the viscosity of the contrast agent approaches that of blood, mixing of the two fl uids takes place more easily. This is of importance in cardiovascular examinations (15).
As a limitation of this study, we did not evaluate vessels with atherosclerotic disease with extensive calcifi cation, because these factors may affect the visualization and attenuation of vessels. We administered a total amount of 70 ml of contrast material at a fl ow rate of 4 ml/s regardless of the BMI and height of patients. However, the BMI did not differ signifi cantly in the two groups; therefore no bias in the comparison between the groups was expected.
In conclusion , the image quality of CCTA using the same amount of iopromide 370 or iomeprol 400 at the same injection rate is similarly excellent with different iodine fl uxes. With standardization for an iodine fl ux, the attenuation is signifi cantly higher when using iopromide 370.
* Data are mean values with standard deviations, and numbers in parentheses are ranges. † Two-group chi-squared test (two-sided). ‡ Two-group Wilcoxon rank sum test (two-sided).
Acta Radiol 2010(9)
This study was supported by a grant from
The authors report no confl icts of interest. The authors alone are responsible for the content and writing of the paper.
Samsung Biomedical Research Institute, Samsung Medical Center, Seoul, Korea.
Although drive counts are frequently used to estimate the size of deer populations in forests, little is known about how counting methods or the density and social organization of the deer species concerned influence the accuracy of the estimates obtained, and hence their suitability for informing management decisions. As these issues cannot readily be examined for real populations, we conducted a series of 'virtual experiments' in a computer simulation model to evaluate the effects of block size, proportion of forest counted, deer density, social aggregation and spatial auto-correlation on the accuracy of drive counts. Simulated populations of red and roe deer were generated on the basis of drive count data obtained from Polish commercial forests. For both deer species, count accuracy increased with increasing density, and decreased as the degree of aggregation, either demographic or spatial, within the population increased. However, the effect of density on accuracy was substantially greater than the effect of aggregation. Although improvements in accuracy could be made by reducing the size of counting blocks for lowdensity, aggregated populations, these were limited. Increasing the proportion of the forest counted led to greater improvements in accuracy, but the gains were limited compared with the increase in effort required. If it is necessary to estimate the deer population with a high degree of accuracy (e.g. within 10% of the true value), drive counts are likely to be inadequate whatever the deer density. However, if a lower level of accuracy (within 20% or more) is acceptable, our study suggests that at higher deer densities (more than ca. five to seven deer/100 ha) drive counts can provide reliable information on population size.
Population size and status assessment are important for game and wildlife management. In the case of rare species, wildlife managers often try to increase population size; in medium-sized, harvested populations, their densities determine hunting plans, while in populations considered overabundant, reduction of density may be judged necessary. According to Leopold et al. (1938) 'any wildlife management worthy of the name will be difficult or impossible until we develop satisfactory methods of inventory'.
Although ungulate census may be relatively easy in open areas (Lowe 1969), it is much harder in forest habitats. A classic example of this difficulty was demonstrated by the study of Andersen (1953), in which roe deer (Capreolus capreolus L.) population size was estimated at 70 individuals, but shooting aimed at eliminating all individuals revealed that there were at least 213 roe deer (a few animals remained). Similar results were obtained by Ueckermann (1964) and Pielowski and Bresiński (1982). Among existing methods, those considered reliable often require much effort (e.g. capture-mark-resighting -Strandgaard 1967) or expensive equipment (e.g. thermal imaging- Gill et al. 1997;Smart et al. 2004).
One quite commonly used method is that of drive counts (Hosely 1956;Overton 1969;Pucek et al. 1975;McCullough 1979;Koster and Hart 1988;Short and Hone 1988;Jędrzejewska et al. 1994;Dzięciołowski et al. 1995;Lancia et al. 1996;Noss et al. 2006). Usually, an area having welldefined boundaries (e.g. forest roads) is driven by a line of beaters who start from one side of the area, and drive deer towards stationary observers placed along the remaining sides. The count within a given block is the number of animals leaving the block through the line of drivers plus those passing through the observers' lines. If population size rather than density is of interest (which is usually the case in wildlife or game management), the number of animals counted in all blocks is extrapolated to the total forest area.
In spite of the popularity of drive counts, so far there have been few attempts to evaluate their applicability for estimating animal density/population size. McCullough (1979) used drive counts in his study of a white-tailed deer (Odocoileus virginianus Zimmermann) population. He compared results of drive counts with population size estimated from the age or death of individuals, and concluded that at low population density drive counts underestimated the population relative to age reconstruction, while at high densities it tended to overestimate. However, McCullough (1979) tested drive counts on an enclosed population (driven individuals remained within the area), which probably limits his conclusions for freeranging populations. Pucek et al. (1975) compared drive counts with snow tracking, and concluded that the latter provides lower density estimates than drive counts, although they did not test the efficiency of drive counts as such. Cederlund et al. (1998) in general found that drive counts and other methods derived from hunting practices were unreliable, and pointed out that double counting, especially at high densities, is hard to avoid. Staines and Ratcliffe (1987) found that deer could be hard to flush from cover, and suggested that drive counting be limited to small areas owing to difficulties in co-ordinating large numbers of beaters and counters. On the basis of existing knowledge, it is therefore hard to draw any clear conclusions regarding the effectiveness of drive counts. In Poland, the method is recommended for use by game managers (Nasiadka 1994), and is probably the most commonly used method for estimating deer populations and trends.
The total number of animals counted fleeing from a particular block when it is driven depends on the group sizes and the spatial locations of groups at the time of the count. These two factors cannot be distinguished from the counts, as groups may fragment or coalesce during the animals' flight. Roe and red deer (Cervus elaphus L.), the two most widely distributed deer species in Europe, differ in their social organization systems, and as a result, group sizes formed in forest environments typically differ (e.g. Dzięciołowski 1979). Roe deer are usually solitary or form small family groups (Hewison et al. 1998), while red deer are gregarious and exhibit much larger group size (Clutton-Brock et al. 1982). Therefore, the statistical distribution of the numbers within each group will likely differ between species, as red deer are more aggregated. Aggregation within groups due to social behaviour may be further enhanced by spatial auto-correlation between groups due to differential habitat use (Welch et al. 1990;Palmer and Truscott 2003;Borkowski 2004). On the other hand, deer group size tends to be affected by activity and period of day (Borkowski and Furubayashi 1998). Drive counts are conducted during daylight when deer are predominately inactive and rest in small groups (Dzięciołowski 1979;Thirgood and Staines 1989;Carranza et al. 1991), which may reduce the difference in aggregation between the two species.
Usually, at least 10% of the total forest area is recommended to be covered by drive counts (Pucek et al. 1975;Nasiadka 1994). However, little is in fact known of how the total area and number of blocks driven influence the results. Similarly, there is no information on how population density and group size affect the results of drive counts. Answering these questions through field studies, however, would be challenging. Even if it were logistically possible to conduct field experiments to examine the effects of such variables, their influence on accuracy cannot be determined unless the true population size is known, which is rarely the case (e.g. Daniels 2006). However, counting methods may be compared in a computer simulation in which total population size is controlled (Smart et al. 2004).
Here, we use computer simulation to evaluate effects of block size, proportion of forest counted, density, social aggregation, and spatial auto-correlation on accuracy of drive counts in a series of 'virtual experiments'. We base the experimental treatments on an analysis of drive count data from commercial forest districts in Poland.
Drive counts were conducted within four commercial forest districts in Poland: Pszczyna, Rudy Raciborskie, Strzałowo, and Iława. Depending on the forest district, the counts were done for one to three consecutive years (
Table 1). Pszczyna and Rudy are located in the Silesian Upland near Gliwice city, south-western Poland (50°45′ N, 18°40′ E), while Iława and Strzałowo are in the Mazurian region near Olsztyn city, northern Poland (53°47′N, 20°30′E). The climate of these regions is typical for central Europe, where oceanic and continental climate types meet. However, in the Silesian Upland, mean annual temperature is higher (ca. 9 C) than in the Mazurian Distict (ca. 6.6 C). Mean annual precipitation in both regions is similar (ca. 600 mm).
Drive counts were used to estimate winter numbers (February-March) of red and roe deer. Each individual area driven was a block of one to a few adjacent forest compartments (on average ca. 60 ha). Usually there were 15-20 beaters and the same number of observers participating in the counts. The observers (either foresters or hunters) had sufficient experience to determine deer species, sex, and age (young/adult). Each observer recorded on an observation form the species and number of individuals of each group (and if possible also the group composition) leaving (or entering) the driven block on his right side. A coordinator collected the same information on animals seen by the beaters. After beating each block, the coordinator collated information from all observers and immediately resolved any possible inconsistencies, in order to minimise the likelihood of double counting and inaccurate group sizes. In the majority of cases, the same blocks were beaten from year to year.
We examined the degree of dispersion of red and roe deer block counts by fitting to generalised linear mixed models (GLMM) having a Poisson error term, logarithmic link function and the logarithm of block area (ha) as an offset. District was fitted as a fixed effect, and block and district× year as random effects. 'Year' was not the same at each site, and was therefore not included in the model as a fixed effect. The roe deer count was subsequently added to the red deer model, and similarly the red deer count to the roe deer model, to test whether there was an interrelation between the two species at the block level.
The suitability of the negative binomial distribution for representing aggregation in a population simulation model was assessed. To do so, block counts were standardised to a 60 ha block area (real block sizes ranged from 30 to 118 ha and district means from 45 to 81 ha) in order that the arithmetic mean, variances and coefficient of variation (c.v.) at the block scale could be estimated for the two species in each district/year combination. The mean and variance were then used to estimate the negative binomial aggregation parameter (k) using the method of moments (Taylor et al. 1979).
We adapted the method of Travis and Palmer (2005) to generate simulated populations of red and roe deer in a virtual forest comprising a set of contiguous square 20 ha compartments. The forest (total area 18,000 ha, similar to typical Polish forest districts) comprised 30×30 compartments. Two types of simulated population were generated using the Macro Facility of SAS (version 9.1):
1. Spatially unstructured. For each deer species independently, the first animal was located within a random compartment; all subsequent individuals were either placed in a random compartment with probability z or, with probability 1-z, placed within the same compartment as the previous individual. Thus, the parameter z controls the degree of demographic aggregation (i.e. within group), and the smaller its value, the greater the demographic aggregation (at the compartment scale) within the population. 2. Spatially auto-correlated. An additional spatial autocorrelation parameter s was introduced, which behaved in a similar manner to z, but controlled the aggregation between groups. If the animal was to be placed (as determined above) in a different compartment to the previous individual, then with probability s it was placed in a random compartment and with probability 1-s it was placed in a compartment adjoining that of the previous animal, one of the four cardinal directions being selected at random. Thus, the smaller the value of s, the greater the spatial auto-correlation of groups within the population. Values of s lower than 0.65 typically produced significant spatial auto-correlation as measured by Moran's I statistic (ArcGIS version 9.1).
Prior to conducting sample counts on the virtual populations, the effects of random variation and of the grouping probability z were examined in a series of trials on spatially unstructured populations to determine whether the simulation algorithm could generate realistically distributed deer populations.
A series of virtual counting experiments was then conducted by generating random sets of counting blocks akin to the blocks used in field counts. Each block comprised either a single compartment or a contiguous set of compartments of specified size running either east-west or north-south (selected at random), which constituted the simplest way to simulate blocks having odd numbers of compartments whilst avoiding irregular shapes. Blocks including a compartment previously allocated to another block were discarded; however, there was no bar to two or more blocks sharing a common edge (Fig. 1). Adjacent blocks are unlikely in reality, but their presence in the simulation does not affect the results. Each experiment was replicated across 20 different simulated populations, each of which was counted 100 times to estimate the mean and range of two types of 'accuracy indicators' for each deer species: (1) the percentage of counts where the estimated total population fell within a specified range (±10%, 20%, or 30%) of the true total population and (2) the estimated population expressed as a percentage of the true population. Strictly speaking, the first of these reflects statistical precision (how close repeated measures are to each other) rather than accuracy (how close a measure is to the true value). However, from the point-of-view of the forest manager, who may have sufficient resources for only a single measurement, the difference is purely semantic, and he is interested in how close his estimate likely to be to the true value; hence, we here use the term 'accuracy'.
The first three experiments examined the effects on count accuracy of three factors forest managers cannot control when planning a count, namely deer density, demographic aggregation and spatial auto-correlation. The last two experiments tested whether two factors within managers' control, block size, and the total area counted, can reasonably be manipulated to improve count accuracy.
Block size (three compartments, i.e. 60 ha), total area counted (10% of the forest, i.e. 30 blocks) and degree of aggregation (z=0.5 for red deer, 0.8 for roe deer, which were found to reproduce the degree of aggregation observed in field counts) were held constant. The population density of each species was varied between 2 and 22/100 ha in increments of 4/100 ha.
Experiment 2: the effect of demographic aggregation in spatially unstructured populations Block size (as Experiment 1), total area counted (as Experiment 1) and population density (red deer 10/100 ha, roe deer 7.5/100 ha) were held constant. The degree of aggregation z of each species was varied between 0.30 and 0.60 for red deer and between 0.60 and 0.90 for roe deer in increments of 0.05 to span the values used in Experiment 1.
Experiment 3: the effect of spatial auto-correlation Block size (as Experiment 1), total area counted (as Experiment 1), population density (as Experiment 2) and the aggregation parameter z (as Experiment 1) were all held constant. The spatial auto-correlation parameter s was varied between 0.1 and 0.9 for each species in increments of 0.2 to produce a wide range of possible spatial auto-correlation.
The effect of increasing or decreasing the block size (and altering the number of blocks to count the same total area) was examined for low (red deer 4/100 ha, roe deer 3/100 ha) and high (red deer 12/100 ha, roe deer 20/100 ha) population densities (typical of Polish forests). The populations were (a) highly spatially aggregated (s=0.3 for both species, z=0.4 for red deer, and 0.7 for roe deer) or (b) relatively unaggregated spatially and with group sizes as applied in Experiment 1 (s=0.9 for both species, z=0.5 for red deer, and 0.8 for roe deer). Counting blocks of 1, 2, 3, 5, and 6 compartments (20, 40, 60, 100, and 120 ha, respectively) were employed. The total area counted was held constant at 10% of the forest, i.e. the number of counting blocks was set to 90, 45, 30, 18, and 12, respectively.
Experiment 5: the effect of increasing the total area counted to improve accuracy For populations in which there is a high degree of overdispersion due to demographic aggregation and/or spatial auto-correlation, count accuracy might be improved by increasing the proportion of the forest counted. In turn, that could be achieved by counting more blocks and/or increasing block size. To test this, we simulated block sizes of 60 and 100 ha to count 10%, 20% and 30% of the forest, at the low and high population densities and levels of aggregation specified for Experiment 4.
Deer density estimates based on the drive count method varied considerably between years in the three districts 2). The GLMM residuals for roe deer were moderately over-dispersed (scale disper-sion=2.3; n=132) and for red deer were highly overdispersed (scale dispersion=7.4; n=132). When data were restricted to blocks with non-zero counts only, the same patterns were observed (roe 1.8; n=109, red 4.2, n=95). Thus, over-dispersion across all blocks was not simply due to some blocks being unoccupied (e.g. unsuitable habitat, disturbance) and deer being distributed between all occupied blocks at random. Rather, it was a genuine result of aggregation patterns at the block scale.
There was no difference in density between districts (having fixed the scale dispersion parameter at unity) for either species (red deer: F 3,4 =1.4, P=0.38; roe deer: F 3,6 =0.98, P=0.46). There was no evidence that the count of either species was related to the presence or count of the other species at the block level (effect of roe deer on red deer: F 1,83 =1.9, P=0.18; effect of red deer on roe deer: F 1,83 =0.95, P=0.33). There was no effect of block area on estimated density within the block for either red or roe deer (red deer: F 1,50 =0.55, P=0.46; roe deer: F 1,44 =0.75, P=0.39); nor was there any effect of block area on the probability that at least one animal was recorded within the block (red deer: F 1,31 =0.33, P=0.57; roe deer: F 1,30 =0.12, P=0.73). In nine of ten counts, the c.v. of the density estimate for red deer (range, 83-186%) was higher than for roe deer (58-120%) and k (the negative binomial aggregation parameter) was lower (red, 0.32-1.80 and roe, 0.94-4.55), reflecting the greater degree of aggregation amongst red deer (although the negative binomial parameters were not independent-see Appendix).
Replicated stochastic trials in which mean red deer density was set at 2.0/compartment (equivalent to 10/100 ha) and mean roe deer density at 1.5/compartment (7.5/100 ha) indicated that values of the aggregation parameter z in the range 0.25 to 0.60 for red deer and 0.60 to 0.85 for roe deer gave realistic values of k and c.v. within the ranges observed from field counts. The parameter z was then fixed at 0.45 for red deer and for roe deer at 0.70 and density was varied for each species between 0.5 and 4.0/compartment. For both species, simulated count data fitted Taylor's power law (Taylor et al. 1978(Taylor et al. , 1979; see Appendix) closely (P < 0.001, R 2 = 0.99 in each case). Neither estimated exponent differed significantly from unity, and estimates of the scaling parameter a were 3.35 and 1.81 for red and roe deer respectively. Thus, for both simulated species, the variance increased more rapidly than the mean, but linearly in relation to the mean, in a similar fashion to counts obtained from real forests (see Appendix). Thus we concluded that the simulation algorithm was able to generate population distributions which displayed the characteristics of real-forest deer populations.
For both deer species, count accuracy increased with density at all accuracy levels assessed, i.e. the proportion of estimates falling within 10, 20 or 30% of the true population (F 5,95 >336, P<0.001 in all cases; Fig. 2). At all but the lowest density, 2 deer/100 ha, counts of both species fell within 20% of the true population most of the time (at least 81% for red deer and 92% for roe deer). However, the expectation of an estimated count falling within 10% of the true total declined quite sharply as density decreased. Below 5 deer/100 ha, fewer than around half of roe deer counts and fewer than 40% of red deer counts would be expected to be that accurate. In the worst-case forests (i.e. the individual population replicates having the lowest accuracy index at each deer density), only 19% of red deer and 31% of roe deer counts achieved the 10% accuracy threshold at the lowest density. Moreover, at low density, the estimates were highly inaccurate, ranging from 19% to 203% of the true total for red deer, and 42% to 169% for roe deer. In comparison, at the highest density (22 deer/ 100 ha), accuracy ranged from 72% to 132% for red deer and from 81% to 118% for roe deer. At the 10% level of assessment, the accuracy of counts differed between the two species at all densities (pairwise t tests implemented in linear mixed model: t 105 >7.4, P<0.0001 in all cases; compare solid lines in Fig. 2).
Year Iława Pszczyna Rudy Strzałowo Red deer Roe deer Red deer Roe deer Red deer Roe deer Red deer Roe deer 1993 10. 0 6.9 ------1994 --10.2 3.5 10.3 8.3 12.7 7.4 1995 13.5 4.5 5.3 11.3 5.1 8.3 --1996 8.6 4.9 8.0 3.3 6.1 13.6 --Mean 10.7 5.4 7.8 6.0 7.2 10.1 --
Table 2 Deer density estimates (deer/100 ha) using drive counts in four Polish forests
Experiment 2: The effect of demographic aggregation in spatially unstructured populations
Increasing the degree of demographic aggregation (reducing the value of z) reduced the accuracy of the count at all levels of accuracy assessment (F 6,114 >10.5, P<0.001 in all cases; Fig. 3). However, the magnitude of the effect of varying aggregation across its full range of likely values for each species at fixed density was substantially less than the magnitude of varying density across a tenfold range at fixed aggregation (Table 3). Overall, the worst-case estimates, occurring at the highest levels of aggregation, were 56% and 177% of the true red deer population and 62% and 146% of the true roe deer population.
Increasing the degree of spatial auto-correlation (reducing the value of s) reduced the accuracy of the count at all levels of accuracy assessment (F 4,76 >67.0, P<0.001 in all cases; Fig. 4). Although the accuracy of roe deer counts was significantly higher than that of red deer counts at the same level of spatial auto-correlation (owing to the lower degree of demographic aggregation in roe populations), the magnitude of the difference was quite small (Fig. 4; Table 3), suggesting that there was no important interaction between demographic aggregation and spatial auto-correlation.
Although there were significant improvements in accuracy by reducing the block size to 20 ha at all levels of assessment for both species at low-density and high spatial aggregation (F 4,76 >4.5, P<0.01 in all cases), the gains were limited (Fig. 5a, b). For example, changing the block size from 60 to 20 ha (and increasing the number of blocks (a) Red deer 0 20 40 60 80 100 0 5 10 15 20 25 Deer density (animals / 100ha) 0 5 10 15 20 25 Deer density (animals / 100ha) Accuracy index (%) Within 30% -mean Within 20% -mean Within 10% -mean Within 10% -worst case (b) Roe deer 0 20 40 60 80 100 Accuracy index (%) Within 30% -mean Within 20% -mean Within 10% -mean Within 10% -worst case
Fig. 2 The accuracy of simulated estimated counts of a red deer and b roe deer in relation to density. The accuracy index shows the proportion of counts falling within a specified percentage of the true population total. Each count covered 10% of the forest using 30 blocks of 60 ha each. Means were derived from 20 replicate virtual forests threefold) would be expected to increase the percentage of total population estimates within 10% of the true red deer total from 27% to 32%; the corresponding figures for roe deer were 34% and 41%. In no case did increasing the block size to 100 or 120 ha make any difference in the accuracy attained with 60 ha blocks. Counts of roe deer (having the lower degree of demographic aggregation) were more accurate for a given block size than those of red deer, but only at the 20% level of assessment was there any significantly greater effect of altering block size on roe deer than on red deer counts (F 4,76 =4.2, P<0.01). Even at the 20 ha block size, in these low-density forests with highly spatially aggregated populations, the proportion of counts falling within 10% of the true populations were as low as 20% for red deer and 34% for roe deer, and estimates ranged from 26% to 201% of the true red deer population and from 44% to 170% of the true roe deer population.
At a high density of red deer and high spatial aggregation, similar effects of changing block size were observed (Fig. 5c), albeit at levels of accuracy approxi-mately 20% higher than for a low-density population. However, increasing block size to 100 or 120 ha had a small detrimental effect on count accuracy at high density, whereas it had no effect at low density. In contrast, as the high density of roe deer was substantially greater, density compensated for inaccuracies due to aggregation, and at all block sizes nearly all counts were within 20% of the true population (Fig. 5d). Only at the 10% level of accuracy assessment was there a meaningful significant effect of block size (F 4,76 =27.0, P<0.001); reducing block size from 60 to 20 ha would be expected to increase the number of estimates falling within 10% of the true population by about 9%.
In contrast, for populations of either species at low and at high density, and which were relatively unaggregated spatially, changing the block size had negligible beneficial effect on accuracy (not shown; F 4,76 <2.6, P>0.045 in all cases). Experiment 5: the effect of increasing the total area counted to improve accuracy
Although there were significant differences in count accuracy between 60 and 100 ha block sizes (the former being more accurate, in line with the results of Experiment 4), they were relatively small compared with the effect of area counted (maximum difference of 5% between block sizes at the same total area at the 10% accuracy level), and have therefore been averaged for clarity. Increasing the proportion of forest counted increased accuracy for lowdensity populations of both species at high spatial aggregation (F 2,97 >242, P<0.001 in all cases; Fig. 6a, b), although the absolute improvement in accuracy varied between species and accuracy level. For example, (1) doubling the area counted increased the mean proportion of red deer estimated counts falling within 10% of the true population from 28% to 41%, and (2) if a level of accuracy within 20% of the true roe deer population were considered acceptable, then increasing the proportion of forest counted from 10% to 20% gave a greater improvement in the frequency of accurate counts (by 28%) than increasing the proportion from 20% to 30% did (by 11%). For highdensity populations at high spatial aggregation, results were similar to those of Experiment 4, i.e. the gains in the proportions of estimates falling within 30% or 20% of the true total were limited because there was already a high degree of accuracy if only 10% of the forest were counted (Fig. 6c, d). The only meaningful improvement in accuracy by increasing area counted occurred for the proportion of counts falling within 10% of the true total, and was greater for an increase in forest area from 10% to 20% than for an increase from 20% to 30%. Similar changes occurred in response to area counted for populations of both species which were relatively unaggregated spatially (not shown), although at the 30% accuracy level there was no improve-
(a) Red deer, low density 0 20 40 60 80 100 0 20 40 60 80 100 120 Block size (ha) 0 20 40 60 80 100 120 Block size (ha) 0 2 0 4 0 6 0 8 0 1 0 0 1 2 0 Block size (ha) 0 2 0 4 0 6 0 8 0 1 0 0 1 2 0 Block size (ha) Accuracy index (%) Within 30% -mean Within 20% -mean Within 10% -mean Within 10% -worst case (b) Roe deer, low density 0 20 40 60 80 100 Accuracy index (%) Within 30% -mean Within 20% -mean Within 10% -mean Within 10% -worst case (c) Red deer, high density 0 20 40 60 80 100 Accuracy index (%) Within 30% -mean Within 20% -mean Within 10% -mean Within 10% -worst case (d) Roe deer, high density 0 20 40 60 80 100 Accuracy index (%) Within 30% -mean Within 20% -mean Within 10% -mean Within 10% -worst case
Our simulated experiments showed that deer density was the most important factor influencing accuracy of drive counts. For example, if accuracy to within 10% of the true population is expected, then this can vary by as much as 50% between very low and very high-density populations of both species, whereas differences due to demographic and spatial aggregation are likely to result in at most a 25% difference in accuracy (Table 3). At high densities (>10 deer/100 ha), drive counts of spatially uncorrelated red and roe deer populations can be expected to be accurate to within 20% of the true value more than 90% of the time, but at lower densities they can be inaccurate. Nevertheless, it must be mentioned that, according to the simulations, even at low densities about 80% of estimates will fall within 30% of the true red and roe deer population. Using similar simulation methods, Smart et al. (2004) investigated three monitoring methods (faecal standing crop, faecal accumulation rate and distance sampling using thermal imaging) and found that although they differed in accuracy, all performed more poorly at low densities. The results of our simulations do not confirm McCullough's (1979) findings of underestimation at low density and overestimation at high density, since at all densities the true number could have been either under-or overestimated. However, our simulations assumed no measurement error (see below), while the probability of double counting in McCullough's (1979) study could have been high, owing to the population being estimated in a relatively small (ca. 520 ha) fenced area. In the opinion of Nasiadka (1994), drive counts underestimate population size by about 20% and the author suggested adding 20% to the estimated population size. However Nasiadka's (1994) recommendation is arbitrary and has no empirical basis. Our study has shown that at relatively high densities, drive count results are quite accurate and no correcting factor is needed. At low densities, the accuracy is lower and rather variable, and it is therefore difficult to propose any universal correcting factor. proportion of the forest counted. The accuracy index shows the proportion of counts falling within a specified percentage of the true population total. Means were averaged over 60 and 100 ha block sizes
The present study also indicates that drive count accuracy is influenced by demographic aggregation and spatial auto-correlation. At high levels of either, counts will be less accurate, although the effect of aggregation at moderate density (7.5 to 10 deer/100 ha) will not be as great as reducing density. It is not surprising that spatial auto-correlation had a similar effect on accuracy as did demographic aggregation, as both serve to increase variance between counting blocks, demographic aggregation at the scale of individual compartments, and spatial auto-correlation at the scale of neighbouring compartments. In reality, the scale at which spatial auto-correlation occurs could itself vary from one part of the forest to another, and auto-correlation could also be anisotropic (varying between directions, e.g. because of an environmental gradient). Analysis of red deer standing crop faecal counts from a Caledonian pine-wood in Scotland (raw means of 10 years' counts along permanently marked transects) indicated that habitat use by red deer was spatially auto-correlated within distances up to 2 km (Palmer, unpublished data). Spatial auto-correlation at larger scales than this will not pose a problem as long as the counting blocks are well spaced throughout the forest. It is spatial auto-correlation at scales close to the counting scale which serves to increase variance most, and hence decrease accuracy. Field counts record the combined effect of the two behavioural processes. Detailed radio-tracking data from many individuals would be required to determine how demographic aggregation and spatial autocorrelation interact to produce the patterns of numbers observed at the block scale. It should also be noted that we simulated aggregation processes at the compartment scale, whereas observed counts were at the block scale. That will tend to reduce apparent block-scale aggregation. Hence, since we applied observed block-scale aggregation to compartments, we have probably reduced the level of aggregation, and that would mean that real counts would be more inaccurate than we have estimated (but probably not by much).
As already mentioned, our simulations took no account of measurement error, owing to lack of empirical data. In real counts, deer theoretically could be either over-or underestimated. However, when 10% of the area is counted and the blocks are driven towards blocks previously counted, it seems that the risk of double counting is minimal. If blocks are distributed regularly throughout the forest, they will be far enough from each other that fleeing individuals are unlikely to stop in as-yet undriven blocks. Alternatively, there are two sources of error that could lead to underestimation. Firstly, some animals may leave a block if disturbed by observers and beaters taking up position around the edge of the block. If this were so, recorded deer density in larger blocks, in which animals should less likely be disturbed (as there is less edge per unit area), should on average be higher than in smaller ones, all other factors being equal. In such a situation, driving of larger blocks might be recommended. Flight behaviour can vary considerably between, and even within, individuals (Sunde et al. 2009), but flight distances of roe deer have been found generally to be less than 100 m on average (de Boer et al. 2004), substantially less than the dimension of typical counting blocks. Moreover, as we found no effect of block area on estimated density within the block for either species in our analysis of data from Polish forests, nor any effect on the probability of recording at least one animal within the block, underestimation due to observer disturbance seems unlikely. Secondly, animals might remain undetected within a driven block. Unfortunately, we have no data on this issue. This may occur especially in areas where animals are accustomed to human presence and therefore reluctant to flee. On the other hand, even in areas where animals tolerate people well and flight distance is short, flight frequency increases when people behave in an unusual way (for instance walk away from trails) (Borkowski 2001). Moreover, owing to their relative sizes, we consider that this issue may be much less important for red deer that for roe deer. No matter how animals behave in reality, maintaining close proximity between beaters to maintain visual contact between them, even in blocks with relatively poor visibility, and using a dog to flush deer from dense cover, should reduce the chance of this sort of error (but see Staines and Ratcliffe 1987). As the size of blocks has little effect on drive count accuracy, block size should be adjusted depending on visibility and number of participants. Poor-visibility blocks should be smaller and driven by a relatively large number of beaters, while surrounded by fewer observers. To compensate, the number of blocks should be increased to maintain the total proportion of forest counted. In addition, drive counts should be organized during the leafless winter period, when visibility in most areas is better.
Our field data showed that levels of aggregation of red and roe deer differ markedly. Red deer distributions were more clumped than those of roe deer, even though daytime counts probably represented mostly inactive individuals. As the more gregarious species, the degree of aggregation of red deer is probably higher than that of roe deer even in the case of inactive individuals. To some extent, differences in the distributions of both species may arise from dissimilarities in habitat use (spatial auto-correlation) evoked by availability of food (Palmer and Truscott 2003) and/or cover (Borkowski 2004;Borkowski and Ukalska 2008). It has been demonstrated, for instance, that red deer as the larger species may be more demanding toward cover condition than smaller roe deer (Borkowski and Ukalska 2008). This may be especially important for resting individuals during day time, i.e. for animals predominately recorded using drive counts.
One may suggest that drive counts are a more reasonable method for roe than for red deer. Firstly, owing to the lower degree of aggregation in roe deer, drive counts are expected to be slightly more accurate than in the case of red deer. Secondly, at least in some areas within the range of both species, roe deer probably occur at higher densities than red deer, though comparative data are rather limited. For instance, in nearly 200 different hunting districts managed by the State Forest Agency in Poland, according to official statistics, an estimated density of 5 deer/100 ha or higher (our simulations suggest that at density >5 deer/100 ha drive count accuracy increases) was recorded only in 3% of districts for red deer, but in 43% for roe deer (Borkowski, unpubl. data). However, according to drive count results, red deer densities are higher (e.g. see Table 2). Moreover, in three forest districts of Białowieża Forest, Poland, where red deer densities are known to be among the lowest in the country, red deer density was recently estimated (by drive counts) at between 5.1 and 7.2 individuals/100 ha (Borkowski et al. unpubl. data). Also, in 20 Scottish forests, densities of both species were more similar, and a density >5 deer/100 ha was recorded only slightly more often for roe deer (13 forests) than for red deer (ten forests) (Latham et al. 1996, Tab. VI, p. 295). Therefore, it can be concluded that the method seems suitable for both deer species in areas where they occur at densities of at least five to seven animals/100 ha. Due to difference in size between both deer species, as mentioned earlier, the method may be even less accurate for roe deer due to measurement error, but no data on this issue are available. Thus, we urge caution when estimating population density by drive counts, especially at low densities. In such cases, it may be appropriate to assess the accuracy of trend detection by drive counts, in a similar manner to that of Smart et al. (2004).
We have demonstrated in this paper how a 'virtual ecosystem' can be used to examine the effects of system parameters on the behaviour of a clearly defined but complex system for which real experiments would be logistically or economically difficult or impossible. We used a 'virtual ecologist' to obtain replicate samples using simulated field counting techniques from a known population (Green and Sadedin 2005). Virtual ecosystems, frequently incorporating an individual-based model (IBM), are increasing in use and application in ecology (Grimm et al. 1999;Hirzel et al. 2001;Tyre et al. 2001;Harris et al. 2008). Here, we did not employ an IBM as such, although we did allocate the virtual deer to specific compartments on an individual basis. However, it is straightforward to recognise how an IBM might be incorporated, for example to model the spatial behaviour of individual deer in response to conspecifics and/or disturbance. The use of virtual experiments to inform forest management appears to be in its infancy, although Wunder et al. (2008) have recently used the technique to examine how well alternative sampling strategies could estimate growth-mortality relationships. Smart et al. (2004) performed a similar computerbased simulation to ours, but did not refer to it as a virtual experiment.
So, could forest managers improve the accuracy of counts by manipulating block size and the total area counted? At the lowest densities likely to be encountered in Polish forests, reducing block size and increasing the number of blocks counted can compensate to a limited extent for inaccuracies inherent in counting low-density spatially aggregated populations. However, it is unlikely that the limited improvement would justify the increases in logistic effort. Enlarging the size of blocks and decreasing their number would not improve count accuracy. Thus, for high-density populations, we do not recommend increasing the block size, as it would have detrimental effects. Decreasing block size would improve the accuracy of red deer counts slightly, but would have little effect for high-density roe deer populations. Increasing the total area of forest counted would compensate more for inaccuracies in estimating the total population, especially for highly aggregated, low-density populations. Whether the gains in accuracy could justify the extra effort required would have to be evaluated on a case-by-case basis, and would depend on available resources, cost, logistical issues, etc. However, our study suggests that at higher deer densities, drive counts can provide reliable information on population size, subject to appropriate correction for measurement error. Drive counts are also expected to be more accurate in forests where spatial aggregation is likely to be low (owing to either large-scale uniformity or high heterogeneity at small scales) than in forests where it is likely to be higher (comprising large block of uniform structure).
1. Drive counts can be recommended in forests with relatively high deer densities, but are expected to be less reliable in areas with low deer densities. The threshold density for the use of drive counts depends on the level of accuracy which is deemed acceptable. 2. It seems sufficient to drive 10% of the total area for relatively high-density populations. Driving up to 30% of the area brings some increase in accuracy, especially at low deer densities, but whether the gain in accuracy justifies the extra effort required needs careful evaluation. In addition, high total forest area counted may increase the risk of double counting.
3. For a given percent of total area counted, driving more blocks of small size provides slightly higher accuracy than driving fewer larger blocks.
correlated with the true density (Spearman r = 0.58, P<0.10). For roe deer, the estimated exponent was greater at 1.37 (and there was insufficient evidence to be sure that this differed from 1.0, F 1,8 =2.8, P=0.13; Fig. 7). This was also indicative of a type I response, but for roe deer there was no significant relationship between 1/k and the mean (Fig. 7). Thus, for red deer at least, as overall density increases, we expect apparent aggregation, as measured by 1/k, to decrease; i.e. k and the mean are not independent parameters for real-forest deer populations. Our data for red deer, and to a lesser extent for roe deer, support the power law relationship of the variance to the mean (Taylor et al. 1979), although we acknowledge that our estimates of its parameters were made from small samples, which were not fully independent. The negative binomial aggregation parameter is not independent of the mean, and the negative binomial does not therefore constitute a sound basis for analysing deer count or faecal count data (White and Eberhardt 1980;White and Bennetts 1996). The power law relationships arise from a combination of within-and between-compartment variation in group size, the former principally due to herding behaviour and the latter to spatial auto-correlation. Although we might ships of a, b the count variance and c, d the negative binomial aggregation parameter k with the mean count. Count data were standardised to a 60 ha block size
expect the herding behaviour of a deer species to relate to density in the same way across different sites (provided that habitat and perceived predation threats were similar), spatial auto-correlation could be site dependent.
Tyre AJ, Possingham HP, Lindenmayer DB (2001) Inferring process from pattern: can territory occupancy provide information about life history parameters? Ecol Appl 11:1722-1737 Ueckermann E (1964) Der Rehwildabschuss. Verlag Paul Parey, Berlin, pp 1-76 Welch D, Staines BW, Catt DC, Scott D (1990) Habitat usage by red (Cervus elaphus) and roe (Capreolus capreolus) deer in a Scottish Sitka spruce plantation. J Zool Lond 221:453-476 White GC, Bennetts RE (1996) Analysis of frequency count data using the negative binomial distribution. Ecology 77:2549-2557 White GC, Eberhardt LE (1980) Statistical analysis of deer and elk pellet-group data. J Wildl Manage 44:121-131 Wunder J, Reineking B, Bigler C, Bugmann H (2008) Predicting tree mortality from growth data: how virtual ecologists can help real ecologists. J Ecol 96:174-187
Acknowledgements SCFP was supported by the
Open Access This article is distributed under the terms of the Creative Commons Attribution Noncommercial License which permits any noncommercial use, distribution, and reproduction in any medium, provided the original author(s) and source are credited.
The variance of red deer block counts was related to the mean by a power law relationship whose exponent was estimated to be 1.03 (Fig 7), indicating a type I response curve of Taylor et al. (1979), having a single turning point. Such a relationship typically results in an inverse correlation between 1/k and the mean, and this was observed (Spearman r=-0.67, P<0.05; Fig. 7). Hence k was also
The phlebotomine sand fly Lutzomyia longipalpis takes blood from a variety of wild and domestic animals and transmits Leishmania (Leishmania) infantum chagasi, etiological agent of American visceral leishmaniasis. Blood meal identification in sand flies has depended largely on serological methods but a new protocol described here uses filter-based technology to stabilise and store blood meal DNA, allowing subsequent PCR identification of blood meal sources, as well as parasite detection, in blood-fed sand flies. This technique revealed that 53.6% of field-collected sand flies captured in the back yards of houses in Teresina (Brazil) had fed on chickens. The potential applications of this technique in epidemiological studies and strategic planning for leishmaniasis control programmes are discussed.
Zoonotic visceral leishmaniasis (ZVL) represents a serious threat to public health in both the Old and New World. It is the most severe form of leishmaniasis, with fulminating disease usually being fatal if untreated and most prevalent in malnourished children (Cerf et al., 1987). About 90% of cases occur in economically deprived rural and suburban areas in three geographical regions: the Indian subcontinent, East Africa and South and Central America, with Brazil contributing most cases in the New World (Dantas-Torres and Brandão-Filho, 2006).
During the last 20 years ZVL due to Leishmania (Leishmania) infantum chagasi has become increasingly urbanised and is now present in several Brazilian cities (Michalsky et al., 2007). Although more common in socio-economically disadvantaged populations occupying the "vilas" or "favelas", the juxtaposition of these areas with wealthy neighbourhoods means that ZVL is a potential threat to the whole population (Alexander et al., 2002). The sand fly Lutzomyia longipalpis (Lutz and Neiva, 1912) is the urban vector of Le. infantum in Brazil.
Originally associated with open Savannah, it is able to colonize urban areas and other habitats affected by human activities, taking blood meals from a wide variety of hosts, including livestock, dogs and chickens (Lainson and Rangel, 2005). Widely differing observations have been made regarding the degree to which Lu. longipalpis bites humans in different habitats, indicating that this species has no strong innate host preference. The absence of a common term to distinguish Lu. longipalpis from other biting insects in urban ZVL foci suggests that this fly does not constitute a substantial biting nuisance for the inhabitants (Alexander et al., 2002).
Identification of the blood meals of haematophagous insects to date has largely depended on serological techniques such as the precipitin test, latex agglutination test and the enzyme-linked immunosorbent assay (ELISA) (Boorman et al., 1977;Washino and Tempelis, 1983;Beier et al., 1988;Gomes et al., 2001;Mwangangi et al., 2003). Although these methods have yielded important information on the identity of the vertebrate hosts of many blood-feeding arthropods, they are time-consuming and lack sensitivity. PCR-based identification of vertebrate host blood meals is a potentially convenient alternative, which has already been performed on several vectors including ticks (Pichon et al., 2003;Estrada-Peña et al., 2005), triatomine bugs (Bosseno et al., 2006;Pizarro et al., 2007) and mosquitoes. PCR based on primers designed from multiple alignments of the mitochondrial cytochrome b gene have identified avian and mammalian hosts of various species of mosquito (Ngo and Kramer, 2003;Kent and Norris, 2005;Molaei et al., 2006;Kent et al., 2006). PCR-RFLP cytochrome b analysis was also used to identify the origin of blood meals in the tick Ixodes ricinus (Kirstein and Gray, 1996) and tsetse flies (Steuber et al., 2005).
Until recently sand fly host identification by blood meal analysis has been limited to serological studies using ELISA (Gomez et al., 1998;Agrela et al., 2002;Bongiorno et al., 2003;Svobodová et al., 2003;Marassá et al., 2006;Rossi et al., 2008), counter immunoelectrophoresis (Morsy et al., 1993), agarose gel diffusion (Srinivasan and Panicker, 1992), precipitin test (Tesh et al., 1971(Tesh et al., , 1972;;Javadian et al., 1977;Morrison et al., 1993;Nery et al., 2004;Afonso et al., 2005) and a more laborious histological technique (Guzman et al., 1994). The first PCR-based method using the prepronociceptin gene has been recently described (Haouas et al., 2007). One of the major constraints in developing a successful PCRbased method for sand flies is that they are diminutive insects, only able to ingest very small quantities of blood (Rogers et al., 2002). Template stability and the lack of uniformity in DNA template concentration are also a concern, particularly as template degradation during blood digestion has been observed for mosquitoes (Mukabana et al., 2002;Ngo and Kramer, 2003;Kent and Norris, 2005). However, preservation of specimens under field conditions coupled with the need for a rapid and sensitive method to identify sand fly hosts is important for the study of vector-vertebrate host associations, as well as to improve current control interventions targeting Lu. longipalpis.
The FTA filter methodology has been shown to be a useful tool in the storage and PCR detection of DNA in a wide variety of organisms, including trypanosome identification in wild tsetse populations (Adams et al., 2006), diagnosis and surveillance of Trypanosoma vivax in ruminants (Gonzales et al., 2006), PCR detection of Cyclospora cayetanensis and Cryptosporidium parvum from food samples and human faecal specimens (Orlandi and Lampel, 2000), as well as archiving and processing DNA from fresh water protozoans (Hide et al., 2003) and extraction and storage of insect DNA in forensic entomology (Harvey, 2005). In the present study, we adapted a multiplex PCR protocol (Kent and Norris, 2005;Ngo and Kramer, 2003) and combined this with a FTA-based technology to store and analyse sand fly samples from the field, in order to identify putative hosts of this important disease vector in urban habitats.
A laboratory colony of Lu. longipalpis established from sand flies caught in Jacobina (Bahia, Brazil) and kept at the Liverpool School of Tropical Medicine was used in the blood-feeding time course experiments and maintained using standard methods (Modi, 1997). Insects were reared under controlled conditions of temperature (28 ± 1 °C) and humidity (80-95%) and the adult female insects were fed on hamsters twice a week. All procedures involving animals were approved by a local Animal Welfare Committee and performed in accordance with UK Government (Home Office) and EC regulations. Insects were provided with 70% sucrose for 96 h post-emergence until they were blood-fed for DNA degradation experiments.
Field specimens of Lu. longipalpis were collected by hanging CDC light traps (Alexander, 2000) from 18:00 to 08:00 h for a total of 65 trap-nights during May-June 2007 in houses of two neighbourhoods (Satélite and Ininga) of Teresina (5°5′20″S and 42°48′07″W), capital of the Brazilian state of Piauí. Active transmission of both canine and human ZVL occurs in this city (Werneck et al., 2007). Twenty-four traps were set up in Ininga, a relatively high-income neighbourhood near the University of Piauí, whereas 41 were placed in Satélite, a lower income neighbourhood. Live insects were collected in the morning following capture and transported to the lab within 2 h. They were then identified according to Young and Duncan (1994), and were either immediately frozen at -20 °C for subsequent whole DNA extraction or homogenised in 0.15 M saline using a cordless motorized pestle and spotted onto Whatman ® FTA cards for DNA storage/extraction.
Whole genomic DNA was extracted from laboratory-fed single sand flies using the FastDNA SPIN Kit for Soil (Q BIOgene). The manufacturer's protocol was modified and optimized for sand fly genomic DNA extraction. Briefly, 20 human blood-fed and 30 chicken blood-fed sand flies were homogenised individually in 1.5 mL microcentrifuge tubes with 244.5 μL sodium phosphate buffer and 30.5 μL MT buffer (both provided with the kit) using a cordless pestle pellet motor on ice. After centrifugation at 14,000 × g for 8 min, the supernatant liquid was transferred to a clean tube, followed by the addition of 65 μL of protein precipitation solution (PPS). After another centrifugation step (14,000 × g for 7 min), the supernatant was transferred to another tube and mixed with 0.5 mL of binding matrix. Tubes were then placed in a rotator for 2 min to agitate the mixture gently and allow DNA binding with the matrix, and subsequently placed in a rack for 3 min for silica matrix settling. A 250 μL fraction of the supernatant was discarded and the remainder used to resuspend the matrix. The mixture was then applied to a SPIN filter cartridge and centrifuged at 14,000 × g for 1 min. Filters were then washed with 250 μL salt/ethanol wash solution (SEWS-M) and centrifuged at 14,000 × g for 1 min. After an extra centrifugation step to dry the matrix of residual ethanol, genomic DNA was eluted with 50 μL of RNase-free water after centrifugation at full speed for 1 min. DNA yield was quantified using a Nanodrop ND-1000 Spectrophotometer and approximately 10 ng DNA template for each sample was used in a multiplexed PCR reaction and visualised in 1.5% agarose/ethidium bromide gels.
Forty seven human blood-fed, 49 chicken blood-fed and 58 wild-caught blood-engorged sand flies were homogenised individually in 50 μL of 0.15 M saline using a cordless motorized pestle and the whole volume spotted onto a Whatman ® FTA card. The homogenates were allowed to air-dry and the FTA cards kept in sealed plastic bags at room temperature until required. For PCR, discs of diameter 2 mm were punched out and placed in a 1.5 mL microcentrifuge tube. FTA disks were washed three times with 0.5 mL of FTA purification buffer for 5 min, twice with 0.2 mL of 10 mM Tris (pH 8.0) containing 0.1 mM EDTA for 5 min at room temperature and then air-dried on a heating block at 56 °C for 10 min. These washed filters were used directly in a multiplexed PCR reaction and visualised in 1.5% agarose/ ethidium bromide gels.
PCR primers based upon alignment of cytochrome b sequences of mammalian and avian species were obtained from previously published primer sequences (Table 1). The multiplexed PCR was optimized for sand fly templates. An initial denaturation of 3 min at 94 °C was followed by 35 cycles at 94 °C for 1 min, 52 °C for 1 min and 72 °C for 1 min. The final extension step was 72 °C for 10 min. Field-captured engorged and unfed sand flies were also screened for Leishmania infection in a separate PCR reaction using previously published primer sequences, targeting either the kDNA or the small subunit of the rRNA gene (Lachaud et al., 2002).
Chicken blood in Alsevers anticoagulant was purchased from TCS Biosciences and human blood obtained from the National Blood Service (Speke, Liverpool, UK). For time-course experiments, sand flies were fed on chicken and human blood through chick-skin feeders at 35 °C and each group sampled at 24, 48, 72, 96, and 120 h intervals from the day of the feed (day 0). Sand flies were either frozen at -20 °C for genomic DNA extraction or homogenised in 50 μL of 0.15 M saline and spotted onto Whatman ® FTA cards for DNA extraction as above.
Frozen aliquots of Le. infantum amastigotes obtained from spleen homogenates of infected female BALB/c mice were rapidly thawed from liquid nitrogen and gently mixed with 5 mL complete M199 medium (Sigma). Parasites were centrifuged at 1500 × g for 10 min and washed with 5 mL M199 medium twice, before being re-suspended in 10 mL of complete M199 medium containing 10% FBS and transferred to a culture flask and cultured at 26 °C for 48 h. In preparation for infection, 2 mL of heat-inactivated rabbit blood was used to re-suspend cultured parasites to a final concentration of 1 × 10 6 parasites/mL. Rabbit blood seeded with parasites was offered to 5-day-old Lu. longipalpis through a chick-skin membrane and infected female sand flies collected daily from 0 to 5 days post-feeding.
Infected sand flies were homogenised in 50 μL of 0.15 M saline and spotted onto Whatman ® FTA cards. Discs were extracted as described above and part of the parasite small subunit ribosomal RNA (SSU rRNA) was amplified by PCR using the genus-specific primers R221 and R332 as described in Lachaud et al. (2002). To determine the sensitivity of parasite detection, 5 × 10 4 of cultured Le. infantum parasites were serially diluted and spotted onto Whatman ® FTA cards for DNA extraction and the parasite SSU rRNA detected as above.
Published as: Acta Trop. 2008 September ; 107(3): 230-237.
To optimise PCR conditions initial experiments were performed on DNA isolated from sand flies using a genomic DNA extraction kit. Time-course analysis on chicken and human-fed sand flies in the laboratory showed that host DNA could be detected for up to 48 h after the blood meal (Fig. 1). Prominent non-specific low-molecular weight products were observed in all lanes, including control reactions lacking DNA, and are often seen in multiplex PCR reactions. However, with human DNA two specific bands were detected, corresponding to the predicted human-specific 334 bp and universal mammalian 623 bp bands (Fig. 1A). The lack of products in samples taken after 48 h is presumably due to digestion of the DNA by sand fly gut enzymes. Using flies fed on chicken blood the predicted 210 bp band was observed (Fig. 1B). Next the same procedure was repeated, except that sand fly samples were homogenised and spotted onto FTA cards, then discs punched out and used in the multiplex PCR (Fig. 2). As before, specific bands were detected in sand flies up to 48 h after feeding on either human or chicken blood. Although the results were broadly similar, in comparison to directly extracted genomic DNA (Fig. 1), the human-specific bands were consistently weaker, whereas the chicken-specific bands were consistently stronger. In both sets of experiments the absence of specific bands after 48 h corresponds to the end of blood meal digestion, indicating likely template degradation by digestive enzymes.
In addition to blood meal identification, it would also be useful to detect the presence of Leishmania parasites in sand fly homogenates. Therefore, a sensitivity test was performed by preparing parasite serial dilutions, which were then spotted onto FTA paper and processed for PCR using primers targeting the parasite small subunit ribosomal RNA (Lachaud et al., 2002) (Fig. 3A). The predicted 603 bp product was detected, and this showed that the equivalent of as few as 49 parasites/original 50 μl sample could be detected. Given that the assay uses a 2 mm disc from a spot of approximately 15 mm diameter this represents a theoretical sensitivity of approximately a single parasite (∼77 fg DNA). To assess whether Leishmania infection could be detected in samples of lab-reared Lu. longipalpis, flies were experimentally infected with Le. infantum, homogenised individually in 0.15 M saline and spotted onto FTA filter paper. A time course analysis of individual sand flies homogenised at various time points postinfection was able to detect the presence of parasite DNA as early as 24 h (Fig. 3B). As expected, the band intensity increased with time as the infection matured. Nevertheless, the results show that sand flies sampled 24-48 h after feeding could be analysed for both blood meal identification and the presence of parasites.
Finally the FTA method was tested on wild-caught sand flies collected in Teresina, Brazil. A total of 2089 sand flies (1701 males and 392 females) were recovered from 65 CDC traps set up in 61 houses from two distinct neighbourhoods during May-June 2007. Of these, 1739 were identified according to Young and Duncan (1994), 1732 as Lu. longipalpis (99.5%), four as Lu. lenti (0.23%) and one each (0.06%) as Lu. whitmani, Lu. termitophyla and Lu. aragaoi. All 58 blood-engorged female sand flies were Lu. longipalpis and positive blood meal identifications were obtained for 43 of these (74%). An example of the results obtained using the FTA extraction diagnostic assay on wild-caught sand flies is shown in Fig. 4. Sand flies that had fed on dogs (e.g. lane 2) and chickens (e.g. lanes 5-11 and 13) were detected. Some flies did not yield any PCR products (e.g. lanes 4 and 16), whereas others amplified the 623 bp universal band (e.g. lanes 3, 12, 14 and 15). In the example shown, two of these (lanes 3 and 15) also included the 210 bp chicken-specific band and were therefore included as chickenfed sand flies in the overall analysis. Others (e.g. lanes 12 and 14) did not reveal any specific bands and were recorded as unidentified. In total, 41 (70.7%) of the 58 blood-engorged sand flies collected had fed on chickens, 2 (3.4%) had fed on dogs, and 15 (25.9%) could not be identified. These negative results could be due to loss of DNA during blood meal digestion, the presence of a small blood meal or feeding on a host not covered by the multiplex PCR Published as: Acta Trop. 2008 September ; 107(3): 230-237.
Sponsored Document Sponsored Document i.e., non-chicken, dog or human blood meals. Fifteen (53.6%) of 28 engorged sand flies collected outside henhouses were positive for chicken blood, compared to 26/30 (86.7%) of specimens collected inside these shelters. None of the blood-fed sand flies collected seemed to have fed on human hosts. The absence of two specific bands in the PCR reactions suggests that none of the engorged sand flies contained blood from more than one host species. Finally, 205/392 field-collected female sand flies (52.3%) were screened for Le. infantum infection using two different set of primers to amplify fragments of the parasite small subunit ribosomal RNA (SSU rRNA) or kDNA (Lachaud et al., 2002). Parasite DNA could not be detected in any of the wild-caught sand flies analysed, indicating that the infection rate in this sample was less than 0.49% (1 in 205).
We have adapted a PCR methodology first developed for mosquitoes (Kent and Norris, 2005;Ngo and Kramer, 2003) to identify blood meals from wild-collected sand flies in an urban area of Le. infantum transmission. This methodology was further developed for host identification by DNA immobilisation using FTA cards for long-term storage and subsequent PCR. Our results compare favourably with those recently described by Haouas et al. (2007), where genomic DNA was extracted from sand flies using a QiaAmp blood DNA mini Kit. In our study successful amplification was obtained when blood-engorged sand flies were homogenised in saline and spotted onto FTA databasing filter paper. FTA paper has a matrix designed to lyse cells upon contact and sequester DNA within the paper matrix (Smith and Burgoyne, 2004). FTA filters also eliminate laborious isolation and purification steps, preserving DNA integrity and eliminating potential sources of target DNA losses through degradative processes associated with conventional methodologies (Orlandi and Lampel, 2000). Moreover, the easy handling and long-term stability of DNA blood samples stored on FTA filter papers makes this methodology robust and ideal in field conditions, where samples are often collected far away from where they are processed. It also provides a convenient and safe way to send samples between laboratories without a cold chain or transport of liquids. As shown in Figs. 1 and 2, FTA extractions were as effective as the genomic DNA extractions currently used from frozen/dried field samples. This method will be extremely useful for entomologically based projects involving blood-feeding behaviour and vectorial capacity of sand flies in urban and rural endemic areas of VL.
The time-course experiments showed that host DNA was detectable in chicken and human-fed control sand flies up to 48 h after the blood meal (Figs. 1 and 2). Similar results were obtained by Haouas et al. (2007) in their study of Phlebotomus species caught in central Tunisia and for Anopheles species captured in the wild (Kent and Norris, 2005). Boayke et al. (1999) and Ngo and Kramer (2003) detected human blood meals in black flies and avian blood meals in Culex pipiens, respectively, up to 72 h post-blood feeding. Although sand flies recently fed on chickens contain nearly double their normal DNA content compared to human-fed sand flies due to the nucleated avian erythrocytes (unpublished data), chicken blood meals could not be identified after 48 h (Figs. 1B and 2B). This represents one of the major constraints of any methodology based on blood identification within the insect's midgut. Due to blood digestion process inherent to any blood-sucking insect, only recently engorged sand flies may be used for the analysis.
The origin of blood meals in 15/58 of field-caught sand flies could not be identified (including four collected inside a henhouse). For those field specimens where vertebrate host DNA identification failed (Fig. 4), host DNA might have been insufficient for amplification, the process of digestion may have denatured the DNA or the sand fly may have fed on an animal not included in the screening. In some of our field specimens that were captured inside a henhouse, sand flies might have fed on different hosts and used the shelter as a resting site Published as: Acta Trop. 2008 September ; 107(3): 230-237.
Sponsored Document Sponsored Document (Brazil et al., 1991). It is also worthwhile mentioning that the control band of 623 bp (common to all species in this study), derived from PCR reactions in which universal forward and reverse primers for the cytochrome b gene were used (Kent and Norris, 2005), could not be visualised in some of the PCR reactions. This was probably due to partial DNA template degradation in the midgut where one of the universal primers would anneal, probably accounting for reactions which only host-specific cytochrome b bands could be visualised. Sand flies ingest minuscule amounts of blood (Rogers et al., 2002) and there may be insufficient host DNA, especially in partially fed sand flies. Mukabana et al. (2002) showed that blood meals of Anopheles gambiae contained 2-82 ng of human DNA. More recently, Kent and Norris (2005) showed that at least 50 ng was necessary for visible amplification from mosquito abdomens. In our PCR amplifications, as little as 10 ng was sufficient for band visualisation using genomic DNA extractions (data not shown). Sand flies seem to concentrate their blood meals through prediuresis (Sádlová et al., 1998;Sádlová and Volf, 1999), including Lu. longipalpis (R. Dillon, unpublished observation) excreting fluid to concentrate proteins of the blood meals in a similar way to mosquitoes. This phenomenon is helpful, as it will increase the concentration of host DNA in the midgut and the probability of obtaining a good-quality template for PCR amplification.
Recently, PCR-based approaches have allowed diagnosis of infectious diseases, including Leishmania detection in human patients (Dweik et al., 2007;Kumar et al., 2007;Foulet et al., 2007), infected dogs (Solano-Gallego et al., 2007;Gomes et al., 2007;de Andrade et al., 2006) and phlebotomine sand flies (Paiva et al., 2006;Cabrera et al., 2002;Myskova et al., 2008;Ranasinghe et al., 2008). Laboratory infections of Lu. longipalpis with Le. infantum followed by whole sand fly homogenisation in Whatman filter paper (Fig. 3) showed that FTA technology can be also used for parasite DNA storage and parasite detection by PCR. Parasite DNA was successfully amplified using previously described genus-specific primers (Lachaud et al., 2002). This simple methodology could be very useful in epidemiological studies in endemic areas for leishmaniasis as specimens suspected to contain parasites could be easily stored and amplified for parasite detection/identification. In addition, safe parasite storage can be guaranteed as the reagents on the FTA paper are designed to kill pathogens upon contact and the papers protect DNA within the samples for several years at ambient conditions (Smith and Burgoyne, 2004).
Previous studies of Lu. longipalpis, for example in an endemic area of visceral leishmaniasis transmission in Colombia, have suggested this sand fly is an opportunistic feeder and is not highly anthropophilic nor strongly attracted to dogs (Morrison et al., 1993). In the same study it was reported that Lu. longipalpis preferred cows and pigs over chickens. This may be true in a rural environment where large domestic mammalian species occur in abundance, reflected in a reduced Lu. longipalpis vectorial capacity (Morrison et al., 1993). In the current study, 30/58 of the engorged sand flies were collected inside henhouses and the remaining 28 outside these structures. Fifteen of the latter insects had clearly fed on an avian host, most probably chickens (53.6%). This result is similar to previous studies on blood meal preference of Phlebotomus papatasi in an Iranian village where a precipitin test showed that over 57% of the sand flies fed on birds, mainly chickens and pigeons (Javadian et al., 1977). Furthermore Lu. longipalpis populations are maintained artificially high in Lapinha cave near Belo Horizonte, Brazil by the constant presence of live chickens (Lane, 1986). Henhouses seem to form a suitable man-made refuge for sand flies. A high proportion of residents in low-income neighbourhoods where VL is present keep chickens for a wide variety of reasons and this may be an important factor in the urbanisation of the disease, with both positive and negative effects of Le. infantum transmission (Alexander et al., 2002). Regarding positive effects, Lu. longipalpis readily feeds on chickens (Lainson, 1989) and large numbers can be collected from chicken houses (Genaro et al., 1990), suggesting that chickens may promote Le. infantum transmission by sustaining a large vector population near human dwellings. Various epidemiological studies have suggested that proximity to chickens is a risk factor for acquiring human (Corredor et al., 1989;Castellon and Domingos, 1991;Arias et al., 1996;Caldas et al., 2002) or canine VL (Moreira et al., 2003), although results have been mixed. Caldas et al. (2002) found that in some analyses chicken rearing was a VL risk factor but in others it appeared to exert a protective effect. This contradiction is unlikely to be resolved by epidemiological approaches alone.
Understanding the role of chicken rearing in the Le. infantum transmission cycle is important because domestic chickens are frequently kept by the urban poor, who are already at increased risk of VL. Therefore, the outcome of studies aiming to comprehend the vector/chicken host interactions could ultimately lead to a better applicability of control methods by health authorities in endemic areas under intense Leishmania transmission, perhaps suggesting removal of chicken coops from near human dwellings, banning chicken rearing from urban environments or even focal treatment of chicken coops (Alexander et al., 2002). Moreover, blood meal identification in field-caught sand flies can confirm a strong association between sand flies and chicken hosts in urban areas and also help to improve understanding the role of domestic chickens in transmission of Le. infantum in urban foci. Any attempt to determine the relative importance of blood meal sources may be biased by differences in the post-feeding accessibilities of insects that bit particular host species. The greater sensitivity of the method reported here means that information can be obtained from specimens that have ingested relatively small amounts of blood and may have moved some distance from the animals on which they fed.
Published as: Acta Trop. 2008 September ; 107(3): 230-237. Sponsored Document Sponsored Document Sponsored Document Published as: Acta Trop. 2008 September ; 107(3): 230-237. Sponsored Document Sponsored Document Sponsored Document Published as: Acta Trop. 2008 September ; 107(3): 230-237. Sponsored Document Sponsored Document Sponsored Document Published as: Acta Trop. 2008 September ; 107(3): 230-237. Published as: Acta Trop. 2008 September ; 107(3): 230-237. Sponsored Document Sponsored Document Sponsored Document Published as: Acta Trop. 2008 September ; 107(3): 230-237. Sponsored Document Sponsored Document Sponsored Document Published as: Acta Trop. 2008 September ; 107(3): 230-237. Sponsored Document Sponsored Document Sponsored Document Sponsored Document Sponsored Document Sponsored Document
Published as: Acta Trop. 2008 September ; 107(3): 230-237. Sponsored Document Sponsored Document Sponsored Document
Published as: Acta Trop. 2008 September ; 107(3): 230-237.Sponsored Document
We are grateful for the assistance of
Published as: Acta Trop. 2008 September ; 107(3): 230-237.
Background: Surgical castration in male piglets is painful and methods that reduce this pain are requested. This study evaluated the effect of local anaesthesia and analgesia on vocal, physiological and behavioural responses during and after castration. A second purpose was to evaluate if herdsmen can effectively administer anaesthesia.
Methods: Four male piglets in each of 141 litters in five herds were randomly assigned to one of four treatments: castration without local anaesthesia or analgesia (C, controls), analgesia (M, meloxicam), local anaesthesia (L, lidocaine), or both local anaesthesia and analgesia (LM). Lidocaine (L, LM) was injected at least three minutes before castration and meloxicam (M, LM) was injected after castration. During castration, vocalisation was measured and resistance movements judged. Behaviour observations were carried out on the castration day and the following day. The day after castration, castration wounds were ranked, ear and skin temperature was measured, and blood samples were collected for analysis of acute phase protein Serum Amyloid A concentration (SAA). Piglets were weighed on the castration day and at three weeks of age. Sickness treatments and mortality were recorded until three weeks of age. Results: Piglets castrated with lidocaine produced calls with lower intensity (p < 0.001) and less resistance movements (p < 0.001) during castration. Piglets that were given meloxicam displayed less pain-related behaviour (huddled up, spasms, rump-scratching, stiffness and prostrated) on both the castration day (p = 0.06, n.s.) and the following day (p = 0.02). Controls had less swollen wounds compared to piglets assigned to treatments M, L and LM (p < 0.001). The proportion of piglets with high SAA concentration (over threshold values 200, 400 mg/l) was higher (p = 0.005; p = 0.05) for C + L compared to M + LM. Ear temperature was higher (p < 0.01) for controls compared to L and LM. There were no significant treatment effects for skin temperature, weight gain, sickness treatments or mortality.
Conclusions: The study concludes that lidocaine reduced pain during castration and that meloxicam reduced pain after castration. The study also concludes that the herdsmen were able to administer local anaesthesia effectively.
Each year approximately 1.5 million male piglets are surgically castrated in Sweden. The number for all EU countries is approximately 100 million. The castration is mainly performed to eliminate boar taint in the meat, but also to prevent aggressive and sexual behaviour of male pigs. Castration is performed within the piglet's first week of life and is traditionally carried out without anaesthesia and analgesia. As surgical castration induces pain in piglets the procedure is considered an important animal welfare issue [1].
Pain is subjective and therefore difficult to quantify, and there are no specific parameters for measuring it [2]. However, it is widely accepted that piglets may react to pain in three ways: trough vocalisation, physiologically, and behaviourally [3]. Although piglets usually vocalise a lot when they are handled there is a clear difference in their vocalisation between being handled and castrated. Piglets that are castrated without anaesthesia produce a higher number of calls and with a higher frequency compared to piglets castrated with anaesthesia [2,4] or sham-castrated piglets (handled identically but without castration) [5][6][7]. The greatest amount of high-frequency calls are produced when the piglet's spermatic cords are pulled and severed, and is therefore identified as the most painful moment during castration [8].
The sympathetic nervous system is activated during different kinds of stress (pain, anger and fear) and several changes are noted in the body, for example: dilated pupils, increased heart rate and blood pressure, redirected blood from the skin, decreased digestion and dilated bronchioles. During activation adrenocorticotropic hormone is released and induces secretion of cortisol [9]. These changes can be used as possible indicators of pain and several of these changes have been shown in piglets during and after castration [2,10,11].
Stress, trauma, infection or inflammation also triggers the acute phase protein response, which is a part of the body's early defence. Serum amyloid A (SAA) is a major acute phase protein in pigs that can increase quickly and with large amplitude, and SAA level can therefore be used for defining the health and welfare status of pigs [12].
During castration, piglets without anaesthesia produce resistance movements with longer duration and higher intensity than piglets with anaesthesia [13]. After castration, behaviour alterations show that pain responses induced by castration persists over time; up to four to six days after castration according to some studies [5,14,15]. The pig production sector is searching for suitable methods that reduce pain induced by surgical castration, and alternatives to surgical castration. The method must be fast, cost effective, produce minimum stress and pain both during and after castration, and be safe for both the handler and the piglet. The method should also ensure a quick recovery to minimize the risk of the piglet being crushed by the sow. Currently there are essentially two alternatives that meet most of these requirements and which could be accepted in Swedish pig production. One method is immunocastration and the other involves the use of local anaesthesia and analgesia.
The objective of this study was to evaluate painrelated responses of male piglets castrated with/without local anaesthesia and with/without analgesia. A second purpose was to evaluate if herdsmen can effectively administer local anaesthesia by intratesticular injection. If the outcome of the study shows that herdsmen are able to effectively administer local anaesthesia this can lead to change of regulation, which will make it possible for herdsmen in Sweden to anaesthetise their piglets before castration. The use of anaesthetics in EU countries and Norway is currently restricted to veterinarians (Council Regulation No 2377/90). Herdsmen are allowed to administer analgesia after approved education according to Swedish regulation (SJVFS 2010:17, D9).
The study has been approved by the Ethical Committee for Animal Experiments, Uppsala, Sweden (reference number C 164/9). All piglets in the study would have been subjected to castration as a routine procedure, regardless of the study.
The study was conducted between October 2009 and February 2010 in five piglet-producing herds in the south-central part of Sweden. The herds were satellite herds within a sow pool with Landrace x Yorkshire sows, and the sires of the piglets were Hampshire boars. Batch-wise production was applied and in each batch about 45 sows farrowed in individual farrowing pens. The pens had a concrete floor, with a slatted floor in the dunging area and a nest area for the piglets with a heat lamp. Cross-fostering was applied and these piglets were not discriminated in the study. All piglets, except for those in herd 1, received an iron injection on the day of castration. Piglets in herd 1 were given oral iron pasta shortly after birth. No piglets were subjected to teeth clipping or tail docking, which is not allowed according to Swedish regulation.
Castration was performed on 1-7 days old piglets and the majority of the piglets were 3-4 days old when castrated. In all herds except for herd 2, the piglets were fixated in a restraining device and castration was carried out using a scalpel. The scalpel was used to make the initial incisions after which the testicles were severed by cutting the spermatic cords. In herd 2, the piglets were restrained between the herdsmen's legs or under their arm and castration was performed using an emasculator. The emasculator was used to make the initial incisions after which the testicles were removed by cutting the spermatic cords.
The study comprised 557 male piglets, randomly selected from five herds. In these herds, 30, 25, 30, 26 and 30 experimental litters were included.
Four male piglets in each of 141 litters were randomly assigned to one of four treatments: castration without local anaesthesia or analgesia (C, controls), castration with analgesia (M, meloxicam), castration with local anaesthesia (L, lidocaine), or castration with both local anaesthesia and analgesia (LM). All four treatments were represented in each litter. Six litters were made up of only three male piglets, and one litter of only two. Seven litters were therefore incomplete and the total number of male piglets was lower than optimal 564. Before the study started, the herdsmen received instructions from a veterinarian on how to inject local anaesthesia and analgesia.
Two technicians, who were not blind to the treatments due to practical reasons, performed all measurements. The measurements were split between the two technicians with each technician performing the same measurements in all herds.
For local anaesthesia, lidocaine 10 mg/ml with epinephrine 5 μg/ml (Xylocain ® , AstraZeneca, Södertälje, Sweden) was used (treatments L and LM). A total of 0.5 ml was injected in each testicle. While most of it was administered into the testicle, a small amount was injected subcutaneously into the scrotum when pulling the needle out. The action time of lidocaine is approximately one hour [16]. Castration was performed three minutes to 30 minutes after injection of lidocaine.
For analgesia, the nonsteroidal anti-inflammatory drug (NSAID) meloxicam 5 mg/ml (Metacam ® , Boehringer Ingelheim Vetmedica, Malmö, Sweden) was used (treatments M and LM). A dose of 0.2 ml was injected intramuscularly behind the piglet's ear immediately after the castration.
On the castration day, the four males in each litter subjected to the study were weighed and marked with spray colour on their back. Each treatment was represented by a different colour. Lidocaine was injected (by herdsmen) to the piglets subjected to treatments L and LM. The castration was performed in random order within each litter. During castration, piglet vocalisation was measured and resistance movements were judged. Vocal response was measured with a decibel meter (Mini Sound Level Meters) measuring dB(A) and the call with the highest intensity level during the castration was recorded. The decibel meter was held as close to the snout as possible without touching it. Resistance movements were judged on a visual analogue scale (VAS) [17] where a mark closer to the left end of the line corresponds to "low intense" movements and a mark closer to the right end corresponds to "high intense" movements. The piglets were ranked on a 1-4 scale (mostleast) within each litter according to the intensity and duration of their resistance movements. After castration, piglets subjected to treatments M and LM were given meloxicam (by herdsmen).
After castration, piglet behaviour was observed through instantaneous observations every ten minute, during 70 minutes, resulting in seven observations per piglet. Each technician studied ten litters per herd, and together a total of 398 piglets from 100 litters. A detailed ethogram with 23 variables (Table 1) with behaviours suggested by Wemelsfelder and van Putten [5], Hay et al. [14] and Llamas Moya et al. [15] was used. The behaviours were classified into five groups: body position, non-specific behaviour, pain-related behaviour, social cohesion and location. The pigs were studied from the front of the pen.
The following morning behaviour was observed according to the same protocol as on the castration day. Subsequently, castration wounds of the experimental piglets were ranked on the basis of how swollen the wounds were using a 1-4 scale (most-least) within the litter. Temperature was measured the following day using digital infrared thermometers. The thermometer for ear temperature measured with an accuracy of ± 0.1°C. Skin temperature was measured around the castration wounds, the thermometer had an accuracy of ± 0.2°C.
In each herd blood samples were collected from piglets in 15 litters, a total of 296 piglets. Plain vacutainer tubes were used for collecting ~2 ml blood per piglet. The samples were centrifuged at 2000 x g at 5°C for five minutes two to five hours after the collection. The plasma was stored into cryo tubes in -20°C until they were analysed for SAA with commercial solid phase sandwich immunoassay kit (Tridelta Ltd.) in accordance with the manufacturer's instructions. The detection limits were 15.6 -2000 mg/l.
Finally, the piglets were marked with differentcoloured ear tags depending on the treatment and the litter identity was written on the tag.
At three weeks of age, the piglets were weighed again and the ear tags were removed. Journals that the herdsmen had kept for registration of sickness treatments and mortality were collected.
Statistical analysis was performed using SAS Software, version 9.2 (SAS Institute Inc., Cary, NC, USA). Normal distribution was checked using proc univariate. Vocalisation, ear and skin temperature, and weight gain were analysed using analysis of variance (proc mixed). The statistical model applied (Model 1) included the fixed effects of herd (5 classes, herd 1-5), treatment (4 classes, C, M, L and LM), the interaction between herd and treatment, and the random effect of litter nested within herd.
Rank data of resistance movements and swelling of wounds was analysed using proc npar1way. Kruskal-Wallis test was applied and pair-wise comparison between treatments was performed using Wilcoxon tests.
The 23 behaviour variables, and five constructed group variables (Table 1), which for each piglet was the sum of all seven 0/1-observations, were analysed using proc glimmix (Model 1, poisson distribution). The group variables were the sum, for each piglet and day, of observations with at least one behaviour variable occurring within a group. The two observation occasions (the castration day and the following day) were analysed separately.
The effect of lidocaine and meloxicam was tested by considering them two separate treatments: lidocaine administered (L+LM), lidocaine not administered (C +M), meloxicam administered (M+LM) and meloxicam not administered (C+L). The statistical model (Model 2) included the fixed effects of a herd (5 classes, herd 1-5), administration of lidocaine (2 classes, 0/1), administration of meloxicam (2 classes, 0/1), the interaction between administration of lidocaine and meloxicam, and the random effect of litter nested within herd.
The data on SAA concentration was transformed into a number of 0/1-variable. Each value was assigned a 1 when the SAA concentration exceeded a certain threshold value (50, 100, 200, 400 and 600 mg/l). The analysis was performed using proc glimmix (Model 2, binomial distribution).
Sickness treatment and piglet mortality were analysed using X 2 -test. Correlations were calculated using spearman rank correlation. P-values ≤ 0.05 were regarded as significant.
The results are presented for all four treatments. However, treatments C and M were identical during castration, i.e. the piglets were castrated without lidocaine. Treatments L and LM were also identical as these piglets were castrated with lidocaine. The following day, treatment C and L were relatively comparable because the piglets were not given meloxicam, while the piglets in treatments M and LM had been treated with meloxicam. The interaction between herd and treatment did not have any significant impact on the variables but was included in the model to demonstrate this.
As shown in Figure 1, piglets castrated with lidocaine (L and LM) produced calls with a lower intensity level (p < 0.001) than piglets castrated without lidocaine (C and M). There were no significant differences between the two treatments with lidocaine (L and LM) or between the two treatments without lidocaine (C and M).
Figure 2 shows the difference in call intensity between the herds. Significant interaction between treatment and herd was not found for call intensity.
Piglets castrated with lidocaine (L and LM) showed less resistance movements (p < 0.001) than piglets castrated without lidocaine (C and M), (Figure 3). No significant difference in resistance movements was found between the two treatments with lidocaine (L and LM) and the two treatments without lidocaine (C and M).
The correlation between call intensity and resistance movements was r = -0,38 (p < 0.001).
Controls had less swollen castration wounds compared to the other three treatment groups (p < 0.001), (Figure 4). There was no significant difference between treatments M, L and LM.
Ear temperature was significantly higher (p < 0.01) for controls compared to piglets given lidocaine (L and LM), (Table 2). No significant differences were found for skin temperature. Within treatment, SD for temperature (both ear and skin) were higher for L (1.1°C) compared with the other three treatments (0.9°C).
For the SAA 0/1-variables, pair-wise comparisons between the four treatments did not show significant differences (Table 3). However, significant effects for the threshold value 200 and 400 mg/l (p = 0.005; p = 0.05) were found when treatments C+L (no meloxicam) were compared with treatments M+LM (meloxicam). For the threshold value 600 mg/l there was a tendency for a lower proportion of piglets given meloxicam (p = 0.06, n.s.).
Herd 1 had a lower proportion of piglets with high SAA concentrations. The percentage of piglets not given meloxicam that were above threshold value 200 mg/l was 13% in herd 1 compared to 37% in the other herds. For piglets given meloxicam, the percentage of piglets above 200 mg/l was 7% for herd 1 and 21% for the other herds. In herd 1, none of the piglets given meloxicam had SAA concentrations over 400 mg/l. In the other herds, at least one piglet given meloxicam had SAA concentrations over the thresholds 400 and 600 mg/l.
A total of 63 piglets (11%) were treated for health problems between the castration day and three weeks of age. The piglets were equally distributed over the treatments (C:17, M:11, L:17 and LM:18). During the same period, 26 piglets (5%) died but no significant effect of treatment on mortality was found (C:6, M:6, L:4 and LM:10).
The mean weight on the castration day was 2.2 kg (SD = 0.5 kg) for all treatments. There was no significant difference between treatments in weight gain (kg) between the castration day and three weeks of age.
No significant treatment effects in behaviour were found related to any of the 23 behaviour variables when these Mean value for call intensity (dB(A)) for the treatments C, M, L and LM, per herd. There were no significant interactions between treatment and herd. were analysed separately (using Model 1). When the variables were categorised into the five groups (Table 1), a significant difference was found between treatments C and LM for the group pain-related behaviour (huddled up, spasms, rump-scratching, stiffness and prostrated), on the day after castration (p = 0.04), (Table 4).
The effect of lidocaine and meloxicam was tested by considering them as two separate treatments: lidocaine administered (L+LM), lidocaine not administered (C +M), meloxicam administered (M+LM), meloxicam not administered (C+L), (using Model 2), (Table 5). The comparisons showed that piglets given meloxicam (M +LM) displayed less pain-related behaviour than piglets not given meloxicam (C+L) on both the castration day (p = 0.06, n.s.) and the following day (p = 0.02). No significant difference was found between treatments L+LM and C+M for pain-related behaviour. No significant differences were found in the other four groups of behaviour variables.
The method
The present study has, in line with several other studies, shown that castration without anaesthesia causes severe pain [2][3][4]13]. This pain persists for several days [5,14,15] and can cause delayed recovery, reduced feedand water intake, reduced immune capacity and impaired welfare [18]. Traumatic experiences of pain, such as that experienced during castration, can also lead to hypersensitivity [19 cited by 10] and may result in increased stress when piglets associate handling with acute pain [11].
The outcomes of the study show that the herdsmen in the study were able to administer local anaesthesia effectively into the testicles and scrotum so that an adequate anaesthesia was achieved. This method is a possible way forward for improved welfare for male piglets in Swedish pig production. The herdsmen must however be instructed by a veterinarian before being allowed to administer local anaesthesia. Precision of the injection and the waiting time after injection affects the efficiency of the anaesthetics [16] and have to be focused in the training. Table 5 Percentage of displayed behaviours the castration day (day 0) and the following (day 1) when lidocaine respectively meloxicam was administered or not Day L+LM C+M M+LM C+L Body position 0 44.7 48.3 44.8 45.2 1 38.9 41.9 39.3 40.7 Non-specific behaviour 0 75.3 71.1 75.0 74.6 1 74.9 74.7 77.6 73.9 Pain-related behaviour 0 4.7 6.1 4.6 (a) 6.5 (b) 1 3.6 5.8 4.3 a 6.0 b Social cohesion 0 3.7 3.5 2.7 2.4 1 1.6 2.2 2.1 2.3 Location-heat lamp 0 52.3 55.9 52.6 52.9 1 52.2 51.2 48.3 50.4 Piglets given meloxicam (M+LM) showed less pain-related behaviour than piglets not given meloxicam (C+L) on both the castration day (day 0, p = 0.06, n.s.) and the following day (day 1, p = 0.04). Means with different letters indicate significant difference (p < 0.05).
Possible methods that do not involve surgical castration are immunocastration and raising of entire males. However, raising of entire males does also affect the animal welfare negatively because of aggressive behaviour and mounting leading to e.g. increased leg problems [20]. To slaughter entire males before sexual maturity is not economically sustainable in Sweden. Castration under CO 2 -gas anaesthesia is not a suitable method according to Swedish animal welfare legislation. CO 2 -anaesthesia induces a high level of stress in the early induction phase, before surgical anaesthetic depth is reached [21]. In addition, use of anaesthetic gases is strictly regulated by the Swedish work environment act (ASS 2001:7) and is not suitable for field use [22].
Lidocaine was chosen for local anaesthesia as it has been used in several studies concerning castration and beneficial effects have been identified [2][3][4]16,23]. Lidocaine has a rapid onset and low toxicity [24]. The effect of the anaesthesia is prolonged by epinephrine and the risk for systemic reactions, e.g. fever, apathy and inappetence, decreases [24]. Ranheim et al. [16] showed that 40 min after injection the lidocaine concentration in the cords was severely decreased and this may have affected the result of the few piglets with long interval (up to 30 min) between lidocaine injection and castration.
In the present study all piglets received 0.2 ml meloxicam regardless of weight. This can have implications for low and heavy weight piglets and might have affected the result. In practice it will probably not be possible to give the exact dose of meloxicam given the weight. Boehringer Ingelheim recommends a dose of 0.2 ml for a piglet weighing 2.5 kg [25]. To be able to evaluate the effect of the drugs during the castration, meloxicam was given after the castration. In practice it would be possible to give meloxicam in connection with the injection of the local anaesthesia. There was a variation in time between the meloxicam administration and the start of the behaviour observations and this might have affected the results of the behaviour study on the castration day.
In agreement with other studies on piglet vocalisation during castration [3,4], the present study shows that piglets castrated with lidocaine produced calls with a lower intensity than piglets castrated without lidocaine. Marx et al. [4] have shown that calls produced by piglets castrated with lidocaine are similar to those produced by sham-castrated piglets.
Marx et al. [4] have suggested that a parameter that describes a single moment in the call, e.g. peak level, is more representative than parameters that describe a mean level. Therefore, in this study the call with the highest intensity during castration was recorded. It is assumed that this call was produced during the pulling and severing of the spermatic cords, which Taylor and Weary [8] have identified as the most painful moment during castration.
A difference in call intensity between the herds may be explained by the different castration techniques used by the herdsmen. The calls with the highest intensity were recorded in herd 2 where the herdsmen used an emasculator for castration. This can be interpreted that castration with an emasculator causes more pain, but it is more likely because the calls in herd 2 were not suppressed by the restraining device.
However, restraining method has been seen to not influence the pain responses during castration [6]. Taylor and Weary [8] state that it might be the pulling of the spermatic cords more than the severing that is painful. Traction upon the testes is likely to be felt along the spermatic cords and into the inguinal canal. If the testicles and spermatic cords are pulled a long distance before severing, this is likely to cause pain that may not be prevented by local anaesthesia in the testicles. Even after intratesticular injection of lidocaine the piglets still responded with some vocalisation and resistance movements during the castration procedure. Ranheim et al. [16] showed by means of autoradiograms that radiolabelled lidocaine injected into a piglet testicle was evident in both testicle and spermatic cord three minutes after injection. However, the concentration in the cremaster muscle was low ten minutes after injection. As the cremaster muscle is cut off during castration this can explain why piglets show some pain-related behaviour despite receiving lidocaine.
The study showed a correlation between dB-level and resistance movements. High dB-levels were associated with intensive resistance movements. In this study, as well in studies by Leidig et al. [13] and Horn et al. [23], local anaesthesia (procaine and lidocaine) reduced resistance movements during castration.
The study showed that controls had less swollen wounds compared to piglets treated with lidocaine or meloxicam. Small bleedings can occur as a result of the injection of local anaesthetics, which can contribute to an increased swelling (personal communication, Nyman, 2011). Why the swellings also were increased for piglets treated with meloxicam cannot be explained. The result is similar to Kluivers-Poodt et al. [3] where thickening of the scrotum was found on the fourth day after castration in several piglets treated with lidocaine and/or meloxicam.
Rectal temperature is the best indicator of adequate body temperature, but measuring temperature in the ear can also provide reliable estimates of body temperature and was used since it is a very fast method. Body temperature is nearly constant but fever occurs because of infections, or in some cases extensive tissue damage. Skin temperature is on the other hand more influenced by the environment and can therefore vary considerably. During activation of the sympathetic nervous system the blood is redirected from the skin to essential organs, which leads to a lowering of the skin temperature [9], and measurements can give information on the shock reaction [3].
In present study measurements of ear temperature showed that controls had higher temperature than piglets given lidocaine, why cannot be explained. Hypothetically, piglets given meloxicam should have lower ear temperature than piglets not given meloxicam because of the NSAIDs antipyretic effects [26]. No differences between treatments were found in terms of skin temperature and this is probably because the measurements were performed the day after castration, when the skin temperature had returned to normal.
The SAA concentration in the blood is normally very low (bordering on the unmeasurable) [12] but can increase hundredfold after stress, trauma, infection or inflammation as a consequence of increased levels of pro-inflammatory cytokines [27]. The concentration is the highest two to three days after a trauma and return to normal levels after seven to ten days [12,28]. The SAA concentration can act as an general marker of inflammation and has been seen to reflect the intensity of stress, trauma and inflammation [28][29][30]. The results from the present study show that piglets that were given meloxicam had lower SAA concentrations on the day after castration. The percentage of piglets with high SAA concentrations (> 200 mg/l) was halved when meloxicam was administered. The enzyme cyclooxygenase (COX) is a prerequisite for the creation of prostaglandins, the proteins that create pain-mediating substances [26]. NSAIDs act anti-inflammatory trough inhibition of COX [26] and the result of the present study can be seen as an indirect measure of the antiinflammatory effect of the meloxicam. The action time for lidocaine is limited to approximately one hour and is not likely to affect the postoperative inflammation [28].
SAA analysis showed a deviating pattern for piglets given meloxicam in herd 1: the proportion of piglets with high SAA concentrations was much lower compared to other herds. Piglets in herd 1 were given oral iron pasta after birth instead of an iron injection on the castration day. Injection of iron is an unnatural way for piglets to receive iron because there is no regulation system for exudation of iron trough the liver or kidneys. In nature, iron enters the body exclusively through the diet. The iron balance is regulated by the rate of absorption from the small intestine and the risk for extreme concentrations is therefore lower when iron is given orally compared with injection [31]. Addition of iron salts in connection with injection of NSAID might increase the irritating effect on gastrointestinal mucous [32] and that might have caused the higher SAA concentrations in the other herds.
The weight gain did not differ between the treatment groups, which is in accordance with other studies [3,5,14,33].
In the present study, piglets showed specific pain-related behaviour induced by castration which also has been seen in other studies [3,5,14,15,33]. The piglets given meloxicam in the present study showed less pain-related behaviours than piglets not given meloxicam. Similary, Keita et al. [33] have found an effect of meloxicam on pain relief two and four hours after castration. Kluivers-Poodt et al. [3] have also seen that piglets castrated with or without lidocaine showed more pain-related behaviours than sham-castrated piglets. However, less painrelated behaviours were displayed if the piglets with lidocaine were also given meloxicam.
Differences in piglets' non-specific behaviour between the treatments were not shown in this study. Other studies have shown that castrated piglets become more isolated after castration [14,15] and that the time spent by the udder (both more and less) differs between castrated and sham-castrated piglets [7,14,15,34]. However, in present study, no sham-castrated were included, but only castrated piglets. A lack of differences in non-specific behaviour can also be explained by the fact that all treatments were present in the same litter, which may have caused, as suggested by Kluivers-Poodt et al. [3], the piglets to influence each other's social behaviour.
This study concludes that lidocaine injected intratesticularly reduced pain responses during castration and that meloxicam reduced the pain-related behaviours after castration. It is therefore recommended that both local anaesthesia and analgesia should be given to piglets to reduce pain induced by castration. The study also concludes that the herdsmen in the study, after training, were able to inject local anaesthesia effectively. However, the method requires handling of the piglets on two separate occasions, which contribute to stress.
the Department of Clinical Sciences, SLU, for performing the SAA analysis. Thank you Stina Warnstam Drolet for correcting the English.
Body weight supported by belly Lateral lying, side Body weight supported by side 2. Non-specific behaviour Walking/running Moving walking, trotting or galloping By udder Activity by the udder: suckling, massaging udder or looking for a teat Nosing/chewing/ licking Nosing/chewing or licking material or the littermates/mother Playing Head shaking, springing (sudden jumping or leaping) or running. Can involve partners (gentle nudging or pushing, mounting, chasing, etc.) Desynchronised Activity different from that of most (at least 75%) littermates (e.g. sleeps while most other littermates suckle)
5. Location
Heat-lamp Sitting, standing, lying under the heat lamp
The proportion of piglets with SAA threshold value over 200, 400 and 600 mg/l was higher for piglets not given meloxicam (C+L) compared to piglets given meloxicam (M+LM) ( ** p = 0.005; * p = 0.05; (*) p = 0.06)
This study was financed by The
The author declares that they have no competing interests.
Authors' contributions MH participated in developing the design of the study, performing the field study, analysing data and drafting the manuscript. NL applied for funding of the study, planned the design of the study, helped analyse data and helped to draft the manuscript. GJ and GN participated with veterinarian expertise to the design of the study and education of the herdsmen. All authors have read and approved the final manuscript.
Objective To investigate whether acupuncture reduces the duration and intensity of crying in infants with colic. Patients and methods 90 otherwise healthy infants, 2-8 weeks old, with infantile colic were randomised in this controlled blind study. 81 completed a structured programme consisting of six visits during 3 weeks to an acupuncture clinic in Sweden. Parents blinded to the allocation of their children met a blinded nurse. The infant was subsequently given to another nurse in a separate room, who handled all infants similarly except that infants allocated to receive acupuncture were given minimal, standardised acupuncture for 2 s in LI4. Results There was a difference (p=0.034) favouring the acupuncture group in the time which passed from inclusion until the infant no longer met the criteria for colic. The duration of fussing was lower in the acupuncture group the fi rst (74 vs 129 min; p=0.029) and second week (71 vs 102 min; p=0.047) as well as the duration of colicky crying in the second intervention week (9 vs 13 min; p=0.046) was lower in the acupuncture group. The total duration of fussing, crying and colicky crying (TC) was lower in the acupuncture group during the fi rst (193 vs 225 min; p=0.025) and the second intervention week (164 vs 188 min; p=0.016). The relative difference from baseline throughout the intervention weeks showed differences between groups for fussing in the fi rst week (22 vs 6 min; p=0.028), for colicky crying in the second week (92 vs 73 min; p=0.041) and for TC in the second week (44 vs 29 min; p=0.024), demonstrating favour towards the acupuncture group. Conclusions Minimal acupuncture shortened the duration and reduced the intensity of crying in infants with colic. Further research using different acupuncture points, needle techniques and intervals between treatments is required.
Ten per cent of newborn children in the Western world experience colic. 1 2 The aetiology is unclear but gastrointestinal factors and allergy to cow's milk protein have been suggested as possible causes. 3 Another suggestion is that colic is a behavioural condition resulting from unfavourable parent-infant interaction. 3 In three meta-analyses current medical treatments are evaluated as either ineffi cient (simethicone) or as having serious side effects like seizures, asphyxia and death [3][4][5] (dicyclomine, presently withheld by the manufacturer). In spite of the good prognosis of infantile colic with full spontaneous recovery, 6 colic inhibits optimal family relations [7][8][9] and increases the risk of child abuse. [10][11][12] Acupuncture is widely used and discussed in infantile colic. However, few articles have been published on this subject. Two uncontrolled studies report positive outcomes after acupuncture in children with night crying. 13 14 One qualitative study 15 and one randomised controlled study 16 also indicate that acupuncture has an effect on infants' crying. The objective of this study was to further investigate whether minimal acupuncture reduces the duration and intensity of crying in infants with colic.
A prospective, randomised, controlled, blinded clinical trial was performed at a private acupuncture clinic in Sweden, from November 2005 to February 2007. For the past 15 years this acupuncture clinic has offered acupuncture treatment for adult patients with different symptoms and for infants with colic.
Infants with colic, 2-8 weeks old, whose parents sought help at either a child health centre, the regional hospital's paediatric clinic or at the acupuncture clinic where the trial was performed, were consecutively preselected by health professionals who were informed of the inclusion criteria: healthy infants, born after gestational week 36, not treated with dicyclomine and fulfi lling the modifi ed Wessel criteria for colic: 'crying/fussing for at least 3 h a day, occurring 3 days or more in the same week'. 1 Parents with eligible infants and who were willing to participate reported the extent and degree of their infant's crying and fussing in a diary for at least 3 days. Exclusion of cow's milk from the infant's diet was recommended during the registration period if it had not already been tried. If meeting the criteria, the infant was included in the study and started the structured programme the following Monday or Thursday. Written informed consent was obtained from the parents, and the study was approved by the local research ethics committee (Dnr 583/2005). All infants continued the regular programme at their ordinary child health centre throughout the duration of the study.
A registered nurse skilled in acupuncture, nurse A, was hired specifi cally to perform the randomisation, administer the intervention and be the sole person aware of allocation and with access to the records during the study. Nurse A met the infants alone in the treatment room and was only informed of their age and study number. The randomisation procedure divided the infants into an intervention group with a structured programme including acupuncture (acupuncture group) or to the same structured programme not including acupuncture (control group). As we proposed that age was a prognostic variable that might interfere with the result, restricted randomisation was used to achieve Acupuncture reduces crying in infants with infantile colic: a randomised, controlled, blind clinical study a balance between 2-5 weeks old and 6-8 weeks old infants, respectively, in the groups. Two sets of sealed opaque envelopes, marked '2-5 weeks old' and '6-8 weeks old', respectively, had been prepared by nurse A before the study started. The envelopes contained a card with either 'control group' or 'intervention group', each in equal amounts. The card in the upper envelope in the pile appropriate to the infant's age determined the group to which each infant was assigned. Consequently, all infants had an equal probability of assignment to either group. Each infant remained in the initially allocated group throughout the study.
The study was double blind as neither the parents who registered the infants crying nor the nurse who met the parents (nurse B, the fi rst author) knew to which group the infant belonged. Nurse B enrolled parents of potential patients, informed them of the trial, assessed the infant's eligibility, obtained informed consent and met the parents at the acupuncture clinic. Two closed doors separated the parents from the treatment room and music was always played. Parents were informed that the needle was very thin, usually caused no bleeding or visible marks and that acupuncture does not necessarily provoke crying.
The structured programme consisted of a total of six biweekly visits to the acupuncture clinic. The fi rst visit lasted for 30 min, during which the parents met nurse B who repeated information on the study and collected baseline demographic data. During the following fi ve visits, parents met nurse B for 15 min appointments, and were asked standardised questions such as 'How is it going?', received standardised oral support such as 'Hopefully it will be better soon' and were given time for questions.
At each visit, the infant was carried to the treatment room by nurse B and left there with nurse A. The initial handling of the infants in the treatment room was identical. Nurse A held each infant's hand and spoke soothingly. If starting to cry, the infant was comforted by the nurse in her arms. The infants allocated to have acupuncture subsequently received minimal, standardised acupuncture with a sterilised, disposable acupuncture needle, Vinco MicroClean, 0.20 × 13 mm. The needle was inserted unilaterally and left in place for 2 s at an approximate depth of 2 mm at point LI4 of the hand's fi rst dorsal interossal muscle, a point often used in clinical practice when treating infants with colic and, also used in an earlier randomised controlled trial (RCT) studying acupuncture treatment for colic and known for the generalised analgetic effect. 16 Left and right hands were used alternately. After a maximum of 5 min in the treatment room, nurse A carried infants back to their parents. Infants allocated to the two groups went through exactly the same procedure except for the insertion of an acupuncture needle in the acupuncture group.
Defi nitions of 'fussing' (showing dissatisfaction and whimpering despite being carried), 'crying' (screaming loudly) and 'colicky crying' (crying hysterically and unconsolably) were communicated to the parents both verbally and in writing. Parents reported infants' fussing, crying and colicky crying in a standardised diary form originally developed and validated by Barr et al 17 and modifi ed and tested by Canivet et al. 18 The diary form consisted of sheets, each covering 24 h. Parents fi lled in boxes, each representing 5 min, to indicate when their infant was fussing (marked as F), crying (marked as C) and colicky crying (marked as CC). All marked boxes were counted manually and transferred into a database. Reports were made on at least 3 days during the baseline week preceding possible inclusion and daily during the three intervention weeks, directly following the baseline week. Twice weekly, parents completed a questionnaire modifi ed from Reinthal et al, 16 in which they described any adverse effects they considered to be caused by treatment. Duration of crying in the treatment room and bleeding were noted by nurse A. The primary end point was the number of infants who fulfi lled the colic criteria during each of the intervention weeks. The secondary end point was the total duration of fussing, crying and colicky crying (TC) during the three intervention weeks as reported by parents in the diary.
Based on the assumption that 50% of the infants would go into spontaneous remission without treatment and 75% with acupuncture, 40 patients per group were needed in order to have a 90% chance of detecting a signifi cant difference in remission at a two-sided 5% level. The statistical software SPSS version 17 (SPSS, Chicago, Illinois, USA) was used for calculations. As two parameters were not normally distributed all data were analysed with non-parametric statistics. Kaplan-Meier analysis was performed to assess the time for each infant's crying to fall below 180 min, indicating that the infant no longer fullfi lled the criteria for colic. To evaluate differences between intervention and control groups the log rank test was performed. Mann-Whitney U test was used to analyse crying and fussing times, and the relative difference in crying and fussing between the baseline and the intervention weeks was measured as a percentage. p Values <0.05 were considered statistically signifi cant.
Of the 210 infants who between November 2005 and February 2007 were suspected to have colic, 90 fulfi lled the colic criteria after completing the diary. Three infants randomised to the control group did not meet the criteria and were excluded, and the procedures for analysing the diaries before randomisation were changed (fi gure 1). Two infants in the acupuncture group who only came to the clinic fi ve times as the symptoms disappeared are counted as fulfi llers as their parents continued to complete the diary. Background data were analysed for infants starting the structured programme (n=86) and for infants who completed the three intervention weeks (n=81) (table 1). Outcomes from the intervention weeks are based on the infants' remaining in the study each week and drop outs are reported as missing values. Infants were stratifi ed by age and age at inclusion was similar in both groups (table 1). However, owing to small numbers in the subgroups, age groups were analysed together.
There were no signifi cant differences between the groups for background characteristics such as parents being born in Sweden, educational level, smoking and mother's complications during pregnancy or delivery; nor were there differences between their baseline levels of fussing and crying (tables 1 and 2).
There was a difference (p=0.034) between groups in the time which passed from inclusion until the infant had a mean value
Table 1 Baseline data for infants Background characteristics Infants starting the intervention (N=86) Infants completing 3 weeks (N=81) Acupuncture group (n= 46) Control group (n=40) Acupuncture group (n= 43) Control group (n=38) Firstborn, n (%) 22 (48) 22 (55) 21 (49) 21 (55) Gender, female, n (%) 22 (48) 19 (48) 21 (49) 19 (50) Gestational age, weeks, mean (SD) 39.2 (1.5) 39.5 (1.3) 39.3 (1.4) 39.5 (1.3) Age when colic started, weeks, mean (SD) 1.9 (1.3) 1.5 (1.0) 2 (1.3) 1.5 (1) Age at inclusion, weeks, mean (SD) 5.0 (1.9) 5.3 (1.7) 5.1 (1.9) 5.2 (1.6) Solely breastfed, n (%) 35 (76) 26 (65) 32 (74) 25 (66) Having a parent and/or sibling with food intolerance/allergy, n (%) 17 (37) 18 (45) 15 (35) 17 (45) Having a parent and/or sibling who had had infantile colic, n (%) 29 (63) 23 (58) 25 (58) 20 (53) for TC of <180 min/day for the fi rst time, indicating that the infant no longer met the criteria for colic. Figure 2 demonstrates this difference by showing the proportion of infants with a mean TC <180 min/day for each of the six treatment periods consisting of 3 or 4 days depending on whether treatment was given on a Monday or a Thursday. Median time until criteria for colic were no longer fulfi lled was 7 days in both groups.
The duration of fussing was shorter in the acupuncture group during the fi rst (p=0.029) and second (p=0.047) intervention weeks. The duration of colicky crying was shorter (p=0.046) in
Proportion of infants with a mean total duration of fussing, crying and colicky crying (TC) under 180 min/day for each of the six treatments.
Table 3 Fussing, crying, colicky crying and total duration of fussing, crying and colicky crying (TC) during the three intervention weeks for the infants still remaining in the trial at each of the intervention weeks
Categories of fussing and crying, min/day First intervention week Second intervention week Third intervention week Acupuncture group (n=46) Control group (n=40) p Value Acupuncture group (n=44) Control group (n=39) p Value Acupuncture group (n=43) Control group (n=38) p Value Fussing, median (q1-q3) 74 (53-154) 129 (80-183) 0.029 71 (41-123) 102 (60-148) 0.047 69 (36-109) 85 (63-151) 0.119 Crying, median (q1-q3) 76 (45-103) 61 (30-102) 0.428 52 (27-88) 55 (24-73) 0.964 54 (21-87) 46 (22-98) 0.846 Colicky crying, median (q1-q3) 20 (6-53) 26 (9-48) 0.634 9 (0-27) 13 (4-49) 0.046 3 (0-18) 9 (0-18) 0.087 TC/day, median (q1-q3) 193 (143-253) 225 (178-316) 0.025 164 (103-201) 188 (149-273) 0.016 149 (92-193) 169 (119-267) 0.062 Difference baseline -fi rst intervention week Difference baseline -second intervention week Difference baseline -third intervention week Acupuncture group (n=46) Control group (n=40) p Value Acupuncture group (n=44) Control group (n=39) p Value Acupuncture group (n=43) Control group (n=38) p Value Fussing, difference in % (min-max in %)
the acupuncture group during the second intervention week. However, TC was lower in the acupuncture group than in the control group as early as the fi rst intervention week (p=0.025) and in the following intervention week (p=0.016) (table 3). A subanalysis showed TC to already be lower (p=0.005) in the acupuncture group after the fi rst treatment. The relative difference between groups, measured as the percentage decrease of crying and fussing from baseline to intervention weeks 1, 2 and 3 showed differences between groups for fussing the fi rst week (p=0.028), for colicky crying the second week (p=0.041) and for TC the second week (p=0.024) (table 4).
Slight bleeding (one drop) was detected after needling in one of the 256 acupuncture treatments administered. Thirty-two infants (74%) in the acupuncture group cried for more than 10 s during one to four interventions in the treatment room compared with 14 infants (37%) in the control group (p = 0.009) (table 5). Crying lasted more than a minute in 37 out of 256 needling occasions (14%). No infant cried for more than 2 min. No other adverse events were reported.
In this study where both acupuncture and control groups were allotted six visits with support and counselling as an intervention beside their ordinary child health centre visits, there was an expected decrease in TC in both groups. 19 However, the decrease was slightly faster in the acupuncture group as shown by measuring both absolute and relative differences between groups. There was a small but signifi cant difference between groups already after the fi rst treatment and in the duration until the infants no longer fullfi lled the colic criterion. Spontaneous healing might explain the lack of difference between groups during the third intervention week. The results of this study are in agreement with the only RCT on acupuncture in infantile colic published, 16 in which 40 infants were included, of whom 20 were needled in LI4 bilaterally for 20 s. Spontaneous remission was more likely to occur in that study as some of the infants were older than 8 weeks. Furthermore, the parents were blinded but not the nurse meeting the parents and administering the acupuncture.
The strengths of our study are the randomisation, the blinding of the parents, the small number of drop-outs and strict protocol, including an extensive diary validated in several studies. 18 20 21 Furthermore, the infants were included before their eighth week in order to minimise the risk of spontaneous healing during the study period, and infants recovering after a 5-day period excluding cow's milk were not included. Blinding patient and practitioner and fi nding an inert control are considerable methodological problems in acupuncture research. [22][23][24][25][26][27] As parents could easily be infl uenced by the acupuncturist's enthusiasm, an advantage of this study was that the nurse they met was blinded to the infants' allocation. The structured programme, ensuring equal support and advice to all participating families infl uenced both groups equally, is a strength. Infants in both groups lacked expectations and had limited communication skills, thereby eliminating any difference in placebo effect in them and in their blinded parents.
No test of blinding was done after the three intervention weeks, which is a limitation. More infants in the acupuncture group than in the control group started to cry in the treatment room. Parents might have heard the infants cry and thus suspected that the infant had received acupuncture. However the fussing/crying lasted for <10 s in most cases. On one occasion one infant cried for more than a minute after the acupuncture treatment but none cried for more than 2 min, indicating that this light acupuncture treatment was well tolerated by the infants.
The safety of acupuncture is a major concern, particularly during early infancy when responses are diffi cult to evaluate. In a review, acupuncture was considered a safe modality for paediatric patients, but the authors advised that fewer needles should be used when treating children. 28 In accordance with this our study used one single point with light stimulation. As different acupuncture points result in different effects 29-32 55 the option of choosing points individually after analysing all symptoms presented in an ordinary clinical setting may increase effi cacy of future acupuncture treatment of colic. The six treatments in this study may be more than needed.
Most basic acupuncture research is conducted with electroacupuncture on animals, and cannot be generalised to manual acupuncture in infants. However, it is known that acupuncture in animals inhibits somatic [33][34][35][36] and visceral pain 37 38 and has an effect on the autonomous system. 25-27 39-44 Stimulating LI4 bilaterally resulted in more immediate effect than unilateral stimulation. 39 The motility in the intestinal tract and the gastric acid secretion increased or decreased depending on which points were needled. 29 30 32 45 46 In human adults [47][48][49] and children 50 acupuncture had a benefi cial effect on visceral symptoms like nausea. Acupuncture increased bowel movement in children, 51 altered gastric motility 52 and affected gastric emptying in adults with motility disorders 53 but caused no effect on gastric motility in healthy individuals. 54 Manual acupuncture applied to LI4 induced an increase in the sympathetic and parasympathetic nervous systems in 12 healthy individuals. 55 It is possible that infantile colic derives from distension of the intestines and activation of the autonomic nervous system and that acupuncture can infl uence both visceral pain and the autonomic nervous system. Thus it is plausible that even modest stimulation of LI4, as performed in this study, can infl uence either or both mechanisms and thereby alleviate infantile colic.
This study includes infants with eczema, a rash from a Von Rosen splint, a temperature, a hand burned by boiling water and infants whose mothers had a high level of anxiety or depression. In this aspect the participants represent clinical reality, and these affl ictions were equally distributed among the groups. Parents who were negative about exposing their children to acupuncture or who lacked the ability to complete the diaries did not participate and infants born prematurely were excluded. This leaves the included sample and the results of this study as reasonably representative of the general population.
Parents have described colic as a strain on the family. [7][8][9] As no safe and effective cure is known we assume that even a short reduction of the colicky period can make a difference. Of the 210 infants estimated by the parents to have colic, only 90 fulfi lled the criteria after registration of their symptoms in the diary. This indicates that parents have a tendency to overestimate the crying, and a diary in which parents note their infant's crying could be a valuable diagnostic tool. Another explanation may be that the defi nition of colic does not refl ect the parent's experience of what they consider to be colic.
Standardised, light stimulation of the acupuncture point LI4 twice a week for 3 weeks reduced the duration and intensity of crying more quickly in the acupuncture group than in the control group. No serious side effects were reported. Future research is needed to validate the results and to investigate the effi cacy of other acupuncture points and modes of stimulation for the treatment of infantile colic.
Acupunct Med 2010;28:174-179. doi:10.1136/aim.2010.002394
Thanks to
The authors thank
This study was conducted with the approval of the Lund University, Research Ethics Committee (Dnr 583/2005).
Provenance and peer review Not commissioned; externally peer reviewed.
▶ Previous reports suggested acupuncture might reduce infantile colic. ▶ We conducted a randomised controlled trial in 90 infants. ▶ Acupuncture showed a small but signifi cant effect on some outcomes.
ABSTRACTa db_231 82..91
Recently, we demonstrated that the central ghrelin signalling system, involving the ghrelin receptor (GHS-R1A), is important for alcohol reinforcement. Ghrelin targets a key mesolimbic circuit involved in natural as well as druginduced reinforcement, that includes a dopamine projection from the ventral tegmental area (VTA) to the nucleus accumbens. The aim of the present study was to determine whether it is possible to suppress ghrelin's effects on this mesolimbic dopaminergic pathway can be suppressed, by interrupting afferent inputs to the VTA dopaminergic cells, as shown previously for cholinergic afferents. Thus, the effects of pharmacological suppression of glutamatergic, orexin A and opioid neurotransmitter systems on ghrelin-induced activation of the mesolimbic dopamine system were investigated. We found in the present study that ghrelin-induced locomotor stimulation was attenuated by VTA administration of the N-methyl-D-aspartic acid receptor antagonist (AP5) but not by VTA administration of an orexin A receptor antagonist (SB334867) or by peripheral administration of an opioid receptor antagonist (naltrexone). Intra-VTA administration of AP5 also suppressed the ghrelin-induced dopamine release in the nucleus accumbens. Finally the effects of peripheral ghrelin on locomotor stimulation and accumbal dopamine release were blocked by intra-VTA administration of a GHS-R1A antagonist (BIM28163), indicating that GHS-R1A signalling within the VTA is required for the ghrelin-induced activation of the mesolimbic dopamine system. Given the clinical knowledge that hyperghrelinemia is associated with addictive behaviours (such as compulsive overeating and alcohol use disorder) our finding highlights a potential therapeutic strategy involving glutamatergic control of ghrelin action at the level of the mesolimbic dopamine system.
Ghrelin, a gastric hormone, exerts orexigenic and proobesity effects by interacting with key brain circuits involved in appetite and energy balance (Tschöp, Smiley & Heiman 2000;Wren et al. 2000;Cummings et al. 2001;Nakazato et al. 2001). The cloned receptor for ghrelin (GHS-R1A) is present, however, not only in discrete hypothalamic cell groups regulating energy balance but also in a number of other central nervous system (CNS) sites, such as the hippocampus, brainstem and reinforcement areas (Guan et al. 1997) implying a role for ghrelin in brain reinforcement. Thus, ghrelin activates a key mesolimbic circuit involved in natural as well as drug-induced reinforcement, the cholinergicdopaminergic reward link (Jerlhag et al. 2006(Jerlhag et al. , 2007(Jerlhag et al. , 2008)). This link encompasses the well-described dopamine projection from the ventral tegmental area (VTA) to the nucleus accumbens (N.Acc.) that forms part of the mesolimbic dopamine system, together with a cholinergic projection from the laterodorsal tegmental area (LDTg) to the VTA. By activating this reinforcement link, ghrelin may increase the incentive value of motivated behaviours such as food and drug seeking (Abizaid et al. 2006;Jerlhag et al. 2006;Jerlhag et al. 2009). Indeed, alcohol reinforcement was absent in pharmacological PRECLINICAL STUDY Addiction Biology doi:10.1111/j.1369-1600.2010.00231.x and genetic models of suppressed ghrelin signalling (Jerlhag et al. 2009).
Peripherally injected ghrelin also stimulates the mesolimic dopamine system (Jerlhag 2008;Quarta et al. 2009), hypothesizing that peripherally produced ghrelin reaches deeper brain structures. Moreover, an important premise for the study is that the afferents regulate ghrelin's effects on the mesolimbic dopamine system involving GHS-R1A in VTA. In the present experiments therefore we sought to determine whether VTA administration of a GHS-R1A antagonist suppresses the locomotor stimulatory and accumbal dopamine releasing effects of peripheral ghrelin. The activity of VTA dopamine neurons is modulated by various afferents to the VTA including glutamate, opioids and orexin (Kalivas, Churchill & Klitenick 1993;Wise 2002). Previously, a role for NMDA, opioid as well as orexin A receptors in different aspects of ghrelin-induced activation of the mesolimbic dopamine systems has been suggested (Toshinai et al. 2003;Naleid et al. 2005;Abizaid et al. 2006). Inspired by the possibility to suppress natural and chemical drug reinforcement by agents that interrupt central ghrelin signalling, we also sought to determine whether ghrelin's ability to activate the mesolimbic dopamine system, as measured by locomotor stimulation and accumbal dopamine release, can be interrupted by pharmacological suppression of glutamatergic, opioid and orexin A systems.
Adult post-pubertal age-matched male NMRI mice (8-12 weeks old and 25-30 g body weight; B&K Universal AB, Sollentuna, Sweden) were used for studies of locomotor activity and dopamine release as such studies are welldocumented in this strain (Jerlhag et al. 2006;Jerlhag et al. 2007;Jerlhag et al. 2008). Upon arrival the mice were allowed to habituate in groups of eight in standard cages (Macrolon III: 400 ¥ 250 ¥ 150 mm), for at least one week before initiation of the experiment. All mice were maintained at 20°C with 50% humidity and a 12/12 hour light/dark cycle (lights on at 7 am). Tap water and food (Normal chow; Harlan Teklad, Norfolk, England) were supplied ad libitum, except during the experiments. All experiments were conducted during the day time when the mice are less active. Studies were approved by the Ethics Committee for Animal Experiments in Gothenburg, Sweden.
Acylated rat ghrelin (Bionuclear; Bromma, Sweden) was diluted in 0.9% sodium chloride (saline vehicle) and was administrated intraperitoneally (i.p.) (10 ml/kg body weight). The selected dose, 0.33 mg/kg, was determined previously, as it increases locomotor activity and accumbal dopamine release as well as induces a conditioned place preference in mice (Jerlhag 2008). Ghrelin was administered 10 minutes prior to the initiation of the experiment (locomotor activity or microdialysis).
The dose of BIM28163 (Ipsen Biomeasure Inc, Milford, MA, USA), a GHS-R1A antagonist, has also been determined previously (Halem et al. 2004;Jerlhag et al. 2009). BIM28163 was diluted in Ringer solution (NaCl 140 mM; CaCl2 1.2 mM; KCl 3.0 mM and MgCl2 1.0 mM) (Merck KgaA, Darmstadt, Germany) and was administered at a dose of 2.5 mg/side (uni-or bilaterally into the VTA) at 40 minutes prior to i.p. ghrelin/vehicle exposure. Previous studies have established that this compound is a GHS-R1A antagonist and fully inhibits ghrelin-induced GHS-R1A activation (Halem et al. 2004).
The selected dose of AP5 (Sigma-Aldrich, Stockholm, Sweden), an N-methyl-D-aspartic acid (NMDA) receptor antagonist, was determined in a dose-response study where 0.5 mg/side (uni-or bilaterally into the VTA) was the highest dose not to affect locomotor activity per se (Fig. 1). AP5 or Ringer vehicle were administered 10 minutes prior to i.p. ghrelin/vehicle administration. AP5 does not affect nicotinic acetylcholine receptors in the CNS (Davies & Watkins 1982).
The selected dose of SB334867 (Tocris, Bristol, United Kingdom), an orexin A receptor antagonist, was determined in a dose-response study where 5 mg/side (bilaterally into the VTA) was the highest dose not to affect locomotor activity per se (data not shown). Doses in a similar range have previously been shown to block the cue-induced reinstatement of cocaine seeking (Smith, See & Aston-Jones 2009). SB334867 or vehicle (10%-DMSO in Ringer vehicle; Merck KgaA) were administered 10 minutes prior to i.p. ghrelin/vehicle exposure.
Naltrexone, an unselective opioid receptor antagonist with some selectivity to the m receptor, was diluted in saline vehicle. Naltrexone (1 mg/kg, i.p.) or saline vehicle were injected 30 minutes prior to i.p. ghrelin/vehicle. The dose was determined from previous studies in which doses in a similar range have been shown to block the reinforcing properties of alcohol in rodents (Herz 1997). The rationale for administering by the i.p. route is that direct mesolimbic effects of nalrexone to interrupt ghrelin-induced reinforcement are unlikely, based on previous studies in which this antagonist had no effect on ghrelin-induced food intake when administered into discrete mesolimbic sites (Naleid et al. 2005).
Intra-VTA injections were made using a volume of 0.5 ml/side via chronically implanted catheters, over 60 seconds. A 5 ml syringe (Kloehn, microsyringe; Skandinaviska Genetec AB, V. Frölunda, Sweden) was used to facilitate drug administration into the VTA. After each injection, the cannula was left in place for a further 60 seconds to facilitate diffusion. Potentially, the volume administered during intra-VTA injections could raise concerns about specificity due to possible leakage into neighboring structures. Previously, however, we found that only 'on-target' placements (in the VTA and LDTg but not in closely adjacent sites) resulted in significant effects of ghrelin or GHS-R1A antagonists on locomotor stimulation and dopamine release using a larger volume (1 ml) for mice (Jerlhag et al. 2007(Jerlhag et al. , 2009)). Supportively, in the present study neither AP5 nor BIM28163 blocked ghrelin-induced locomotor stimulation nor accumbal dopamine release in a few mice in which the canulae were misplaced in neighbouring structures (data not shown). All drug challenges were part of a balanced design with regard to both the treatment order and the number of subjects per treatment.
Peripheral administration of ghrelin has previously been shown to stimulate locomotor activity in mice (Jerlhag 2008). Locomotor stimulation is, at least in part, mediated by an increase in the extracellular concentration of accumbal dopamine (Engel et al. 1988). Locomotor stimulation has been suggested to be an homologous effect evolving from a common mechanism involving the mesolimbic dopamine system, implying that locomotor activity reflects reinforcement induced by drugs of abuse (Imperato & Di Chiara 1986;Wise & Bozarth 1987;Engel et al. 1988). Thus, accumbal dopamine measurement experiments were conducted only after first establishing that aforementioned antagonist compounds suppress ghrelin-induced locomotor stimulation. All mice, except those treated with naltrexone, were implanted with bilat-eral guide cannulaes aiming at the VTA. The mice were anesthetized with isofluran (Isofluran Baxter; Univentor 400 Anaesthesia Unit, Univentor Ldt., Zejtun, Malta), placed in a stereotaxic frame (David Kopf Instruments; Tujunga, CA, USA) and kept on a heating pad to prevent hypothermia. The scull bone was exposed and two holes for the guide cannulas (stainless steel, length 10 mm, with an o.d./i.d. of 0.6/0.45 mm) and one for the anchoring screw were drilled. The VTA coordinates were 3.4 posterior to bregma, Ϯ0.5 mm lateral to the midline and 1.0 mm below the brain surface (Franklin & Paxinos 1996). All guide cannulae were surgically implanted four days prior to the experiment. After surgery the mice were kept in individual cages (Macrolon III). At the time of the experiment, a cannula for drug administration was inserted and extended another 3.8 mm ventrally beyond the tip of the guide cannula, aiming at the VTA.
Locomotor activity was registered in eight sound attenuated, ventilated and dim lit locomotor boxes (420 ¥ 420 ¥ 200 mm, Kungsbacka mät-och reglerteknik AB, Fjärås, Sweden). Five by five rows of photocell beams, at the floor level of the box, creating photocell detection allowed a computer-based system to register the activity of the mice. Locomotor activity was defined as the accumulated number of new photocell beams interrupted during a 60-minute period.
Before initiating the experiments, a dummy cannula was carefully inserted and retracted into the guide cannula to remove clotted blood and hamper the spreading depression. Mice were then allowed to habituate to the locomotor activity box one hour prior to drug challenge. In separte experiments, the effects of i.p. administered ghrelin on locomotor stimulation was investigated following intra-VTA administration of BIM28163, AP5 or SB334867 to mice. In subsequent experiments, naltr- exone was injected i.p. prior to ghrelin. In the first locomotor activity experiment, BIM28163 (2.5 mg/side) or an equal volume (0.5 ml/side) of vehicle solution (Ringer) was administered locally and bilaterally into the VTA. Ghrelin (0.33 mg/kg) or an equal volume of vehicle solution (saline vehicle) was thereafter injected. The same experimental protocol was used for AP5 (0.5 mg/side), SB334867 (5 mg/side) and naltrexone (1 mg/kg, i.p.). All mice received drug treatment only twice (antagonist/ vehicle and ghrelin/vehicle). Neither water nor food was available to the mice during the locomotor experiments. The activity registration started five minutes after the last injection and was subsequently measured for a 60-minute period. For intra-VTA administration only mice with guide cannulae placements in the VTA were included in the statistical analysis.
For measurements of extracellular dopamine levels and overflow (that reflect dopamine release), mice were implanted unilaterally with a microdialysis probe positioned in the N.Acc. shell and a guide cannulae into the VTA. The probe and the guide cannula/e were positioned ipsilateral, and the location was randomly alternated to either the left or right side. The surgery was preformed two days prior to the experimental day as described above (see Locomotor activity experiments). The coordinates for the N.Acc. shell were: 1.5 mm anterior to the bregma, Ϯ 0.7 lateral to the midline and 4.7 mm below the surface of the brain and the coordinates for the VTA were 3.4 posterior to bregma, Ϯ0.5 mm lateral to the midline and 1.0 mm below the brain surface (Franklin & Paxinos 1996). At the time of the experiment a cannula for drug administration was inserted and extended another 3.8 mm ventrally beyond the tip of the guide cannula, aiming at the VTA. The exposed tip of the dialysis membrane (20 000 kDa cut off with an o.d./i.d. of 310/220 mm, HOSPAL, Gambro, Lund, Sweden) of the probe was 1 mm.
In separate experiments, the effects of intra-VTA administration of AP5, or in separate experiments BIM28163, on ghrelin-induced accumbal dopamine release were determined, involving microdialysis in freely moving mice. On the day of the experiment, a dummy cannula was carefully inserted and retracted into the guide cannula. The probe was thereafter connected to a microperfusion pump (U-864 Syringe Pump; AgnThós AB) and perfused with Ringer solution at a rate of 1.5 ml/minute. After one hour of habituation to the microdialysis set-up, perfusion samples were collected every 20 minutes. The baseline dopamine level was defined as the average of three consecutive samples before the first drug/vehicle challenge. After the baseline samples, the antagonist (AP5 or BIM28163) was administered locally into the VTA followed subsequently by a ghrelin (i.p.) injection. The dopamine levels in the dialysates were determined by HPLC with electrochemical detection. A pump (Gyncotec P580A; Kovalent AB; V. Frölunda, Sweden), an ion exchange column (2.0 ¥ 100 mm, Prodigy 3 mm SA; Skandinaviska GeneTec AB; Kungsbacka, Sweden) and a detector (Antec Decade; Antec Leyden; Zoeterwoude, the Netherlands) equipped with a VT-03 flow cell (Antec Leyden) were used. The mobile phase (pH 5.6), consisting of sulfonic acid 10 mM, citric acid 200 mM, sodium citrate 200 mM, 10% EDTA, 30% MeOH, was vacuum filtered using a 0.2 mm membrane filter (GH Polypro; PALL Gelman Laboratory; Lund, Sweden). The mobile phase was delivered at a flow rate of 0.2 ml/minute passing a degasser (Kovalent AB), and the analyte was oxidized at +0.4 V.
After completion of the microdialysis experiments, the locations of the probe and guide cannulae were verified. Neither water nor food were available to the mice during the microdialysis experiment. Only mice with probe placement in the N.Acc. and guide cannulae in the VTA were included in the statistical analysis.
After the locomotor activity and microdialysis experiments were completed, the location of the probe and/or cannula/e were verified. The mice were decapitated, probes were perfused with pontamine sky blue 6BX to facilitate probe localization, and the brains were mounted on a vibroslice device (752M Vibroslice; Campden Instruments Ltd, Loughborough, UK). The brains were cut in 50 mm sections and the location of the probe and/or cannula/e was determined by gross observation using light microscopy. The exact position (some correct and some misplaced) of the probe and/or guide cannula/e was verified (Franklin & Paxinos 1996).
All locomotor activity data were evaluated by a two-way ANOVA followed by Tukey's HSD post-hoc tests comparing treatments. The microdialysis experiments were evaluated by a two-way ANOVA for repeated measures followed by Tukey's HSD post-hoc test for comparisons between different treatments and specifically at given time points. Data are presented as mean Ϯ SEM. A probability value of P < 0.05 was considered as statistically significant.
First, the role of GHS-R1A receptors in the VTA for the reinforcing effects of ghrelin by tests of ghrelin-induced locomotor stimulation and, in separate studies, by measurement of ghrelin-induced dopamine release were investigated. The locomotor stimulatory and accumbal dopamine releasing effects of ghrelin were attenuated by local administration of the GHS-R1A antagonist BIM28163 into the VTA (Fig 1a,b), at a dose shown previously to have no effect on locomotor stimulation and accumbal dopamine release per se (Jerlhag et al. 2009). Thus, ghrelin-induced locomotor stimulation (P < 0.01) was attenuated by VTA administration of BIM28163 (P < 0.01) in mice (F(3,25) = 5.45, P = 0.005: n = 6-8). In the microdialysis experiments a significant effect of systemic ghrelin to increase dopamine release in comparison to vehicle treatment was observed (P = 0.003). Pre-treatment with BIM28163 attenuated the ghrelininduced increase in dopamine release compared with vehicle pre-treatment in mice (P = 0.001) (treatment F(3,26) = 6.39, P = 0.002; time F(13,338) = 1.77, P = 0.047; treatment-time interaction F(13,338) = 4.01, P < 0.001). This difference was evident at the time intervals 20-100 minutes (P < 0.001: n = 7-8).
The ghrelin-induced locomotor stimulation (P < 0.01) was not affected by VTA administration of the orexin A receptor antagonist SB334867 (P > 0.05) in mice (F(3,24) = 8.44, P = 0.005: n = 6-8) (Fig. 2a). Likewise, the ghrelin-induced locomotor stimulation (P < 0.01) was not suppressed by i.p. injection of the opioid receptor antagonist naltrexone (P > 0.05) in mice (F(3,28) = 6.01, P = 0.003: n = 8) (Fig. 2b).
Intra-VTA administration of the NMDA receptor antagonist, AP5, abolished the ghrelin-induced locomotor stimulation and accumbal dopamine release (Figs 3a,b), at a dose that had no effect per se (Table 1). Specifically, the ghrelin-induced locomotor stimulation (P < 0.01) was attenuated by VTA administration of AP5 (P < 0.001) in mice (F(3,27) = 8.06, P < 0.001: n = 7-8). Moreover, systemic ghrelin increased dopamine release (P < 0.001) and pre-treatment with AP5 attenuated the (a) Locomotor activity Counts/60 min (b) Locomotor activity Counts/60 min Veh SB Veh SB Veh Veh Ghr Ghr 0 100 200 300 *** ** n.s. 0 50 100 150 Veh Nal Veh Nal Veh Veh Ghr Ghr *** ** n.s. Control experiments showed that neither VTA administration, the volume infused nor the antagonist per se had any effect on locomotor activity (Figs 1a, 2a,b and 3a) or accumbal dopamine release (Figs 1b and 3b).
For all locomotor activity experiments eight mice in each treatment group were implanted with bilateral guide cannulae, whereas 11 mice undertook surgery for each treatment group for the microdialyis experiments. All these mice were included in the experiments (locomotor activity or microdialysis studies). After the experiment the location of the probe and/or guide cannulae was verified and only mice with probe placement in the N.Acc. shell and/or cannula/e in the VTA were included in the statistical analysis (Fig. 4). Neither AP5 nor BIM28163 suppressed ghrelin-induced locomotor stimulation or accumbal dopamine release in a few mice in which the canulae were misplaced in neighbouring structures. It should also be emphasized that in a few mice the probe was located outside the N.Acc. shell and in these mice no effect of ghrelin on accumbal dopamine release was observed (data not shown).
The present study demonstrates that the stimulatory effects of peripheral ghrelin on locomotor stimulation and accumbal dopamine release involve VTA GHS-R1A signalling as both effects were suppressed by VTA administration of a GHS-R1A antagonist (BIM28163). Moreover, ghrelin's ability to activate the mesolimbic dopamine system was suppressed by pharmacological blockade of glutamatergic receptors but not by blockade of opioid or orexin A receptors. Thus, the locomotor stimulating effect of ghrelin was not affected by intra-VTA administration of the orexin A antagonist (SB334867) or by peripheral administration of an opioid receptor antagonist (naltrexone). Finally, ghrelininduced locomotor stimulation as well as accumbal dopamine release were suppressed by VTA administration of an NMDA receptor antagonist (AP5). Taken together these data suggest that systemic ghrelin activates the mesolimbic dopamine system via GHS-R1A in the VTA and that glutamate rather than orexin or opioid signalling is required for ghrelin to stimulate the mesolimbic dopamine system.
Given that transport of ghrelin across the blood-brain barrier into the brain is somewhat limited (Banks et al. 2002), there have been suggestions that ghrelin may exert its central effects via vagal afferents (Date et al. 2000) or by gaining access at circumventricular organs, such as the arcuate nucleus and area postrema. Studies using Fos protein to map ghrelin's central actions have shown that whereas peripheral administration of ghrelin and GHS-R1A agonists activate a rather limited population of cells in the arcuate nucleus (Dickson, Leng & Robinson 1993;Hewson & Dickson 2000) and area postrema (Bailey et al. 2000), additional hypothalamic cell groups were recruited after central administration (Lawrence et al. 2002). In the present study, it was demonstrated that the actions of peripheral ghrelin on the midbrain Bregma +1.5 mm (a) Bregma -3.4 mm (b) Figure 4 Verification of cannula/e and/or probe placement. A coronal mouse brain section showing ten representative probe placements (illustrated by vertical lines) in the N.Acc. (a) or guide cannula/e placements in the VTA (b) of mice used in the present study (Franklin & Paxinos 1996).Ten representative placements are illustrated, but all other placements were within the N.Acc. shell or in the VTA. Placements outside either of these areas were not included in the statistical analysis.The number given in each brain section indicates millimetres anterior (+) and posterior (-) from bregma.
dopamine system, reflected by locomotor stimulation and accumbal dopamine release, can be blocked by VTA administration of a GHS-R1A antagonist. Taken together with previous neuroanatomical and accumbal dopamine measurement studies showing that the target cells for ghrelin in the VTA include the dopaminergic cell group (Abizaid et al. 2006;Jerlhag et al. 2006;Jerlhag 2008;Kawahara et al. 2009;Quarta et al. 2009), it seems likely that peripheral ghrelin directly activates the VTA dopamine system via GHS-R1A. The activation of the mesolimbic dopamine system by ghrelin may be due to the reported ability of the GHS-R1A to dimerize with the dopamine D1 receptor, both receptors expressed on dopamine neurons in the VTA, and thereby amplifies the dopamine signalling (Jiang, Betancourt & Smith 2006). Central ghrelin signalling system, including the GHS-R1A, appears to be required for reinforcement induced by addictive drugs including cocaine (Wellman, Davis & Nation 2005;Davis, Wellman & Clifford 2007;Tessari et al. 2007) and alcohol (Jerlhag et al. 2009). The mesolimbic dopamine system appears to be a likely target for ghrelin also in man, as evidenced from functional MRI studies in which peripheral ghrelin altered the response of the ventral striatum to visual food cues (Malik et al. 2008). Indeed, ghrelin may, via VTA GHS-R1A, increase the incentive value for natural as well as chemical reinforcements.
Neurotransmitters, including orexin, opioids as well as glutamate, have previously been shown to modulate the intake of natural and chemical reinforcers of the mesolimbic dopamine system as well as to regulate the activity of these VTA neurons (Hoebel et al. 1989;Engel et al. 1992;Wise 2002;Thiele et al. 2003). Specifically, AP5 and SB334867, in a similar dose range, attenuates the reinforcing properties of drugs of abuse (Herz 1997;Taber & Fibiger 1997;Smith et al. 2009). Here, neither the opioid receptor-nor the orexin A receptorantagonist affected ghrelin-induced locomotor stimulation, indicating that these systems do not interfere with these effects of ghrelin. Consistent with our findings, the orexigenic response to ghrelin when administered into key mesolimbic dopamine structures such as the VTA has previously been shown to be independent of opioid receptor signalling (Naleid et al. 2005). Peripheral injection of naltrexone has been shown to blocks the reinforcing properties of alcohol in rodents (Herz 1997), suggesting that naltrexone passes the blood-brain barrier and has central effects. Even though opioid receptors, specifically the m receptor, mediate drug-induced reinforcement this receptor appear to be less important for the ability of ghrelin to activate the mesolimbic dopamine system. Orexin-containing neurons have previously been suggested to regulate ghrelin-induced feeding (Toshinai et al. 2003), but orexin A receptors do not appear to be crucial for the locomotor stimulatory effects of ghrelin. Collectively, these data suggest that ghrelin-induced activation of the mesolimbic dopamine system appears to be regulated via other mechanisms than its orexigenic properties. In the present study, it was shown that ventral tegmental NMDA receptors are required for ghrelininduced locomotor stimulation and accumbal dopamine release. Supportively, the effects of ghrelin to increase the electrical activity of dopaminergic neurons in the VTA appears to be dependent on the excitatory glutamatergic input and also blockade of NMDA receptors in the VTA reduces food-induced accumbal dopamine release (Taber & Fibiger 1997;Abizaid et al. 2006). NMDA receptors have also been shown to mediate the accumbal dopamine release observed when animals consume food after ghrelin administration (Kawahara et al. 2009). As shown previously, ghrelin-induced reinforcement also involves nicotinic acetylcholine receptors in the VTA (Jerlhag et al. 2006;Jerlhag et al. 2008). Collectively, these data suggest that neurotransmitters including acetylcholine and glutamate are required for ghrelin-induced reinforcement, which it has in common with natural as well as chemical reinforcers of the mesolimbic dopamine system. The mechanisms for the interaction between ghrelin, acetylcholine and glutamate in the VTA are still unclear. However, presynaptic nicotinic acetylcholine receptors have been shown to modulate the release of glutamate in the VTA, which via postsynaptic NMDA receptors regulate accumbal dopamine release (Schilström et al. 1998;Schilström et al. 2000). Tentatively, ghrelin may increase the acetylcholine release, which via such mechanisms may indirectly activate the mesolimbic dopamine system. However, the possibility of the existence of presynaptic NMDA receptors on cholinergic neurons cannot be excluded (Corlew et al. 2008). The possibility that ghrelin has an ability to rearrange the excitatory NMDAmediated synaptic input in a manner that would increase the probability of activation of the dopamine neurons by other inputs should also be considered; thus findings in the hippocampus as well as VTA show that ghrelin has effects on synaptic plasticity (Abizaid et al. 2006;Diano et al. 2006). The glutamatergic afferents to the VTA mainly originate from the prefrontal cortex, lateral hypothalamus, bed nucleus of stria terminalis, superior coliculus and the LDTg. This input regulates the activity of dopamine via NMDA receptors (Schilström et al. 1998;Georges & Aston-Jones 2002;Geisler & Zahm 2005). It should however be emphasized that other afferents including GABA, serotonin and noradrenalin, known to modulate the activity of dopamine projections to the N.Acc. (Wise 2002), may also have important roles for the ghrelin-induced activation of the mesolimbic dopamine system. Although, GABAA receptors in the VTA does not mediate the increase in accumbal dopamine observed after food consumption induced by ghrelin (Kawahara et al. 2009).
In summary, the present study shows that the effects of peripheral ghrelin on locomotor stimulation and accumbal dopamine release in mice (that reflect direct actions of ghrelin at the level of the mesolimbic dopamine system, specifically the VTA) can be suppressed by an NMDA antagonist and are therefore likely to be under glutamatergic control. These data may have clinical implications since hyperghrelinemia is associated with addictive behaviours including compulsive overeating and alcohol use disorder (Cummings et al. 2001;Kim et al. 2005;Kraus et al. 2005). It may therefore be proposed that ghrelin-responsive circuits at the level of the VTA, that appear to be sensitive to cholinergic and glutamatergic input, may serve as a novel pharmacological target for treatment of such addictive behaviours.
The NMDA receptor antagonist, AP5, decreased the locomotor activity at a dose of 1 mg/side and of 2 mg/side bilaterally into the VTA per se compared to vehicle treatment in mice. However, a dose of 0.5 mg/side bilaterally into the VTA did not affect the locomotor activity per se (F(3,12) = 12.72, P = 0.0005, n = 4 in each group. **P < 0.01, ***P < 0.001, n.s. P > 0.05, Tukey's HSD post-hoc test).
© 2010 The Authors, Addiction Biology © 2010 Society for the Study of AddictionAddiction Biology, 16, 82-91
Acknowledgements Supported by the
EJ conducted the experiments and analyzed the data. All authors contributed to the writing of the manuscript. JAE designed research. All authors have critically reviewed content and approved final version submitted for publication.
Aim Swedish studies have shown that experience of using snus is associated with an increased probability of being a former smoker. We examined whether this result is also found in Norway. Design Seven cross-sectional data sets collected during the period 2003-08. Setting Norway. Participants A total of 10 441 ever (current or former) smokers Measurements Quit ratios for smoking were compared for people with different histories of snus use. Motive for snus use was examined among combination users (snus and cigarettes). Smoking status was examined among snus users. Findings Compared to smokers with no experience of using snus, the quit ratio for smoking was significantly higher for daily snus users in six of seven data sets, significantly higher for former snus users in two of five data sets and significantly lower for occasional snus users in six of seven data sets. Of combination users who used snus daily, 55.3% [confidence interval (CI) 44.7-65.9] reported that their motive for using snus was to quit smoking totally. This motive was reported significantly less often by combination users who used snus occasionally (35.7%, CI 27.3-44.2). Former smokers made up the largest proportion of daily snus users in six of seven data sets. In the remaining data set, that included only the age group 16-20 years, people who had never smoked made up the largest segment of snus users. Conclusions Consistent with Swedish studies, Norwegian data shows that experience of using snus is associated with an increased probability of being a former smoker. In Scandinavia, snus may play a role in quitting smoking but other explanations, such as greater motivation to stop in snus users, cannot be ruled out.
In the European Union, with Sweden as the only exception, snus-a low-nitrosamine smokeless tobacco product-has been banned since 1992. However, with regard to health, there is disagreement about whether the ban is appropriate. The main argument against the ban, raised by bodies such as the Royal College of Physicians of London [1] and the European Respiratory Society [2], is that the ban deprives smokers who are seriously addicted to nicotine a harm-reducing alternative to cigarettes. In the United States, the American Association of Public Health Physicians [3] and a series of eminent tobacco researchers are in favour of including snus within the arsenal of harm-reducing nicotine products. Their recommendations are based upon the conviction that use of snus can contribute to cessation of, or a dramatic reduction in, smoking-the most harmful form of nicotine intake.
However, few randomized controlled trials (RCT) have been carried out to assess the use of snus on smoking cessation. In the absence of experimental studies, observational data have been used to illuminate this issue. In an authoritative report from the European Commission [4], it is claimed that trends in the use of tobacco in Norway indicate that the availability of snus cannot have had much influence on the number of smokers. The reason given is that although snus users are greatly over-represented by men, the rate of reduction in smoking for men is not much greater than that for women. The organization Physicians for a Smoke-free Canada [5] has claimed the same, on the grounds that the reduction in smoking in Norway and Sweden is no greater than in countries such as Canada and Finland, where snus is seldom used.
This type of comparative analysis of trends in the prevalence of smoking for different genders or different countries can, at best, give only a weak indication of the importance of snus for smoking cessation. In epidemiological research of tobacco behaviour, a more precise measure-the quit ratio-can be used to calculate the proportion of people in a population who have stopped smoking [6]. Swedish prospective [7] and retrospective [8] studies have shown that the quit ratio is higher for smokers who use, or have used, snus than for smokers who have never used snus. We have examined whether this same result is also found in a series of recent Norwegian cross-sectional studies, in which it has been possible to measure the quit ratio for daily smoking. We have also studied which motives people who smoke and use snus (combination users) give for their use of snus.The implications of these results for public health are then discussed.
The quit ratio for smoking is an expression of the number of former daily smokers as a proportion of the total number of people who have ever smoked daily in a population. It is a statistical measure of smoking cessation activity that is recommended for use in tobacco behaviour research [6]. In order to identify significant differences in quit ratio between different groups of snus users, 95% confidence intervals (CIs) and P-values were calculated. As snus use is a predominantly male phenomenon in Norway, no gender-specific results were presented. In the samples with sufficient females for a test, no interaction by gender was observed with regard to quit ratios for smoking across snus status.
The Norwegian Institute for Alcohol and Drug Research collect data on risk-related behaviour in the population on a regular basis. Seven studies in the Institute's databank, carried out since 2003, included questions that made it possible to calculate quit ratios for smoking cigarettes across snus use status (daily, occasional, former and never user). Smoking habits were assessed using the question: 'Do you smoke?' (yes, daily/ yes, occasionally/no, not at all). In all seven surveys the definition of former smoking was based upon subjective response to a question presented to current non-smokers at the time of the survey: 'Have you ever smoked daily?' (yes/no). Current and former snus use was assessed using similar questions.
Study population 1 included 3604 current or former smokers of both genders in the age group 16-74 years. The subjects were identified from a data pool from yearly representative surveys of tobacco behaviour carried out by Statistics Norway (SSB) by telephone for the years 2003-08, and included 7500 respondents in total. The mean response rate for the period was 67%. The material and methods for these surveys have been described previously [9]. The motives for using snus among dual users of snus and cigarettes in this study were assessed by asking how well-on a scale from 1 (apply fully) to 5 (do not apply at all)-the statements listed in Table 4 described their situation. This question was included in the surveys for the period 2005-08. The percentages for those who scored 1 or 2 on the scales are displayed in Table 4. In order to identify significant differences in motives for snus use between daily and occasional snus users, 95% CIs were calculated.
Study population 2 included 423 current or former smokers of both genders in the age group 16-20 years. This sample was drawn from a national representative survey of tobacco behaviour carried out by telephone by Synovate, Norway in September 2007 among 2415 people. The material and methods have been described previously [10].
Study population 3 included 790 women and men born later than 1970 who were students at Oslo University (UiO) in 2007, and who reported that previously or at the time of the survey they had smoked cigarettes daily. They were identified from a survey of 1655 students with information about use of alcohol, drugs and tobacco. The survey was carried out by a mailed questionnaire by the Foundation for Student Life in Oslo (SIO), in cooperation with the Norwegian Institute for Alcohol and Drug Research (SIRUS). The response rate was 57%. The material and methods have been described previously [11].
Study population 4 included 2018 smokers and former smokers of both genders in the age group 15-91 years in 2007. These people were selected from Synovate, Norway's national representative Norwegian Monitor Survey of 3683 people carried out by home visits. The material and methods have been described previously [12].
Study population 5 included 729 smokers and former smokers of both genders in the age group 21-30 years, selected from a national survey of 2362 people carried out by mailed questionnaire by SIRUS in 2006. The response rate was 42%. The material and methods have been described previously [13].
Study population 6 included 639 smokers and former smokers of both genders in the age group 21-30 years selected from a survey of 2270 young adults in the Norwegian capitol Oslo, carried out by mailed questionnaire by SIRUS in 2006. The response rate was 45%. The material and methods have been described previously [13].
Study population 7 consisted of 2572 male ever smokers aged 20-50 years selected randomly from a national representative web panel. Of those invited to participate, 7170 men (48.6%) responded to an electronic questionnaire. The material and methods have been described previously [14].
Table 1 presents information about the seven surveys with 27 955 respondents in total, in which questions about use of tobacco and snus were asked. A total of 10 441 people (both users and non-users of snus) reported that they had either smoked daily previously but had quit (5144) or that they smoked daily at the time of the survey (5207) (Table 2). In the student population in Oslo (study 3), the quit ratio for smoking was 67.4% (95% CI 63.1-71.7). This was significantly higher than in the other samples. The lowest quit ratio of 32.2% (95% CI 27.8-36.7) was observed in the nationally representative sample of young people in the age group 16-20 years (study 2). The quit ratio was 52% in both of the two national representative samples that had the same age and gender distribution as the normal population (Table 2).
In six of the seven studies the quit ratio for smoking was significantly higher for daily snus users than for people with no experience of using snus (Table 3). The same finding was observed for people who at the time of the study reported that they had used snus previously in five of the six studies, but the differences were significant in only two of the studies. The quit ratio for smoking for people who used snus occasionally was significantly lower than for people who had no experience of snus use in six of seven studies.
For snus users who also smoked cigarettes on a daily basis, 43.8% of them reported that they used snus to quit smoking, 56.3% to reduce their smoking and 59.1% to replace cigarettes in places where smoking is not allowed (Table 4). All three motives for using snus were more common for smokers who used snus on a daily basis compared to smokers who used snus occasionally, but this was significant only for those reporting using snus to quit smoking: 55.3% daily users (95% CI 44.7-65.9) compared with 35.7% occasional users (95% CI 27.3-44.2).
In the survey that included only young people in the age group 16-20 years (study 2), 42.8% of daily snus users were without previous smoking experience, while only 17.9% were former smokers (Table 5). In the other six surveys the proportion of snus users who were former smokers formed the largest group among snus users (varying between 34.4 and 42.5% of all snus users). The proportion of snus users who had never smoked cigarettes varied from 18.8 to 31% in the same six surveys. A small minority of snus users (3.1-10.6%) smoked daily, but many more smoked occasionally (16-35%).
Snus appeared to play a role in quitting smoking: daily snus use and, to a lesser extent, former snus use, were associated with being a former smoker across most of the included studies; occasional snus use was less likely to be associated with being a former smoker. However, the results must be interpreted with caution, as with nonrandomized observational studies we cannot exclude the danger of selection bias in the groups that are compared.
In contrast to Sweden, where nearly all snus users are daily users, almost half of Norwegian snus users use snus only occasionally [15]. In this group, the quit ratio for smoking was low (Table 3), and the proportion of combination users of snus and cigarettes was accordingly high. A possible interpretation is that occasional snus users have a stable combination use, in which snus is used more as a substitute for cigarettes, for example in the steadily increasing social arenas where cigarette smoking is undesirable. An alternative interpretation of this relationship is that many of these low-frequent snus users at the time of the interview were in an incomplete transition phase of stopping smoking daily, and that they will replace cigarettes with daily use of snus later. A clarification of the two hypotheses requires longitudinal data.
Use of snus can damage health in a number of different ways, some serious and some less serious, but systematic literature reviews [4] and register-based prospective studies [16] have concluded that snus is much less hazardous than smoking. It is estimated that the reduction in risk by switching from cigarettes to snus is at least 90% [17]. If snus increases the rate of smoking cessation, what are the consequences for public health if, overall, the use of snus increases? The extent and nature of the impact on public health will depend upon the relative risk hazard of snus and smoking, and the relative uptake and use by smokers and non-smokers.
To identify the net effect of snus use from a public health perspective is a complicated task. However, the conditions for carrying out this task are best in countries such as Norway and Sweden, using our observational data on the transition between cigarettes and snus.
In all the studies described here, except the one of people under 20 years of age, people who had quit smoking formed the largest group of snus users. The proportion of snus users who had no previous experience of smoking was less than the proportion of ex-smokers in all the studies, except for the youngest respondents. One frequently reported study [18] found that 14-25 nonsmokers would have to begin to use snus in order to cancel out the positive health effects from each smoker who transferred from cigarettes to snus. The results from the studies presented here suggest that in Norway, as in Sweden, snus users are more likely to be recruited from smokers than from non-smokers. This supports the hypothesis that availability of snus has more positive effects on public health than negative effects-at least in these two countries.
A small minority of snus users (3.1-10.6%) smoked daily, while many more smoked occasionally (16-35%). As we lack data about reduced smoking frequency among these combination users-as shown in Sweden [19]-it is difficult to draw any conclusions about whether this combination use is more or less damaging than the amount of smoking that would have taken place without snus use.
The picture was different in the survey that included only young people in the age group 16-20 years (survey 2). In this survey as many as 42.8% of snus users had no smoking experience, while only 17.9% were former smokers. This result probably emanates from the fact that the general quit ratio for smoking is very low in such a young age group (Table 2). Thus snus, probably along with medicinal products, is not used as a method for quitting smoking to the same degree among young people as among older established smokers.
The largest increase in use of snus during the last few years has been among young people [15]. If almost half these recruits are primary users, as our surveys indicate, there is cause for concern that increased use of snus does not result in harm reduction in this cohort, but increased health risk. Conversely, we might assume that a segment of primary users are young people who would otherwise have begun to smoke if snus had not been available and would therefore have been exposed to a product with greater health risk. Studies have shown that young people who initiate tobacco use through using snus have many of the same predisposing factors as young smokers [20], and that this may indicate that the products recruit users from the same segment of the population. Indeed, a longitudinal study from Sweden [21] has shown that use of snus reduces the risk of starting to smoke among young people, when known predictors for starting to smoke are controlled for. Data from the United States are more inconsistent [22,23]. At any rate, a large proportion of people must be recruited to be primary users of snus in order to cancel the health benefits for every young person who starts to use snus instead of smoking [18,24].
Even assuming that the comparison of the quit ratio for smoking by use of snus gives a much more valid indication of the importance of snus for quitting smoking than gender-specific or country-specific comparisons of the trend in the prevalence of smoking, the quit ratio is still a fairly crude measure of quitting activity [6]. Although former smoking was defined consistently across these seven studies, our definition did not distinguish between former smokers according to how long ago they had stopped. This is a weakness.
In Norway, as in Sweden, the quit ratio for daily smoking is higher for daily snus users than for people who have never used snus. People who use snus occasionally have the lowest quit ratio for smoking. Former smokers make up the largest segment of Norwegian snus users, except among the 16-20-year-olds. These results indicate that for many smokers use of snus may represent an exit from cigarette smoking. The larger number of primary snus users in the younger age group may lead to a change in the future composition of snus users according to their previous smoking experience.
© 2010 The Authors, Addiction © 2010 Society for the Study of AddictionAddiction, 106, 162-167
SIRUS: Norwegian Institute for Alcohol and Drug Research; SIO: Foundation for Student Life in Oslo; SSB: Statistics Norway.
Funding for this research has been provided by the
None.
An unresolved issue in the field of implementation research is how to conceptualize and evaluate successful implementation. This paper advances the concept of ''implementation outcomes'' distinct from service system and clinical treatment outcomes. This paper proposes a heuristic, working ''taxonomy'' of eight conceptually distinct implementation outcomes-acceptability, adoption, appropriateness, feasibility, fidelity, implementation cost, penetration, and sustainability-along with their nominal definitions. We propose a two-pronged agenda for research on implementation outcomes. Conceptualizing and measuring implementation outcomes will advance understanding of implementation processes, enhance efficiency in implementation research, and pave the way for studies of the comparative effectiveness of implementation strategies.
A critical yet unresolved issue in the field of implementation science is how to conceptualize and evaluate success. Studies of implementation use widely varying approaches to measure how well a new mental health treatment, program, or service is implemented. Some infer implementation success by measuring clinical outcomes at the client or patient level while other studies measure the actual targets of the implementation, quantifying for example the desired provider behaviors associated with delivering the newly implemented treatment. While some studies of implementation strategies assess outcomes in terms of improvement in process of care, Grimshaw et al. (2006) report that metaanalyses of their effectiveness has been thwarted by lack of detailed information about outcomes, use of widely varying constructs, reliance on dichotomous rather than continuous measures, and unit of analysis errors.
This paper advances the concept of ''implementation outcomes'' distinct from service system outcomes and clinical treatment outcomes (Proctor et al. 2009;Fixsen et al. 2005;Glasgow 2007a). We define implementation outcomes as the effects of deliberate and purposive actions to implement new treatments, practices, and services. Implementation outcomes have three important functions. First, they serve as indicators of the implementation success. Second, they are proximal indicators of implementation processes. And third, they are key intermediate outcomes (Rosen and Proctor 1981) in relation to service system or clinical outcomes in treatment effectiveness and quality of care research. Because an intervention or treatment will not be effective if it is not implemented well, implementation outcomes serve as necessary preconditions for attaining subsequent desired changes in clinical or service outcomes.
Distinguishing implementation effectiveness from treatment effectiveness is critical for transporting interventions from laboratory settings to community health and mental health venues. When such efforts fail, as they often do, it is important to know if the failure occurred because the intervention was ineffective in the new setting (intervention failure), or if a good intervention was deployed incorrectly (implementation failure). Our current knowledge of implementation is thwarted by lack of theoretical understanding of the processes involved (Michie et al. 2009). Conceptualizing and measuring implementation outcomes will advance understanding of implementation processes, enable studies of the comparative effectiveness of implementation strategies, and enhance efficiency in implementation research.
This paper aims to advance the ''vocabulary'' of implementation science around implementation outcomes through four specific objectives: (1) to advance conceptualization of implementation outcomes by distinguishing implementation outcomes from service and clinical outcomes; (2) to advance clarity of terminology currently used in implementation science by nominating heuristic definitions of implementation outcomes, yielding a working ''taxonomy'' of implementation outcomes; (3) to reflect the field's current language, conceptual definitions, and approaches to operationalizing implementation outcomes; and (4) to propose directions for further research to advance knowledge on these key constructs and their interrelationships.
Our objective of advancing a taxonomy of implementation outcomes is comparable to the work of Michie et al. (2005Michie et al. ( , 2009)), Grimshaw et al. (2006), the Cochrane group, and others who are working to develop taxonomies and common nomenclature for implementation strategies. Our work is complementary to these efforts because implementation outcomes will provide researchers with a framework for evaluating implementation strategies.
Our understanding of implementation outcomes is lodged within a previously published conceptual framework (Proctor et al. 2009) as shown in Fig. 1. The framework distinguishes between three distinct but interrelated types of outcomes-implementation, service, and client outcomes. Improvements in consumer well-being provide the most important criteria for evaluating both treatment and implementation strategies-for treatment research, improvements are examined at the individual client level whereas improvements at the population-level (within the providing system) are examined in implementation research. However, as we argued above, implementation research requires outcomes that are conceptually and empirically distinct from those of service and clinical effectiveness.
For heuristic purposes, our model positions implementation outcomes as preceding both service outcomes and client outcomes, with the latter sets of outcomes being impacted by the implementation outcomes. As we discuss later in this paper, interrelationships among these outcomes require conceptual mapping and empirical tests. For example, one would expect to see a treatment's strongest impact on client outcomes as an empirically supported treatment's (EST) penetration increases in a service setting-but this hypothesis requires testing. Our model derives service outcomes from the six quality improvement aims set out in the reports on crossing the quality chasm: the extent to which services are safe, effective, patientcentered, timely, efficient, and equitable (Institute of Medicine Committee on Crossing the Quality Chasm 2006; Institute of Medicine Committee on Quality of Health Care in America 2001).
The paper's methods were shaped around its overall aim: to advance clarity in the language used to describe outcomes of implementation. We convened a working group of implementation researchers to identify concepts for labeling and assessing outcomes of implementation processes. One member of the group was a doctoral student RA who coordinated, conducted, and reported on the literature search and constructed tables reflecting various iterations of the heuristic taxonomy. The RA conducted literature searches using key words and search programs to identify literature on the current state of conceptualization and measurement of these outcomes, primarily in the health and behavioral sciences. We searched in a number of databases with a particular focus on MEDLINE, CINAHL Plus, and PsycINFO. Key search terms included the name
Client Outcomes Satisfaction Function Symptomatology Service Outcomes* Efficiency Safety Effectiveness Equity Patientcenteredness Timeliness Implementation Outcomes Acceptability Adoption Appropriateness Costs Feasibility Fidelity Penetration Sustainability *IOM Standards of Care of the implementation outcome (e.g., ''acceptability,'' ''sustainability,'' etc.) along with relevant synonyms combined with any of the following: innovation, EBP, evidence based practice, and EST. We scanned the titles and abstracts of the identified sources and read the methods and background sections of the studies that measured or attempted to measure implementation outcomes. We also included information from relevant conceptual articles in the development of nominal definitions. Whereas our primary focus was on the implementation of evidence based practices in the health and behavioral sciences, the keyword ''innovation'' broadened this scope by also identifying studies that focused on other areas such as physical health that may inform implementation of mental health treatments. Because terminology in this field currently reflects widespread inconsistency, we followed leads beyond what our keyword searches ''hit'' upon. Thus we read additional articles that we found cited by authors whose work we found through our electronic searches. We also conducted searches of CRISP, TAGG, and NIH reporter and studies to identify funded mental health research studies with ''implementation'' in their titles or abstracts, to identify examples of outcomes pursued in current research.
We used a narrative review approach (Educational Research Review), which is appropriate for summarizing different primary studies and drawing conclusions and interpretation about ''what we know,'' informed by reviewers' experiences and existing theories (McPheeters et al. 2006;Kirkevoid 1997). Narrative reviews yield qualitative results, with strengths in capturing diversities and pluralities of understanding (Jones 1997). According to McPheeters et al. (2006), narrative reviews are best conducted by a team. Members of the working group read and reviewed conceptual and theoretical pieces as well as published reports of implementation research. As a team, we convened recurring meetings to discuss the similarities and dissimilarities. We audio-taped and transcribed meeting discussions, and a designated individual took thorough notes. Transcriptions and notes were posted on a shared computer file for member review, revision, and correction.
Group processes included iterative discussion, checking additional literature for clarification, and subsequent discussion. The aim was to collect and portray, from extant literature, the similarities and differences across investigators' use of various implementation outcomes and definitions for those outcomes. Discussions often led us to preserve distinctions between terms by maintaining in our ''nominated'' taxonomy two different implementation outcomes because the literature or our own research revealed possible conceptual distinctions. We assembled the identified constructs in the proposed heuristic taxonomy to portray the current state of vocabulary and conceptualization of terms used to assess implementation outcomes.
Through our process of iterative reading and discussion of the literature, we worked to nominate definitions that (1) achieve as much consistency as possible with any existing definitions (including multiple definitions we found for a single construct), yet (2) serve to sharpen distinctions between constructs that might be similar. For several of the outcomes, the literature did not offer one clear nominal definition.
Table 1 depicts the resultant working taxonomy of implementation outcomes. For each implementation outcome, the table nominates a level of analysis, identifies the theoretical basis to the construct from implementation literature, shows different terms that are used for the construct in the literature, suggests the point or stage within implementation processes at which the outcome may be most salient, and lists the types of existing measures for the construct that our search identified. The implementation outcomes listed in Table 1 are probably only the ''more obvious,'' and we expect that other concepts may emerge from further analysis of the literature and from the kind of empirical work we call for in our discussion below. Many of the implementation outcomes can be inferred or measured in terms of expressed attitudes and opinions, intentions, or reported or observed behaviors. We now list and discuss our nominated conceptual definitions for each implementation outcome in our proposed taxonomy. We reference similar definitions from the literature, and also comment on marked differences between our definitions and others proposed for the term.
Acceptability is the perception among implementation stakeholders that a given treatment, service, practice, or innovation is agreeable, palatable, or satisfactory. Lack of acceptability has long been noted as a challenge in implementation (Davis 1993). The referent of the implementation outcome ''acceptability'' (or the ''what'' is acceptable) may be a specific intervention, practice, technology, or service within a particular setting of care. Acceptability should be assessed based on the stakeholder's knowledge of or direct experience with various dimensions of the treatment to be implemented, such as its content, complexity, or comfort. Acceptability is different from the larger construct of service satisfaction, as typically measured through consumer surveys. Acceptability is more specific, referencing a particular treatment or set of treatments, while satisfaction typically references the general service experience, including such features as waiting times, scheduling, and office environment. Acceptability may be measured from the perspective of (Aarons 2004). Aarons and Palinkas (2007) used semi-structured interviews to assess case managers' acceptance of evidencebased practices in a child welfare setting. Karlsson and Bendtsen (2005) measured patients' acceptance of alcohol screening in an emergency department setting using a 12-item questionnaire.
Adoption is defined as the intention, initial decision, or action to try or employ an innovation or evidence-based practice. Adoption also may be referred to as ''uptake.'' Our definition is consistent with those proposed by Rabin et al. (2008) and Rye and Kimberly (2007). Adoption could be measured from the perspective of provider or organization. Haug et al. (2008) used pre-post items to capture substance abuse providers' adoption of evidence-based practices, while Henggeler et al. (2008) report interview techniques to measure therapists' adoption of contingency management.
Appropriateness is the perceived fit, relevance, or compatibility of the innovation or evidence based practice for a given practice setting, provider, or consumer; and/or perceived fit of the innovation to address a particular issue or problem. ''Appropriateness'' is conceptually similar to ''acceptability,'' and the literature reflects overlapping and sometimes inconsistent terms when discussing these constructs. We preserve a distinction because a given treatment may be perceived as appropriate but not acceptable, and vice versa. For example, a treatment might be considered a good fit for treating a given condition but its features (for example, rigid protocol) may render it unacceptable to the provider. The construct ''appropriateness'' is deemed important for its potential to capture some ''pushback'' to implementation efforts, as is seen when providers feel a new program is a ''stretch'' from the mission of the health care setting, or is not consistent with providers' skill set, role, or job expectations. For example, providers may vary in their perceptions of the appropriateness of programs that co-locate mental health services within primary medical, social service, or school settings. Again, a variety of stakeholders will likely have perceptions about a new treatment's or program's appropriateness to a particular service setting, mission, providers, and clientele. These perceptions may be function of the organization's culture or climate (Klein and Sorra 1996). Bartholomew et al. (2007) describe a rating scale for capturing appropriateness of training among substance abuse counselors who attended training in dual diagnosis and therapeutic alliance.
Cost (incremental or implementation cost) is defined as the cost impact of an implementation effort. Implementation costs vary according to three components. First, because treatments vary widely in their complexity, the costs of delivering them will also vary. Second, the costs of implementation will vary depending upon the complexity of the particular implementation strategy used. Finally, because treatments are delivered in settings of varying complexity and overheads (ranging from a solo practitioner's office to a tertiary care facility), the overall costs of delivery will vary by the setting. The true cost of implementing a treatment, therefore, depends upon the costs of the particular intervention, the implementation strategy used, and the location of service delivery.
Much of the work to date has focused on quantifying intervention costs, e.g., identifying the components of a community-based heart health program and attaching costs to these components (Ronckers et al. 2006). These cost estimations are combined with patient outcomes and used in cost-effectiveness studies (McHugh et al. 2007). A review of literature on guideline implementation in professions allied to medicine notes that few studies report anything about the costs of guideline implementation (Callum et al. 2010). Implementing processes that do not require ongoing supervision or consultation, such as computerized medical record systems, may carry lower costs than implementing new psychosocial treatments. Direct measures of implementation cost are essential for studies comparing the costs of implementing alternative treatments and of various implementation strategies.
Feasibility is defined as the extent to which a new treatment, or an innovation, can be successfully used or carried out within a given agency or setting (Karsh 2004). Typically, the concept of feasibility is invoked retrospectively as a potential explanation of an initiative's success or failure, as reflected in poor recruitment, retention, or participation rates. While feasibility is related to appropriateness, the two constructs are conceptually distinct. For example, a program may be appropriate for a service setting-in that it is compatible with the setting's mission or service mandate, but may not be feasible due to resource or training requirements. Hides et al. (2007) tapped aspects of feasibility of using a screening tool for co-occurring mental health and substance use disorders.
Fidelity is defined as the degree to which an intervention was implemented as it was prescribed in the original protocol or as it was intended by the program developers (Dusenbury et al. 2003;Rabin et al. 2008). Fidelity has been measured more often than the other implementation outcomes, typically by comparing the original evidencebased intervention and the disseminated/implemented intervention in terms of (1) adherence to the program protocol, (2) dose or amount of program delivered, and (3) quality of program delivery. Fidelity has been the overriding concern of treatment researchers who strive to move their treatments from the clinical lab (efficacy studies) to real-world delivery systems. The literature identifies five implementation fidelity dimensions including adherence, quality of delivery, program component differentiation, exposure to the intervention, and participant responsiveness or involvement (Mihalic 2004;Dane and Schneider 1998). Adherence, or the extent to which the therapy occurred as intended, is frequently examined in psychotherapy process and outcomes research and is distinguished from other potentially pertinent implementation factors such as provider skill or competence (Hogue et al. 1996). Fidelity is measured through self-report, ratings, and direct observation and coding of audio-and videotapes of actual encounters, or provider-client/patient interaction. Achieving and measuring fidelity in usual care is beset by a number of challenges (Proctor et al. 2009;Mihalic 2004;Schoenwald et al. 2005). The foremost challenge may be measuring implementation fidelity quickly and efficiently (Hayes 1998). Schoenwald and colleagues (2005) have developed three 26-45-item measures of adherence at the therapist, supervisor and consultant level of implementation (available from the MST Institute www.mstinstitute.org). Ratings are obtained at regular intervals, enabling examination of the provider, clinical supervisor, and consultant. Other examples from the mental health literature include Bond et al. (2008) 15-item Supported Employment Fidelity Scale (SE Fidelity Scale) and Hogue et al. (2008) Therapist Behavior Rating Scale-Competence (TBRS-C), an observational measure of fidelity in evidence based practices for adolescent substance abuse treatment.
Penetration is defined as the integration of a practice within a service setting and its subsystems. This definition is similar to (Stiles et al. 2002) notion of service penetration and to Rabin et al.s' (2008) notion of niche saturation. Studying services for persons with severe mental illness, Stiles et al. (2002) apply the concept of service penetration to service recipients (the number of eligible persons who use a service, divided by the total number of persons eligible for the service). Penetration also can be calculated in terms of the number of providers who deliver a given service or treatment, divided by the total number of providers trained in or expected to deliver the service. From a service system perspective, the construct is also similar to ''reach'' in the RE-AIM framework (Glasgow 2007b). We found infrequent use of the term penetration in the implementation literature; though studies seemed to tap into this construct with terms such a given treatment's level of institutionalization.
Sustainability is defined as the extent to which a newly implemented treatment is maintained or institutionalized within a service setting's ongoing, stable operations. The literature reflects quite varied uses of the term ''sustainability,'' but our proposed definition incorporates aspects of those offered by Johnson et al. (2004), Turner and Sanders (2006), Glasgow et al. (1999), Goodman et al. (1993), and Rabin et al. (2008). Rabin et al. (2008) emphasizes the integration of a given program within an organization's culture through policies and practices, and distinguishes three stages that determine institutionalization: (1) passage (a single event such as transition from temporary to permanent funding), (2) cycle or routine (i.e., repetitive reinforcement of the importance of the evidence-based intervention through including it into organizational or community procedures and behaviors, such as the annual budget and evaluation criteria), and (3) niche saturation (the extent to which an evidence-based intervention is integrated into all subsystems of an organization). Thus the outcomes of ''penetration'' and ''sustainability'' may be related conceptually and empirically, in that higher penetration may contribute to long-term sustainability. Such relationships require empirical test, as we elaborate below. Indeed Steckler et al. (1992) emphasize sustainability in terms of attaining long-term viability, as the final stage of the diffusion process during which innovations settle into organizations. To date, the term sustainability appears more frequently in conceptual papers than actual empirical articles measuring sustainability of innovations. As we discuss below, the literature often uses the same term (niche saturation, for example) to reference multiple implementation outcomes, underscoring the need for conceptual clarity as we seek to advance in this paper.
Advancing the conceptualization, measurement, and empirical understanding of implementation outcomes requires research on several critical issues. We propose two major themes for this research-(1) conceptualization and measurement, and (2) theory building-and identify important issues within each of these themes.
Research on several fronts is required to advance the conceptual and measurement properties of implementation outcomes, five of which we identify and discuss.
For each outcome listed in Table 1, we found literature using different and sometimes inconsistent terminology. Sometimes studies used different labels for what appear to be the same construct. In other cases, studies used one term for a label or nominal definition but a different term for operationalizing or measuring the same construct. This problem was pronounced for three implementation outcomes-acceptability, appropriateness, and feasibility. These constructs were frequently used interchangeably or measured under the common generic label as client or provider perceptions, reactions, and attitudes toward, or satisfaction with various aspects of the innovation, EST, or clinical practice guidelines. For example, Graham et al. (2007) assessed doctors' attitudes and perceptions toward clinical practice guidelines with a survey that tapped all three of these outcomes, although none of them were explicitly labeled as such: acceptability (e.g. perceived quality of and confidence in guidelines), appropriateness (e.g. perceived usefulness of guidelines), and feasibility (e.g. these guidelines provide recommendations that are implementable). Other studies interchanged the terms for acceptability and feasibility within the same article. For example, Wilkie et al. (2003) begin by describing the measurement of ''usability'' (of a computerized innovation), including its ''acceptability'' to clients but later use the findings to conclude that the innovation was feasible.
While language inconsistency is typical in most stilldeveloping fields, implementation research may be particularly susceptible to this problem. No one discipline is ''home'' to implementation research. Studies are conducted across a broad range of disciplines, published in a scattered set of journals, and consequently are rarely cross referenced. Beyond mental health, we found articles referencing these implementation outcomes in physical health, smoking cessation, cancer, and substance abuse literatures, addressing a wide variety of topics.
Clearly, the field of implementation science now has only the beginnings of a common language to characterize implementation outcomes, a situation that thwarts the conceptual and empirical advancement of the field but could be overcome by use of a common lexicon. Just as Michie et al. (2009) state the ''imperative that there be a consensual, common language'' (p. 4) to describe behavior change techniques, so is common language needed for implementation outcomes.
Referent for Rating the Outcome Several of the proposed implementation outcomes could be used to rate (1) a specific treatment; (2) the implementation strategy used to introduce that treatment into the care setting; or (3) a broad effort to implement several new treatments at once. A lingering issue for the field is whether implementation processes should be tackled and studied specifically (one new treatment) or in a more generalized way (the extent to which a system's care is evidence-based or guideline congruent). Understanding the optimal specificity of the referent for a given implementation outcome is critical for measurement. As a beginning step, researchers should report the referent for all implementation outcomes measured.
Implementation of new treatments is an inherently multilevel enterprise, involving provider behavior, care organization, and policy (Proctor et al. 2009;Raghavan et al. 2008). Implementation outcomes are important at each level of change, but the research has yet to determine which level or unit of analysis is most appropriate for particular implementation outcomes. Certain outcomes, such as acceptability, may be most appropriate for individual level analysis (for example, providers, consumers), while others, such as penetration may be more appropriate for aggregate analysis, at the level of the health care organization. Currently, very few studies reporting implementation outcomes specify the level of measurement, nor do they address issues of aggregation within or across levels.
Construct validity. The constructs reflected in Table 1 and the terms employed in our taxonomy of implementation outcomes derive largely from the research literature. Yet it is important to also understand outcome perceptions and preferences through the voice of those who design and deliver health care. Qualitative data, reflecting language used by various stakeholders as they think and talk about implementation processes, is important for validating implementation outcome constructs. Through in-depth interviews, stakeholders' cognitive representations and mental models of outcomes can be analyzed through such methods as cultural domain analysis (CDA). A ''cultural domain'' refers to a set of words, phrases, and/or concepts that link together to form a single conceptual subject (Luke 2004;Bates and Sarkar 2007), and methods for CDA, such as free-listing and pile-sorting, have been used since the 1970s (Bates and Sarkar 2007). While primarily used in anthropology, CDA is aptly suited for health services research that endeavors to understand how stakeholders conceptualize implementation outcomes, informing the generation of definitions of implementation outcomes. The actual words used by stakeholders may or may not reflect the terms used in academic literature and reflected in our proposed taxonomy (acceptability, appropriateness, feasibility, adoption, fidelity, penetration, sustainability and costs). But such research can identify the terms and distinctions that are meaningful to implementation stakeholders.
The literature reflects a wide array of approaches for measuring implementation outcomes, ranging from qualitative, quantitative survey, and record archival. Michie et al. (2007) studied perceived difficulties implementing a mental health guideline, coding respondent descriptions of implementation difficulties as 0, 0.5, or 1. Much measurement has been ''home-grown,'' with virtually no work on the psychometric properties or measurement rigor. Measurement development is needed to enhance the portability and usefulness of implementation outcomes in real world settings of care. Measures used in efficacy research will likely prove too cumbersome for real-world studies of implementation. For example, detailed assessment of fidelity through coding of encounter videotapes would be too time-intensive for a multi-agency study assessing fidelity of treatment implementation.
Research is also needed to advance our theoretical understanding of the implementation process. Empirical studies of the five issues we list here will inform theory, illuminate the ''black box'' of implementation processes, and help shape models for developing and testing implementation strategies.
Any effort to implement change in care involves a range of stakeholders, including the treatment developers who design and test the effectiveness of ESTs, policy makers who design and pay for service, administrators who shape program direction, providers and supervisors, patients/clients/consumers and their family members, and interested community members and advocates. The success of efforts to implement evidence-based treatment may rest on their congruence with the preferences and priorities of those who shape, deliver, and participate in care. Implementation outcomes may be differentially salient to various stakeholders, just as the salience of clinical outcomes varies across stakeholders (Shumway et al. 2003). For example, implementation cost may be most important to policy makers and program directors, feasibility may be most important to direct service providers, and fidelity may be most important to treatment developers. To ensure applicability of implementation outcomes across a range of settings and to maximize their external validity, all stakeholder groups and priorities should be represented in this research.
The implementation of any new treatment or service is widely recognized as a process, involving a sequence of activities, beginning with initial considerations of what and how to change current care. Chamberlain has identified ten steps for the implementation of an evidence-based treatment, Multidimensional Treatment Foster Care (MTFC), beginning with consideration of adopting MTFC and concluding when a service site meets certification criteria for delivering the treatment (Chamberlain et al. 2008). As we suggest in Table 1, certain implementation outcomes may be more important at some phases of implementation process than at other phases. For example, feasibility may be most important once organizations and providers try new treatments. Later, it may be a ''moot point,'' once the treatment-initially considered novel or unknown-has become part of normal routine.
The literature suggests that studies usually capture fidelity during initial implementation, while adoption is often assessed at 6 (Waldorff et al. 2008), 12 (Adily et al. 2004;Fischer et al. 2008), or 18 months (Cooke et al. 2001) after initial implementation. But most studies fail to specify a timeframe or are inconsistent in choice of a time point in the implementation process for measuring outcomes. Research is needed to explore these issues, particularly longitudinal studies that measure multiple implementation outcomes before, during, and after implementation of a new treatment. Such research may reveal ''leading'' and ''lagging'' indicators of implementation success. For example, if acceptability increases for several months, following which penetration increases, then we may view acceptability as a leading indicator of penetration. Leading indicators can be useful for managing the implementation process as they signal future trends.
Where leading indicators may identify future trends, lagging indicators reflect delays between when changes happen and when they can be observed. For example, sustainability may be observed only well into, or even after the implementation process. Being aware of lagging indicators of implementation success may help managers avoid over-reacting to slow change and wait for evidence of what may soon prove to be successful implementation.
Our team's observations of implementation suggest that implementation outcomes are themselves interrelated in dynamic and complex ways (Woolf 2008;Repenning 2002;Hovmand and Gillespie 2010;Klein and Knight 2005) and are likely to change throughout an agency's process to adopt and implement ESTs. For example, the perceived appropriateness, feasibility, and implementation cost associated with an intervention will likely bear on ratings of the intervention's acceptability. Acceptability, in turn, will likely affect adoption, penetration, and sustainability. Similarly, consistent with Rogers' theory of the diffusion of innovation, the ability to adopt or adapt an innovation for local use may increase its acceptability (Rogers 1995). This suggests that when providers believe they do not have to implement a treatment ''by the book'' (or with precise fidelity), they may rate the treatment as more acceptable.
Modeling the interrelationships between implementation outcomes will also inform their definitional boundaries and thus shape the taxonomy. For example, if two outcomes which we now define as distinct concepts are shown through research to always occur together, the empirical evidence would suggest that the concepts are really the same thing and should be combined. Similarly, if two of the outcomes are shown to have different empirical patterns, evidence would confirm their conceptual distinction.
Once researchers have advanced consistent, valid, and efficient measures for implementation outcomes, the field will be equipped to conduct important research treating these constructs as dependent variables, in order to identify correlates or predictors of their attainment. Their measurement will enable research to determine which features of a treatment itself or which implementation strategies help make new treatments acceptable, feasible to implement, or sustainable over time. The diffusion of innovation literature posits that the implementation outcome, adoption of an EST, is a function of such factors as perceived need to do things differently (Rogers 1995) perception of the new treatment's comparative advantage (Frambach and Schillewaert 2002;Henggeler et al. 2002) and as easy to understand (Berwick 2003). Such suppositions require empirical test using measures of implementation outcomes.
Using Implementation Outcomes to Model Implementation Success Reliable, valid measures of implementation outcomes will enable empirical testing of the success of efforts to implement new treatments, and pave the way for comparative effectiveness research on implementation strategies. In most current initiatives to move evidence-based treatments into community care settings, the success of the implementation is assumed and evaluated from data on clinical outcomes. We believe that an exclusive focus on clinical outcomes thwarts understanding the process of implementation, as well as the effects of contextual factors that must be addressed and that are captured in implementation outcomes.
Established evidence for a ''proven'' treatment does not ensure successful implementation. Implementation also requires addressing a number of important contextual factors, such as provider attitudes, professional behavior, and the service system. Constructs in the proposed taxonomy of implementation outcomes have potential to capture those provider attitudes (acceptability) and behaviors (adoption, uptake) as well as contextual factors (system penetration, appropriateness, implementation cost).
For purposes of stimulating debate and future research, we suggest that successful implementation be considered in light of a ''portfolio'' of factors, including the effectiveness of the treatment to be implemented and implementation outcomes such as included in our taxonomy. For example, implementation success (I, in the equation below) could be modeled to reflect (1) the effectiveness (E) of the treatment being implemented, plus (2) implementation factors (IO's), which heretofore have been insufficiently conceptualized, distinguished, and measured and rarely used to guide implementation decisions.
For example, in situation ''A'', an evidence-based treatment may be highly effective but given its high cost, only mildly acceptable to key stakeholders and low in sustainability. The overall potential success of implementation in this case might be modeled as follows:
Implementation success = f of effectiveness = high ð Þ + acceptability = moderate ð Þ + sustainability low ð Þ:
In situation ''B'', a given treatment might be only moderately effective but highly acceptable to stakeholders because current care is poor, the treatment is inexpensive, and current training protocols ensure high penetration through providers. This treatment's potential might be modeled in the following equation:
Thus using implementation outcomes, the success of implementation may be modeled and tested, thereby making decisions about what to implement more explicit and transparent.
To increase the success of implementation, implementation strategies need to be employed strategically. For example, implementation strategies could be employed to increase provider acceptance, improve penetration, reduce implementation costs, and achieve sustainability of the treatment being implemented. Understanding how to achieve implementation outcomes requires the kind of work now underway by Michie et al. (2009) advance a taxonomy of implementation strategies and reflect their demonstrated effects.
Open Access This article is distributed under the terms of the Creative Commons Attribution Noncommercial License which permits any noncommercial use, distribution, and reproduction in any medium, provided the original author(s) and source are credited.
The science of implementation cannot be advanced without attention to implementation outcomes. All studies of implementation should explicate and measure implementation outcomes. Given the rudimentary state of the field, we chose a narrative approach to reviewing the literature and constructing a taxonomy. Our purpose is to advance the clarity of language, provoke debate, and stimulate more systematic work toward the aims of advancing the conceptual, linguistic, and methodological clarity in the field. A taxonomy of implementation outcomes can help organize the key variables and frame research questions required to advance implementation science. Their measurement and empirical test can help specify the mechanisms and causal relationships within implementation processes and advance an evidence base around successful implementation.
Various enzyme identification protocols involving homology transfer by sequence-sequence or profile-sequence comparisons have been devised which utilise Swiss-Prot sequences associated with EC numbers as the training set. A profile HMM constructed for a particular EC number might select sequences which perform a different enzymatic function due to the presence of certain fold-specific residues which are conserved in enzymes sharing a common fold. We describe a protocol, ModEnzA (HMM-ModE Enzyme Annotation), which generates profile HMMs highly specific at a functional level as defined by the EC numbers by incorporating information from negative training sequences. We enrich the training dataset by mining sequences from the NCBI Non-Redundant database for increased sensitivity. We compare our method with other enzyme identification methods, both for assigning EC numbers to a genome as well as identifying protein sequences associated with an enzymatic activity. We report a sensitivity of 88% and specificity of 95% in identifying EC numbers and annotating enzymatic sequences from the E. coli genome which is higher than any other method. With the next-generation sequencing methods producing a huge amount of sequence data, the development and use of fully automated yet accurate protocols such as ModEnzA is warranted for rapid annotation of newly sequenced genomes and metagenomic sequences.
Emergence of next-generation sequencing technologies [1] and complete genome sequencing projects have greatly facilitated the process of unraveling the full repertoire of biological functions that an organism possesses. For a pathogenic organism, such a compilation of functions has a direct implication in identifying potential drug targets, for example, selecting genes or functions which are unique to the pathogen and not present in the host organism. Metabolic enzymes have for long been considered as a promising group from which such drug targets can be identified [2]. Enzymes form a sizable part of the druggable genome. Druggability is equated with the presence of certain folds in the proteins which can favour interactions with drug-like chemical compounds. The active sites or ligand binding pockets of enzymes are prime examples of druggable folds such that around 47% of the small molecule drugs available in the market have an enzyme target [3,4]. Enzyme drug targets have been identified and exploited in a number of pathogenic protozoans [5][6][7], bacteria such as M. tuberculosis [8,9], and fungal pathogens such as C. albicans [10][11][12]. Metagenome projects resulting from recent advances in environmental shotgun sequencing also provide opportunities for metabolic enzyme mapping as well as studying the environmental impact on evolution of metabolism [13,14].
Metabolic reconstruction is a process that aims to develop a complete overview of the metabolic capabilities of an organism from the genome or metagenome sequence. There are various databases to aid metabolic reconstruction which integrate and curate data from different sources, for example KEGG [15] which combines genomic, chemical, and network information, PUMA2 [16] which provides an array of tools for comparative genomics, and MetaCyc [17] which is a multiorganism pathway/genome database. Accurate identification of metabolic enzymes from a fully sequenced genome is a very important step towards such a reconstruction. The most common approach for detection of an enzyme function in a given genome is on the basis of sequence similarity with homologues whose function is known, which can be accomplished by using either sequencesequence comparison methods such as BLAST [18] and FASTA [19] or profile-sequence comparison methods like PSI-BLAST [20] and HMMER [21]. Different protocols have been developed using these methods that detect enzymatic function at the level of the EC number. PRIAM [22], for example, generates position-specific scoring matrices for collections of sequences which are associated with the same EC number which are then used to score sequences in a genome using reverse positionspecific blast (RPS-BLAST) [23]. MetaSHARK [24] improves on the sensitivity of PRIAM enzyme prediction by using hidden Markov models to identify corresponding enzymes directly from the genomic sequence. Other approaches for enzymatic function inference include checking for the presence of a conserved pattern or motif and identification of functionally critical residues. EFICAz [25] is a protocol which combines these two approaches along with an iterative HMMER-based procedure for generating multiple alignments for EC number families and pairwise sequence comparison using a familyspecific identity threshold. It has recently been extended by including additional components based on support vector machine (SVM) models [26].
Metabolic enzymes belonging to certain core pathways are conserved in all three domains of life (namely archaea, bacteria, and eukaryotes) and can thus be easily identified using sequence homology [27]. However, there are a total of 4905 unique EC number entries in the enzyme [28] database release of 19 January, 2010, out of which only 2507 entries have one or more sequence associated with them. This has led to development of various algorithms and methods which try to associate sequences with enzymatic activities hitherto unannotated in an organism (hence, a pathway "hole" [29]) and can be collectively termed as "hole-filling" algorithms. A variety of other information is used in addition to the sequences, for example topology of the metabolic network [30,31] or genomic evidences such as chromosomal clustering of operons [32], a combination of both [29], or an ensemble of various kinds of methods such as gene coexpression, phylogenetic profile cooccurence, protein fusion, and so forth [33,34]. Whereas these methods can be used in conjunction or as a complement to the profile-sequence comparison methods to obtain a complete overview of the metabolic reactions of an organism, the latter group of methods, nevertheless, still retains its importance.
A profile HMM constructed using Swissprot sequences for a particular EC number might score sequences belonging to other EC groups very highly if the enzymatic activities have developed in the same protein fold and therefore share certain fold-specific residues [35]. For example, the Alpha/Beta hydrolase fold (SCOP [36] classification) has 35 protein families which span through various enzymatic functions in terms of EC numbers. A group of 9 EC numbers, all of which are carboxylic ester hydrolases (EC 3.1.1) with different substrate specificities, are included in this fold group (Supplementary Table ST1 available online at doi:10.1155/2011/743782) and would be expected to have common fold-specific signals which would be conserved in sequences belonging to these EC numbers. As a result, the profile HMM for the EC 3.1.1.8 used with default parameters selects sequences from 5 other EC groups (Supplementry Figure SF1, inset available online at doi:10.1155/2011/743782). Figure SF1 also depicts the overall structural similarity between representatives from each of the six EC groups.
We have earlier described the use of negative training sequences (i.e., sequences of different functions related to the training sequences by virtue of sharing a common fold) to both optimise the discrimination threshold as well as modify the emission probabilities of the profile HMMs to increase its specificity. We have used relative entropy of the amino acid probabilities of the positive and negative training sequences to select residues in the positive alignment which are responsible for its specific function as opposed to residues which are similarly conserved in both the negative and positive training sequence sets [35]. In this paper, we describe ModEnzA (HMM-ModE Enzyme Annotation), where we apply HMM-ModE to create profile HMMs which are specific at the functional level as defined by the EC classification. We enrich the training set by mining sequences from the nonredundant (NR) database for increased coverage of the EC numbers and thus, increased sensitivity. We use the Markov clustering algorithm (MCL) [37] to partition the EC sequence sets into clusters corresponding to nonorthologous sequences or oligomeric subunits. The ModEnzA protocol is used to annotate metabolic enzymes from completely sequenced reference genomes. We present a comparative analysis of our protocol with other methods used for genome-wide enzyme identification such as PRIAM, MetaShark, and EFICAz.
The expasy enzyme database (release of 19 Jan 2010) had a total of 4905 unique EC number entries out of which 2507 were associated with a total of 180315 sequences. These 2507 entries were divided into 2 groups, those having 3 or more sequences and those having just 1 or 2 sequences. The 1910 EC numbers which had 3 or more sequences in Swiss-Prot were designated as Tier I (Figure 1) while the sequences in the latter group were used as queries to mine similar sequences
Scan the proteome of organisms List of Enzymes Apply filters to select similar sequences with unambiguous annotations Less than 3 sequences Less than 3 sequences Discard EC More than 3 sequences with original query as Best hit (reciprocal test) Tier I Tier II Tier III MCL clustering Construct HMM-ModE profiles Training sequences from Swiss-Prot for each EC number BLASTP with the non redundant database (slow, sensitive search) from the nonredundant protein database (NR) database as follows.
There were 597 EC numbers that had just 1 or 2 sequences associated with them in the Swiss-Prot database. These were used as BLASTp queries to mine similar sequences from the NR database at NCBI (http://www.ncbi.nlm.nih.gov/sites/ entrez?db=Protein). We used an E-value cut-off of 10 -35 , a percent identity range of 50 < × < 99 and a query coverage of 80% as the primary filters. If a sequence appeared as BLAST hit in more than one EC numbers, it was removed from all the EC number files. We removed sequences with ambiguous annotations such as "hypothetical", "unnamed", "unknown", "unclassified", and "unidentified". While we did not follow a systematic testing approach to fix these parameters, we nevertheless tried various combinations of the parameters (e-values of 10 -30 and 10 -35 , percent identity ranges 50-99, 75-99, 75-95, etc). The final set of parameters was decided upon after visual inspection of the annotations of the gathered sequences looking for the least amount sequences with ambiguous sequence descriptions and the most number of EC groups with more than 3 sequences.
After the filtering step, we were left with 450 EC number groups with 3 or more sequences. To these we further applied a reciprocal BLAST best hit criterion, wherein we discarded those NR sequences which did not have the original query swissprot sequence as their top hit. We obtained 364 EC groups having 3 or more sequences which fit this criterion. These were designated as Tier II profiles (Figure 1). The remaining 86 EC groups which also had 3 or more sequences but did not have sufficient reciprocal best hits were designated as Tier III (Figure 1). [37]. The sequences belonging to each EC number in Tier I were clustered into subgroups with MCL using pairwise BLAST scores as input. This resulted in 2313 distinct subgroups having 3 or more sequences. These were considered as separate profiles called as Tier I profiles, for example the 7 subgroups of EC number 3.1.1.4 were designated as 3.1.1.4 1 through 3.1.1.4 7 and converted into separate HMMs. There were 117 EC numbers which had subgroups containing just 1 or 2 sequences. These subgroups were discarded after manual inspection. Similarly, the sequences in the other tiers were also clustered using MCL. We obtained 370 subgroups with more than three sequences in Tier II and 86 subgroups in Tier III making a total of 2769 distinct ModEnzA profiles.
The subgroups for each EC number were used to construct HMM profiles and HMM-ModE profiles as previously described [35]. Briefly, HMMs are generated for each subgroup using hmmbuild from the HMMER package [21] after aligning the sequences using MUSCLE [38]. The discrimination threshold is optimised by performing an n-fold cross-validation routine, partitioning the training sequences into n train and test sets such that each sequence is part of at least one test set. For each test set t, a profile HMM created from the remaining (n -1) sets is used to score the sequences to get a True Positive (TP) score distribution. False positives (FP) are identified from the Swiss-Prot sequences (i.e those sequences that perform an enzymatic function which is different from that of the EC number for which the HMM was generated) using hmmsearch from HMMER. The FPs are also partitioned into n sets such that each FP sequence is part of at least one set. The profile HMM created for each (n -1) subset is also used to score the corresponding FP set to get an FP score distribution. The sensitivity, specificity, and Matthews correlation coefficient (MCC) [39] distributions for each of n sets is calculated using these TP and FP score distributions. The optimal discrimination threshold is identified as the mid-point of the MCC distribution averaged over the n sets.
The number of cross-validation sets n is defined as follows:
If more than 3 false positives are identified, then these are again aligned and converted into HMMs and used to modify the true positive profile HMM as described earlier [35]. In case of a large number of false positives from the preclassified training set, to avoid issues of multiple alignments of very large datasets, we restrict the number of false positives to 200. The false positives are first clustered using MCL, and then sequences are randomly selected from the clusters proportionate to the size of the clusters such that the final number is 200. An optimized threshold is then calculated as above. In case no false positives are selected by the original HMM, it is used with default parameters.
All methods described above were automated using scripts written in-house, to form the workflow described in Figure 1, and are available from the authors upon request.
For the bacterial genomes (E. coli, B. aphidicola, and M. pneumoniae) we used the corresponding HAMAP [40] annotations as the benchmark for comparing the various enzyme identification methods. The genome sequences for these and the corresponding annotations were downloaded from http://expasy.org/sprot/hamap/bacteria.html. We used hmmsearch to score the protein sequences for a given genome with the ModEnzA profiles generated above.
The sequence IDs identified by each of the methods which had a 4-digit EC number annotation in HAMAP were considered as true positives (TP). False negatives (FN) were those predicted sequence IDs which had EC annotations in HAMAP but were not selected by a method whereas false positives (FP) were predicted sequences which actually did not have an EC number annotation in HAMAP. Similarly, for assigning EC numbers, true positives were those predicted EC numbers which had corresponding sequences in the HAMAP genome annotations, false positives, those that did not and false negatives were EC numbers which were present in HAMAP but not predicted by a method. For the P. falciparum genome, we used EC annotations in PlasmoDB as the benchmark. The sequence annotation and EC number information for P. falciparum can be obtained from http:// plasmodb.org/plasmo/. True positives, false positives, and false negatives were defined as earlier. For the ROC curves for P. falciparum we used the KEGG annotations as the benchmark.
The PRIAM program [22] and the corresponding profiles were downloaded from http://priam.prabi.fr/REL JUL06/ index jul06.html and used with an RPS-BLAST e-value cutoff of 10 -30 on the protein sequences of the genomes. The MetaShark package [24] was obtained from http:// bmbpcu36.leeds.ac.uk/shark/ This was run using an e-value cutoff of 10 -30 on the genomic DNA sequences of the organisms (including plasmids in case of B. aphidicola) obtained from NCBI ftp://ftp.ncbi.nih.gov/ The EFICAz enzyme annotations [25] for various organisms were obtained from http://cssb2.biology.gatech.edu/EFICAz/.
The percentage sensitivity and specificity of each method was calculated as
* 100,
2.5. Receiver-Operator Characteristic Curves. For comparison of ModEnzA with PRIAM and MetaShark the ModEnzA profiles for all EC numbers were rebuilt using the July 2006 version of ENZYME database which is used by the current versions of both PRIAM and MetaShark. The ROC curves (1-specificity versus sensitivity) were plotted for each of the three methods (PRIAM, MetaShark, and ModEnzA) on the four genomes by comparing against the benchmark sets mentioned above. The three programs were run on the genomes using an E-value threshold (RPS-BLAST, PSI-BLAST, and hmmsearch E-values for PRIAM, MetaShark, and ModEnzA, resp.) of 10. For PRIAM and MetaShark, the resulting E-values were binned and a threshold sweep was used to calculate the sensitivity and specificity at each E-value threshold. For ModEnzA, similar calculations were performed using a sweep through the hmmsearch scores (Figure 2(a)). The MetaShark output consists only of EC numbers, hence its sensitivity and specificity values were calculated only on the basis of EC number comparisons, whereas for PRIAM and ModEnzA, both the correct EC 0 0.2 0.4 0.6 0.8 1 -0.1 0 0.1 0.2 0.3 0.4 0.5 Sensitivity 1-Specificity Priam ModEnzA-RT MetaShark 0 0.2 0.4 0.6 0.8 1 -0.1 0 0.1 0.2 0.3 0.4 0.5 Sensitivity 1-Specificity 0 0.2 0.4 0.6 0.8 1 -0.1 0 0.1 0.2 0.3 0.4 0.5 Sensitivity 1-Specificity 0 0.2 0.4 0.6 0.8 1 -0.1 0 0.1 0.2 0.3 0.4 0.5 Sensitivity 1-Specificity Priam ModEnzA-RT MetaShark B. aphidicola E. coli M. pneumoniae P. falciparum Advances in Bioinformatics numbers predicted as well as the correct sequences assigned were used.
One of the crippling issues in biological data modeling is the incomplete nature of the training data. We have used both Swiss-Prot [41] and NR sequences for generating profiles specific for EC numbers. There were 1910 EC numbers which had 3 or more Swiss-Prot sequences associated with them. These form our Tier I profiles, which can be used to annotate sequences with a very high degree of confidence, since they were generated from curated training sequences. We considered three sequences as the minimum number required to generate a multiple alignment and consequently for constructing HMMs. Even though these profiles with 3 sequences (and in general all profiles with very few training sequences) are not expected to be representative of the protein function, they are, nevertheless, included in the ModEnzA protocol as place holders which would become more accurate as they are populated with more sequences in future.
The training set for EC numbers which had less than 3 Swiss-Prot sequences was enriched by mining sequences from the NR database using BLASTp with parameters suitably modified for increased sensitivity (namely, BLOSUM62 substitution matrix and a lower than default value for " f ", the hit extension parameter). The hits were screened with a stringent E-value cut-off of 10 -35 . Hits having more than 99% identity or those having less than 50% identity with the query sequence were left out to remove identical sequences and potential false positives while ensuring variability in the protein family. This resulted in 450 EC groups (which had more than three sequences) out of which in 364 EC groups (Tier II), all the mined sequences picked up the original swissprot query sequence as the reciprocal best BLAST hit. The remaining 86 were put in a separate group designated as Tier III. By enriching the training data set, we have populated an additional 450 enzyme functions with highquality sequences, providing increasing coverage over the set of enzyme functions. This is still not comprehensivethe coverage that our profiles provide, including Tier II and III, is still only 48% of the total number of enzyme functions known. However, the same method developed for known functions may be applied to new functions as they are mapped with sequences in future versions, of which this enrichment exercise serves as an example. The coverage within an organism is expected to be much higher, as most common functions are mapped with known sequences.
Sequences belonging to the same EC group and thus performing the same enzymatic function, might have complex relationships amongst them in terms of sequence similarity due to the presence of (1) heteromeric multiple subunits of the enzymes, (2) nonorthologous or unrelated sequences performing the same function, or (3) sequences that can perform more than one enzymatic function [22]. Clustering the sequences for a particular EC number based on a similarity score is therefore an important requirement to separate these multiple subunits or nonorthologous sequences.
We used MCL to cluster the sequences of an EC group into subgroups after removing sequences which had "fragment" as part of the annotation. To validate our use of automated clustering, we compared our MCL clusters for some of the cases mentioned above with available structural information from PDB [42]. For example, MCL clusters the swissprot sequences for the DNA-directed RNA polymerase (EC 2.7.7.6) into 14 subgroups (PRIAM has 51 separate profiles for EC 2.7.7.6). We tested the ModEnzA profiles for EC 2.7.7.6 on a set of sequences of the structural subunits for the RNA polmerase downloaded from PDB. The yeast polymerase (PDB ID: 3H0G) has 12 subunits while the prokaryotic polymerase from E. coli (PDB ID: 3LU0) has 5 subunits (excluding the sigma subunit). In each case, the sequences corresponding to different subunits were picked up by a different ModEnzA subgroup profile (Supplementary Table ST2 available online at doi:10.1155/2011/743782). We also manually inspected the "singlet" sequences, that is the sequences which the MCL clustering procedure cannot assign to any cluster bigger than size 3 and thus have to be discarded while making the profiles. We obtained 2313 distinct MCL subgroups with more than 3 sequences in Tier I. There were 117 EC numbers (with a total of 215 "singlet" sequences) with subgroups having less than 3 sequences. These were subjected to a BLASTp analysis against the NR database and the results were manually inspected to ascertain the functional neighbours (top BLAST hits) of these sequences. Out of the 215 sequences that were so tested, 23% (55) matched with sequences which had a different functional annotation than the original EC group suggesting that these might have had possible errors in annotation. Some of the sequences (14 or 6%) were shorter than most of the other sequences in that EC group (probably fragments or truncated proteins) whereas a fraction (3%) were longer than the rest. Another 15 sequences (6%) were not clustered because they belonged to a different sub-unit of the enzyme, indicating the ability of MCL to clearly distinguish between heteromeric subunits of enzymes. Whereas 23% sequences did not show any significant hits in the BLAST search, a sizable fraction (30%) were such that they had the correct annotation but not a sufficient number of neighbours to form a separate cluster. These were not included in the profile construction; however, they could be used for mining similar sequences from the NR database if they match sequences with the same function as the EC group. As mentioned above, all of the 215 sequences were not included in the profile generation step. This exercise demonstrates the efficacy of using an MCL-based clustering approach in separating enzyme subgroups.
The PRIAM method identifies the longest homologous subsequences shared within the set of sequences associated with EC number (EC group) as a single module or domain. It initially takes the shortest sequence in the group as a module and then proceeds to identify similar subsequences within a given EC group using PSI-BLAST. The matching subsequences are removed from the corresponding sequences and the shortest sequence identified to start the next iteration [22]. Complete linkage clustering using a 30% sequence identity cutoff has also been employed to create subgroups as in the EFICAz protocol [25,26]. The most populated and most divergent subgroup is then converted into a profile HMM which is used to add sequences with E-value < 0.01 which have at least one conserved potential active site residue. These subgroups are termed as CHIEFc (conservation-controlled HMM iterative procedure for enzyme family classification) families [25]. The Markov clustering algorithm (MCL) provides for an accurate, unsupervised, and fully automated protocol for clustering the sequences of a given EC number. MCL avoids potential problems associated with module architecturebased clustering as well as the pairwise protocols of dealing with similarity relationships which have been used earlier [43]. It also circumvents the requirement of using arbitrary sequence identity cutoffs as described above or manual assignment. Instead, the graph-theoretic representation of the similarity scores allows for the detection of global patterns of sequence similarity in a single step [37].
The training sequences for each EC number were used to generate HMM profiles using the HMM-ModE protocol as described in Methods. The average sensitivity and specificity distributions of the n sets in the n-fold cross validation exercise provide us with a confidence measure with which to use the profiles. If the original profile HMM selects false positives from within the negative training sequences, the information from these is used to modify the emission probabilities of the HMM so that it becomes more specific.
The ModEnzA profiles were used to scan the complete genomes of organisms to assess the performance of our method. Following the example of [22,24], we chose three bacterial genomes (E. coli, B. aphidicola, and M. pneumoniae) which have been extensively annotated using both manual and automatic methods as part of the high-quality automated and manual annotation of microbial proteomes project [40] and one eukaryotic genome (P. falciparum), for which detailed descriptions of function are available from the PlasmoDB resource [44]. The enzymatic function annotations from these sources were used as benchmarks for comparing the sensitivity and specificity of ModEnzA with other enzyme identification methods PRIAM, MetaSHARK, and EFICAz (see Section 2).
Identification of functional residues or residues that are responsible for imparting functional specificity is a critical step toward constructing function specific profiles. Information theoretic measures such as positional entropy [45] and mutual information [46] have earlier been used for this purpose. EFICAz employs an evolutionary footprinting approach which scores each position in an alignment of a family of sequences by a combination of entropy based conservation scores in (1) the "homofunctional" alignment, that is, consisting of members of the same family and (2) a "heterofunctional" alignment which consists of similar sequences (which might have different functions), mined from a nonredundant database using the homofunctional profile HMM. The alignment positions are then ranked using a Z-score of the conservation degree to identify functionally discriminating residues (FDRs). HMM-ModE, the protocol upon which ModEnzA is based, uses a similar concept of negative training sequences, these are sequences belonging to other functional families which are scored positively by the HMM for a particular family. We use a position-dependent null model which contains conservation information from these negative training sequences. We calculate the relative entropy between the distributions of amino acids in the alignments of the positive and negative training sequences and select residues where this score is higher than the relative entropy between the positive sequences and the null set (i.e., the null probabilities as calculated from the swissprot sequences). The probabilities of the amino acids in the negative set are used to modify the emission probabilities of these selected residues to generate a function specific profile HMM. This avoids the problem of loss of information associated with using a Z-score cutoff [35]. The discrimination potential of a profile is a function of the unique nature of the family at the level of both its fold and specificity determining positions. As a consequence, the discriminating potential of a profile HMM at a functional level would be expected to be dependent on the frequency of occurrence of the parent fold among proteins, that is, an activity that arises in a commonly occurring fold will be more likely to have false positives from the complete training set than an activity that arises in a fold that is not so wide-spread. As such, there is no discernible relationship between the cluster size (no. of training sequences available for a subgroup) and the discrimination potential of the corresponding profile (Supplementary Figure SF2 available online at doi:10.1155/2011/743782). EFICAz uses a family specific Sequence Identity threshold (SIT), which is set after making all pairwise sequence comparisons within the family, for predicting enzyme function [25]. Both PRIAM and MetaShark have options for varying the e-value threshold. An E-value cutoff of 10 -10 in PRIAM results in a high sensitivity (92.61%) but a specificity of only 82.25% for identifying EC numbers from E. coli. On the other hand, using a stringent E-value cutoff of 10 -30 increases specificity (90.42%) with a drop in sensitivity (90.80%) (data not shown). This is a drawback of using arbitrary thresholds for discrimination. ModEnzA uses cross-validation to ensure an optimal threshold for each separate profile [32] thereby eliminating the use of arbitrary thresholds.
The EFICAz webserver at http://cssb2.biology.gatech.edu /cgi-bin/eficaz browse.cgi provides 4 digit EC number predictions for a number of genomes. These predictions were directly compared with ModEnzA. The sensitivity and specificity were calculated with respect to the EC numbers that were correctly assigned to a genome as well as the sequences associated with an enzymatic activity which were identified by the methods. The comparative performance of EFICAz and ModEnzA for identifying EC numbers and enzymatic sequences from the four genomes is shown in Table 1. The sensitivity and specificity values of the ModEnzA profiles are higher compared to EFICAz both in assigning EC numbers Advances in Bioinformatics as well as identifying enzymatic sequences. As expected, the sensitivity is improved by the inclusion of the Tier II and Tier III profiles (the sensitivity in assigning EC-associated sequences increasing up to 97% for B. aphidicola, e.g.) but there is a slight drop in the specificity because the training sequences were mined from the NR database and hence may have errors of annotation. The Tier I ModEnzA profiles by themselves show a higher specificity than EFICAz while selecting enzymatic sequences for each genome tested. The augmentation by Tier II and III profiles still maintains an on par specificity while increasing the sensitivity to a value greater than that of EFICAz for identifying both EC numbers and sequences. Given that the genome databases mentioned above also have an automated annotation component to them, it must be noted that choosing a different set of annotations as the benchmark (e.g KEGG) does not significantly alter the conclusions of the comparisons (supplementary Table ST3 available online at doi:10.1155/2011/743782).
The PRIAM and MetaShark methods use the July 2006 version of the ENZYME database as the training set. To ensure a fair comparison we rebuilt all the ModEnzA profiles using the July 2006 ENZYME version. As demonstrated in the next section, the KEGG database has more EC-sequence associations for P. falciparum than PlasmoDB, and for this reason we replaced PlasmoDB with KEGG as the benchmark for the comparison of ModEnzA with PRIAM and MetaShark. Receiver-operator characteristic (ROC) curves serve as an indicator for the discriminating potential of a classifier method. The ROC curves for the retrained ModEnzA (ModEnzA-RT) profiles as well as PRIAM and MetaShark are shown in Figure 2(a). Since EFICAz uses 4 different methods, each with its own set of parameters, we could not include EFICAz in this analysis. The area under the ROC curve for the ModEnzA profiles is better than both PRIAM and MetaShark for all the four genomes. MetaShark predictions on P. falciparum suffer from very low specificity (Figure 2(a), bottom left panel) probably because it is difficult to map the profile HMMs back onto its DNA sequence which is atypically AT rich and contains lots of repetitive sequences [47]. The ROC curves for ModEnzA scans on the entire genomes of the four organisms (Figure 2(b)) show the high discrimination potential of the ModEnzA profiles.
Eukaryotic species account for only about 16% of the most represented species (31% of all sequences) in terms of number of sequence entries in Swiss-Prot (http://ca.expasy. org/sprot/relnotes/relstat.html). The P. falciparum genome is also not included in the HAMAP project which deals extensively with bacterial, archaeal, and plastid encoded proteins. Hence, this could be considered as an ideal case to compare the various enzyme identification methods. We chose the PlasmoDB enzyme annotations as a benchmark to decide on the true positive, false positive, and false negative predictions. The sensitivity and specificity calculated in terms of these numbers for ModEnzA and EFICAz is shown in Table 1. The relative scarcity of training sequences from the eukaryotic domain is reflected in the low sensitivities of the methods. Again the specificity of the ModEnzA profiles is higher than EFICAz. ModEnzA assigned 22 EC numbers which were not annotated in PlasmoDB (False Positives). We checked for the corresponding functions of these EC numbers in two other databases, namely KEGG (because it is often used as a reference knowledge base) [12] and PlasmoCyc [4] (which also contains a comprehensive annotation of the P. falciparum genome). We found that 13 of the 22 false positive EC numbers have been annotated as belonging to P. falciparum by either of these databases (Table 2).
We were interested in the ability of Tier II and Tier III profiles to annotate novel sequences, especially as they were created from sequence similarity and not expert-curated data sets. It was also interesting to address the fact that in the absence of a profile, ModEnzA could be used to pick a related function. The Tier II ModEnzA profiles selected 14 sequences and 7 EC numbers, respectively, from the P. falciparum genome (Table 3). The cysteine protease falcipain sequences (gene IDs PF11 0161, PF11 0162, and PF11 0165), for instance, do not currently have an EC number associated with them. So in absence of a corresponding EC profile, ModEnzA annotates it with the nearest cysteine endopetidase bromelain (EC 3.4.22.32) but it is gratifying to note that the annotation is correct upto the general level of the first three E.C. digits. Three of the 7 EC numbers predicted by ModEnzA exactly match the corresponding PlasmoDB assignments while 3 others have the same first three digits. Only one EC assignment (EC 3.4.23.2) has more than 1 digit mismatch with the EC annotated in PlasmoDB. Of the 14 sequences annotated, the EC assignments for 8 share the first three digits with the corresponding annotations in PlasmoDB (Table 3). The Tier II and Tier III profiles can annotate sequences up to the first three EC digits with sufficient accuracy. However, as has been mentioned earlier, these should be used with caution, because the training sequences used for these profiles may be prone to annotation errors.
As the results show, there is a discrepancy between the annotations/predictions of different databases. The Mod-EnzA protocol assumes importance because it is a rapid tool that provides a high degree of confidence in assigning EC numbers to a genome. A typical hmmsearch with ModEnzA Tier I profiles on the E. coli genome (4407 proteins) takes ∼7.8 Hrs on a laptop having a 2.4 GHz core2duo Intel processor and 3 Gb RAM. The same search on a workstation with a 2.50 GHz Intel xeon quadcore processor with 8 GB RAM takes ∼2.2 Hrs. Since the hmmsearch and hmmscan program is inherently capable of multithreading, ModEnzA can be expected to be even faster on machines with more processors. For example, we were able to annotate the enzymes in a metagenomic sample with 203240 translated protein fragments with ModEnzA in just around 1.3 hrs by splitting the target sequences into 10 parts and using 10 nodes (each with a dual core processor) to run the hmmscan program on a high-performance computing cluster (unpublished results).
The n-fold cross validation routine built into ModEnzA on the training sequences ensures an optimal threshold which can separate the true positives from the False positives for any given EC number. The modification of emission probabilities of the True positive profiles by using information from the false positive alignment further increases the specificity. We have decided to make this data available for use by the scientific community using HMMER2, even though HMMER3 has since been released [21].
HMMER3 has only local-local alignments, and our method is based on predicting the fold (domain) and hence is implicitly based on global or "Glocal" (align a complete model to a subsequence of the target) alignments. ModEnzA is based on scripts that modify the emission probabilities in the model, and it is unsure if the formats for HMMER3 are stable enough to extract probabilities from the model. The model has changed between the two versions, and it has been advised that profiles built on one should not be used with the other, as parameters are differently optimised. When more alignment modes are available and the format is more stable, newer versions of ModEnzA would migrate to HMMER3 to take advantage of the increased sensitivities and speed.
We present a method for enzyme annotation by enriching existing curated databases and using profile hidden Markov models optimised for specificity using negative training sequences. The protocol shows improved sensitivity and specificity compared to other existing methods for enzyme identification and can be used to accurately map the metabolome of an organism.
"-"-Annotation not present in either PlasmoCyc or KEGG.
* Gene product descriptions and EC annotations obtained from PlasmoDB. # IUBMB EC description.
Musical education has a beneficial effect on higher cognitive functions, but questions arise whether associations between music lessons and cognitive abilities are specific to a domain or general. We tested 194 boys in grade 3 by measuring reading and spelling performance, non verbal intelligence and asked parents about musical activities since preschool. Questionnaire data showed that 53% of the boys had learned to play a musical instrument. intelligence was higher for boys playing an instrument (p < .001). to control for unspecific effects we excluded families without instruments. the effect on intelligence remained (p < .05). Furthermore, boys playing an instrument showed better performance in spelling compared to the boys who were not playing, despite family members with instruments (p < .01). this effect was observed independently of iQ. our findings suggest an association between music education and general cognitive ability as well as a specific language link.
Active music performance relies on a demanding action-perceptionloop calling for long periods of focused attention on dynamic visual, auditory, and motor signals. Given this extra training of high-level cognitive skills in children who learn to play an instrument, it can be asked whether making music enhances children's performance in domains other than music.
Positive relationships between playing an instrument and general cognitive abilities have been observed previously. In a retrospective design with 6-to 11-year-old children, Schellenberg (2006) found a correlation between the duration of music lessons and performance in an verbal and non-verbal IQ test as well as school performance. The effects on IQ and on academic performance were still observable in undergraduates that had been trained to play an instrument in childhood. Forgeard, Winner, Norton, and Schlaug (2008) observed a relationship between playing an instrument and higher cognitive functions in a sample of forty-one 8-to 11-year-old children who had at least 3 years of musical instruction. Beside motor learning and enhanced melodic discrimination, the authors also found enhanced vocabulary and nonverbal reasoning scores.
However, no differences were found in a prospective study investigating 6-year old children between a group of 16 control children and 15 children who had weekly private keyboard lessons for 15 months (Hyde et al., 2009). Nevertheless, the authors were able to show neartransfer effects (motor and auditory skills) as well as structural brain changes for the keyboard group.
In an experimental design, Schellenberg (2004) reported an effect on IQ using Wechsler's WISC-III in 6-year-olds after keyboard or singing lessons for 36 weeks. The music group (+ 7.0 points) showed a larger increase than the control group taking drama lessons in the same time or waiting for piano lessons (+ 4.3 points). This finding contrasts the meta-analysis of Hetland (2000, Analysis 2) including five experimental studies about the effect of musical training on Raven's IQ.
Besides this broad effect of music on general cognitive performance, some studies also found associations with mathematical (Cheek & Smith, 1999;Vaughn, 2000) and spatial abilities (Hetland, 2000; Analysis 1).
Moreover, there seems to be a link between musical training and language abilities since musical training in childhood influences the development of auditory processing in the cortex (Fujioka, Ross, Kakigi, Pantev, & Trainor, 2006;Moreno & Besson, 2006). There is evidence that musical training is linked to language related aspects such as pitch processing (Moreno et al., 2009;Schön, Magne, & Besson, 2004;Wong, Skoe, Russo, Dees, & Kraus, 2007), speech prosody (Thompson, Schellenberg, & Husain, 2004), verbal memory (Chan, Ho, & Cheung, 1998;Ho, Cheung, & Chan, 2003;Jakobson, Cuddy, & Kilgour, 2003;Kilgour, Jakobson, & Cuddy, 2000). Additionally, musical aptitude was found to correlate with second language acquisition (Slevc & Miyake, 2006). Furthermore, associations of musical training and reading performance have been demonstrated in a normal population (Barwick, Valentine, West, & Wilding, 1989;Butzlaff, 2000;Lamb & Gregory, 1993) as well as in dyslexics (Overy, 2003).
The putative link between musical and language abilities is seen in the discrimination of rapid auditory events (Jakobson et al., 2003;Tallal & Gaab, 2006). Musical instrument training should improve auditory information processing, which in turn is crucial for the acquisition of reading and writing skills.
It is no longer the question whether or not musical training is associated with higher cognitive abilities, because there is growing evidence that it is. An unresolved issue however, is the nature and specificity of the link (Schellenberg & Peretz, 2008). It has been proposed that all specific relations observed so far can be explained by a carry-over effect of the relation between musical training and general abilities as measured by IQ (Schellenberg & Peretz, 2008). Indeed, such a dependency was always found in Schellenberg's studies. Most of the previous studies showing a relation between musical training and specific abilities, such as language performance, did not measure general abilities. Therefore these studies could not report on the dependency of both.
Our correlational study addresses this unresolved issue of linkspecificity by looking at a general association as well as at a specific language association of musical training.
We recruited 272 elementary school boys of Grade 3 aged 8 to 9 years from 26 schools in a southern German school district. The recruitment served two purposes. On the one hand, the boys were screened for an electroencephalographic study on auditory processing in normal and dyslexic children (Gust, 2009). Therefore we included only healthy boys who were native German speakers and had not repeated a class.
On the other hand all screening data was used in combination with an additional parents' questionnaire to answer the research question presented here.
The study followed the principles of the declaration of Helsinki and was approved by the local internal review board of the Medical Faculty, University of Ulm.
We tested non-verbal intelligence with the German adaptation of Cattells Cultural Fair Intelligence Test -Scale 1 (CFT-1; Cattell, Weiß, & Osterland, 1997). The CFT-1 consists of five subtests (substitutions, labyrinths, classification, similarities, and matrices) and takes about 45 min to complete. This non-verbal IQ test was chosen to measure intelligence independently from progress in reading and writing.
Reading and spelling performance was tested with the Salzburger Lese-und Rechtschreibtest (SLRT; Landerl, Wimmer, & Moser, 1997).
The SLRT is an individually given test assessing reading accuracy and reading speed for three word and two non-word reading subtests as well as spelling performance with regard to different types of spelling errors.
Parents filled out a questionnaire about the musical experience of their child during preschool and school years, including singing, listening to music, and playing an instrument, either at home or in an institutional setting such as children choir and music school. Additionally, we asked questions about the parental encouragement concerning non-musical activities. It was rated on a scale from 1 (never) to 7 (more than once daily) how often adults engaged with the boys in activities like looking together at picture books, reading books to the boys, telling stories to the boys, encouraging boys to draw and paint, or being at the playground with them. A composite score of "parental investment" was calculated from these ratings.
Lastly, parents were asked if any family member is playing an instrument. We expect that boys who play an instrument differ from boys that do not play an instrument. The existence of family members who play instruments allows to control for any unspecific differences, such as the family value of playing an instrument, or the minimum family income to allow for financing an instrument and lessons.
Descriptive and inferential statistics were computed using STATISTICA 7.1 (StatSoft, Inc. Tusla, OK, USA).
Two hundred and six parents completely answered and sent back the questionnaire on the musical experience of their boys (76% return rate). Table 1 summarizes the overall musical experience of the boys.
One quarter had experience in singing in a choir, and half of the boys learned playing an instrument or did so in the past.
Table 2 provides a breakdown of the boys who learned an instrument according to the age at which boys started musical instrument training. A complete data set on non-verbal IQ, spelling and reading with the SLRT resp. was available for 194 of the 206 boys whose parents returned the questionnaire on musical experience.
In our sample intelligence showed a normal distribution with a mean of 104.5, a standard deviation of 13.6, a minimum of 72 and a maximum of 142. The non-verbal IQ was higher for boys playing an instrument, t(192) = 3.45, p <. 001 (see Figure 1). The size of the effect (d = 0.50) was at a medium level (Cohen, 1988). To control for differences in family values and family income boys who lived in families without musical instruments (n = 58) were excluded. The effect on intelligence remained, t(134) = 2.40, p < .02, d = 0.46 (Figure 1), when we compared the boys not playing (from families with members playing musical instruments) with the boys playing an instrument themselves.
No difference in non-verbal IQ was found between boys who have sung in a choir and those who did not, t(192) = 1.53, p = .127. For boys who took part in a course on "First Experiences With Music" a higher non-verbal IQ was found, t(192) = 2.76, p < .01, d = 0.41. However, when families without musical instruments were excluded, this difference disappeared, t(134) = 1.75, p = .083. "Parental investment" correlated weakly with non-verbal IQ, n = 183, r =.156, p < .05.
Spelling performance was better for boys playing an instrument as measured by the spelling mistakes made in the SLRT, t(192) = 4.22, p < .0001, d = 0.60. This effect remained after excluding the families without instruments, t(134) = 2.78, p < .01, d = 0.51.
A weak correlation between spelling mistakes and non-verbal IQ (r = -.17, p < .05) was found in our sample: The more intelligent the students the fewer spelling mistakes they made. To eliminate the effect of non-verbal IQ an ANCOVA was performed that confirmed the relationship between playing an instrument and spelling independently of non-verbal IQ, F(1, 191) = 13.96, p < .001; also after families without instruments were excluded, F(1, 133) = 5.36, p < .05.
Reading performance was accessed by reading speed and by reading mistakes as measured by the SLRT. Only for the reading time the boys who play an instrument showed an advantage, t(192) = 2.02, p < .05, d = 0.29; but this better performance disappeared when families without musical instruments were excluded, t(134) = 0.53, p = .60, d = 0.09.
The other variables (singing in a choir, taking part in a course on "First Experiences With Music", "Parental Investment") were not associated with reading or spelling performance.
The results so far described were obtained from the whole group of boys. In the following analysis we focus on low-performer in terms of spelling. Low performers were defined as the quarter of boys (n = 51) Note. In the course "First Experiences With Music" the boys were trained to listen, to sing and dance together, and to play on instruments such as glockenspiel and woodblock.
tAble 2.
Note. Instrument types were recorder (n = 56), piano or keyboard (n = 33), guitar (n = 11), drum set, drum, trumpet, French horn, saxophone, accordion, melodica, baritone horn, violoncello, glockenspiel, xylophone. with the highest number of SLRT spelling mistakes. Not all of these low-performers showed a spelling deficit as defined by the ICD-10.
Only 17 of them exhibited a spelling performance on or below 10% of the population accompanied by an IQ within the normal range.
Boys who played an instrument were underrepresented in the lowest quartile of spelling performance: Only 27.5% of boys in the lowest quartile played an instrument whereas 61.5% of boys of the better quartiles were active musicians. Comparison between lowest quartile and all other quartiles combined proved a significant difference, X 2 (1) = 17.52, p < .0001; that turned into tendency towards significance after families without instruments were excluded, X 2 (1) = 6.86, p = .076.
Boys from a non-selected sample of third grade elementary school who play an instrument have shown a higher non-verbal IQ and were better in a formal spelling test compared to boys who did not play an instrument. The effects remained when controlled for musical interests by excluding families without instruments. The positive effect on reading vanished after this exclusion. The effect on spelling was independent of the influence from non-verbal IQ. A closer look at the distribution in spelling performance showed that only students in the lowest quartile differ from the others with respect to playing an instrument.
Our sample is not representative for the whole school population because of our inclusion criteria: male sex, right handedness, native German speakers, no class repetition. For the purpose of the current study the exclusion of girls seems to be a disadvantage. Advantageous, however, is the exclusion of children without native German language as our study needs a homogeneous group in terms of language acquisition. We did not focus on students suffering from dyslexia but included the whole range reading and spelling performance in a normal school population. The obtained return rate of the parents' questionnaire about the musical experience of their child (76%) is within acceptable limits.
A further restriction of our study is the retrospective design. The results do not clarify a causal relation from music education to cognitive performance, they only demonstrate correlations. Any questionnaire covering the past might introduce a positive bias. But this should not be a problem for the core of our results, because we do not expect that parents give a wrong answer to the simple question "Does your child play an instrument?" We did not specify this response: We included children who received recorder group lessons for 30 min per week
at age 8 for some weeks as well as children who played instruments with up to 4 hr individual lessons weekly starting at the age of 5. Given that weak inclusion criterion our results may even underestimate the observed effect.
A positive effect of playing an instrument on general cognitive abilities has been observed previously. Schellenberg (2004) reported an effect on IQ using Wechsler's WISC-III in 6-year-olds after keyboard or singing lessons for 36 weeks. In our study, the differences between the two groups were larger regarding both the IQ difference (Schellenberg delta 2.7 points vs. delta 6.6 points) and the effect size (Schellenberg d 0.35 vs. d 0.52). However, the studies differed considerably. For instance, we did not measure changes over time but compared two groups classified by the parents' questionnaire. Furthermore, we also focused on reading and writing performance and have therefore chosen a non-verbal IQ test to measure intelligence independently from progress in reading and writing. Schellenberg (2004) did not observe a difference between verbal and non-verbal subtests in his investigation.
The putative link between musical and language abilities is seen in the discrimination of rapid auditory events (Jakobson et al., 2003;Tallal & Gaab, 2006). Our results indicate a stronger link of playing an instrument in respect to spelling as to reading performance. This result is in contrast to findings reported in a meta-analysis (Butzlaff, 2000) and has, to the best of our knowledge, not been reported before.
Our findings can be explained in different ways. Firstly, it has to be taken into consideration that the German language has better phoneme-grapheme mapping than the English language. Therefore, it might be feasible that native German speakers benefit from auditory training such as playing an instrument in regard to the discrimination of rapid auditory events which in turn helps them to decipher the spelling of the German words.
Another explanation can be found in the double deficit hypothesis of dyslexia (Wolf & Bowers, 1999). According to this hypothesis, spelling deficits are associated with a phonological deficit whereas dysfluent reading is associated with a naming speed deficit (Wimmer & Mayringer, 2002;Wimmer, Mayringer, & Landerl, 2000). The naming speed deficit is supposedly not due to auditory information processing but rather to lexical access. In contrast, musical training mainly affects sound processing and therefore spelling capability.
In line with these results, studies with dyslexic risk populations of 6-year old children and a dyslexic population of 9-year old children demonstrated a positive effect of musical lessons on spelling performance and phonological abilities but not on reading (Overy, 2003). It therefore appears from these studies that children with particularly low reading and spelling abilities benefit most from playing a musical instrument. We cannot exclude that low performers dislike playing an instrument and therefore cause the observed group differences. Our data (Figure 3) show a leap between the percentages of players in the two lowest quartiles (delta 30%) that cannot be seen between the other quartiles (delta max. 8%). This pattern suggests a kind of threshold.
In our sample, the active participation in a choir or the lessons "First Experiences With Music" did not show the benefits found when children were playing musical instruments. It cannot be ruled out that this negative finding is based on intensity or quality of the musical activities. Another explanation could be the differences in specific motor skill between singing and playing an instrument. Further studies that control for intensity and quality of the musical courses might clarify if there is a specific advantage in playing an instrument.
It has been proposed that all specific relations observed so far can be explained by a carry-over effect of the relation between musical training and general abilities as measured by IQ (Schellenberg & Peretz, 2008). Indeed, in Schellenberg's studies such dependency was observed. Our data contrast these observations. We observed both a relation to general abilities as measured by non-verbal IQ and a relation to spelling performance that was still observed when controlled for general abilities. The relation between musical training and spelling performance may alternatively be explained by an increase of crystallized intelligence. Such an increase was found by Schellenberg (2004Schellenberg ( , 2006)). In our study only non-verbal IQ was tested so that we cannot rule out a mediation of crystallized intelligence on the effect of musical training on spelling. However, in this case we would have also expected an effect on reading. The isolated improvement of spelling but not reading suggests a specific link between musical training and spelling abilities mediated by improvement in auditory analysis.
From our data we propose that there is both an association of musical training and general abilities as well as specific spelling abilities. The link between training and general abilities has been demonstrated to be causal (Schellenberg, 2004). In the retrospective study, Schellenberg (2006) has also shown that the association between duration of musical training and academic average was evident even when IQ was held constant. This, too, suggests a general as well as an additional specific link between musical training and cognitive abilities. Our data justify a prospective study investigating a specific impact of musical training on spelling in languages with shallow orthographies such as German.
• volume 7 • 1-6
We would like to thank
This document was posted here by permission of the publisher. At the time of deposit, it included all changes made during peer review, copyediting, and publishing. The U.S. National Library of Medicine is responsible for all links within the document and for incorporating any publisher-supplied amendments or retractions issued subsequently. The published journal article, guaranteed to be such by Elsevier, is available for free, on ScienceDirect.
The variety of form and function in inositides (inositol lipids and phosphates) stems largely from the degree and isomeric specificity of phosphorylation of their common inositol ring. Thus their synthesis, and its differential regulation, lie primarily in the large family of kinases that phosphorylate the inositol ring. In a previous review in this publication (Irvine et al., 2006) we addressed the inositol phosphate kinases, with a specific focus on the Ins(1,4,5)P 3 3-kinase family. Here we turn to the inositol lipid kinases, with a specific focus on the phosphatidylinositol phosphate kinases. These are the family of enzymes that synthesize phosphatidylinositol 4,5-bisphosphate (PtdIns(4,5)P 2 ) (Fig. 1). The Type I enzymes (EC 2.7.1.68) are now known to catalyze the major route of PtdIns(4,5)P 2 synthesis, the 5phosphorylation of PtdIns4P (Rameh et al., 1997), and are here abbreviated as Type I PtdIns4P 5-kinases (PtdIns4P 5-kinases Iα, Iβ and Iγ). The Type II enzymes (EC 2.7.1.149) are 4-kinases (Rameh et al., 1997), whose preferred substrate (Rameh et al., 1997;Morris et al., 2000;Roberts et al., 2005) is PtdIns5P. They will be abbreviated here as PtdIns5P 4-kinases IIα, IIβ and IIγ.
Although there has been some discussion about the potential functions of the Type II enzymes, there is an emerging consensus that their major function is to regulate the levels of their substrate, PtdIns5P. The amount of PtdIns(4,5)P 2 that they will synthesize relative to the Type I enzymes is likely to be small (mostly because PtdIns5P is present at much lower levels than PtdIns4P (Roberts et al., 2005)). Also, the recent discovery and cloning in Majerus' lab of two isoforms of PtdIns(4,5)P 2 4-phosphatase (EC 3.1.3.78) (Ungewickell et al., 2005;Zou et al., 2007), one of which is partly nuclear (Zou et al., 2007) (see below for the significance of this) generates a very plausible cycle for the generation of PtdIns5P (by a PtdIns(4,5)P 2 4phosphatase) and its removal (by a Type II PtdIns5P 4-kinase). Lecompte et al. (2008) have suggested an interesting evolutionary pathway for the appearance of this cycle -that PtdIns5P generation and removal may occur by different routes involving different combinations of enzymes (e.g. from PtdIns via Type III PtdIns3P 5-kinase and myotubularin) © 2010 Elsevier Ltd. This document may be redistributed and reused, subject to certain conditions. and that the PtdIns(4,5)P 2 4-phosphatase/Type II PtdIns5P 4-kinase 'cycle' may be the latest of these to evolve, and is confined to metazoans. It is worth noting in this context that the significance of a PTEN homologue that is a 5-phosphatse (EC 3.1.3.36), but very specific for PtdIns5P (Pagliarini et al., 2004), remains to be explored fully -is this an alternative way of removing PtdIns5P, and if so, how does it fit in with the Type II PtdIns5P 4-kinases?
The functions of PtdIns5P are being actively explored, and it seems likely that our list of these is far from complete. This is a pertinent place to point out an error in our description of the mass assay for PtdIns5P, which we described in Roberts et al. (2005). This was a refinement of the original assay (Morris et al., 2000), and included, after loading the neomycin beads with lipids, a prior rinse with '50 mM Ammonium formate', before two elutions of inositol lipids with '2 M TEAB' (Triethylammonium bicarbonate). Both these two solutions were described as if they were made up in water, but they should of course have been described as being mixed with chloroform and methanol as with the other bead-elution solutions used in this protocol. Thus they are respectively: CHCl 3 :CH 3 OH:Ammonium Formate 5:10:2 by volume (final concentration of formate 50 mM); and CHCl 3 :CH3OH:2 M TEAB 5:10:2 by volume (final concentration of TEAB 0.55 M). This latter solution (with slightly altered CHCl 3 :CH 3 OH:2 M TEAB proportions), is correctly described by Zou et al. (2007), who also successfully used a more widely available type of glass bead for this experimental procedure.
To return to PtdIns5P functions, in the cytosol its principal effect has been suggested to be to increase Akt activation (Pendaries et al., 2006), perhaps by inhibiting its dephosphorylation (Ramel et al., 2009) or the dephosphorylation of PtdIns(3,4,5)P 3 (Carricaburu et al., 2003). Recently two putative PtdIns5P effectors, Dok-1 and Dok-2 were described in T cells (Guittard et al., 2009), and the significance of these will be an interesting area for future exploration. A better understood aspect of PtdIns5P function is in the nucleus, where ING-2 has been identified as an effector (Gozani et al., 2003). Divecha's group have put together a convincing case that PtdIns5P 4-kinase IIβ regulates the levels of nuclear PtdIns5P (Jones et al., 2006), whereby stressing cells activates p38 MAP kinase, which phosphorylates PtdIns5P 4-kinase IIβ on two Ser residues causing its inhibition, and thus PtdIns5P increases. This complements the demonstration that stress increases the nuclear localization of the PtdIns5P-generating enzyme Type I PtdIns(4,5)P 2 4-phosphatase (Zou et al., 2007). So, what regulates the nuclear localization of PtdIns5P 4-kinase IIβ, and is this the only Type II enzyme involved? We have made some interesting observations in this context recently, which we will briefly summarize here.
Although there have been some indications that endogenous PtdIns5P 4-kinase IIα may be partly nuclear (Divecha et al., 1993;Boronenkov et al., 1998), extensive studies on transfected cells have shown that the IIα and IIβ isoforms are respectively cytosolic and nuclear, with the latter being localized by a unique nuclear localization sequence consisting of an acidic α-helix (Ciruela et al., 2000;Clarke et al., 2007). Moreover, tagging the endogenous PtdIns5P 4-kinase IIβ by genomic tagging in DT40 cells (see below) confirmed that the endogenous enzyme appeared to be entirely nuclear, within the limits of the cell fractionation procedure employed (Richardson et al., 2007). However, a nuclear localization for PtdIns5P 4-kinase IIβ sat at odds with the phenotype of the knockout mouse (Lamia et al., 2004) complemented by transfection studies with PtdIns5P 4-kinase IIβ, both of which pointed to a predominantly cytoplasmic role for this enzyme (Carricaburu et al., 2003).
We have recently used genomic tagging in DT40 cells to begin to resolve some of these contradictions and to throw a new light on the relationship between these two Type II PtdIns5P 4-kinase isoforms (M.W., N. Bond, J. Richardson, K. Lilley, R.F.I. and J.H.C unpublished observations). In brief, we have found that PtdIns5P 4-kinases IIα and IIβ heterodimerize, such that a significant proportion of IIα is nuclear. Moreover, the enzymic activity of the IIα is much greater than that of IIβ, so it is possible that a, or the, major function of the IIβ isoform is simply to target the IIα to the nucleus. This is particularly interesting in the light of the evidence, based on mRNA levels, that in most tissues, indeed, in all that were studied except spleen, the IIβ isoform is more highly expressed than the IIα (Clarke et al., 2008).
Note that spleen is a site of synthesis of B cells (of which DT40s are, indirectly, an example), and T cells, so it is interesting that, as described above, Guittard et al. (2009) have recently discovered two new, cytoplasmic, putative PtdIns5P effectors in T cells. The spleen is also the site of synthesis of erythrocytes and platelets, from which PtdIns5P 4-kinase IIα was first purified and cloned by Boronenkov and Anderson (1995) and Divecha et al. (1995) respectively. Another intriguing link between erythrocytes and PtdIns5P 4-kinase IIα has emerged recently in a pair of α-thalassemic twins with very different β-globin gene expression levels, in which the only other gene showing a similar disparity in expression in their reticulocytes was PtdIns5P 4-kinase IIα (Wenning et al., 2009). This emphasis on PtdIns5P 4-kinase IIα in blood cells makes a contrast with all other tissues so far investigated in this context, where PtdIns5P 4-kinase IIβ mRNA is dominant over that for PtdIns5P 4-kinase IIα (Clarke et al., 2008). It throws open the possibility that in many tissues most PtdIns5P 4-kinase IIα is nuclear, and this will be especially so in muscle and liver (Clarke et al., 2008). This in turn sheds an alternative light on the experiments discussed above where PtdIns5P 4-kinase IIβ levels were manipulated (Carricaburu et al., 2003;Lamia et al., 2004), in that perhaps the major effect of decreasing PtdIns5P 4-kinase IIβ is to increase cytosolic IIα because there is no IIβ to take it to the nucleus. Of course we don't yet know if IIα/IIβ heterodimers are localized within the nucleus in the same place as IIβ/IIβ homodimers, nor whether they have distinct functions, and overall our discovery of this heterodimerization asks more questions than it answers. But that is what makes it interesting.
This has been the 'Cinderella' member of the family until recently, and we have been investigating some of its basic biology and biochemistry, which we summarize here.
When first cloned by Itoh et al. (1998), PtdIns5P 4-kinase IIγ was described as being highly expressed in kidney, localized to the endoplasmic reticulum, and phosphorylated when cells were stimulated by mitogens. We have confirmed that kidney expresses PtdIns5P 4-kinase IIγ more highly than any other tissue -as judged by mRNA levels, kidney is unique in expressing more PtdIns5P 4-kinase IIγ than PtdIns5P 4-kinase IIα and IIβ combined (Clarke et al., 2008). Even more striking is the localization of expression within the kidney, as PtdIns5P 4-kinase IIγ is largely confined to epithelial cells in the thick ascending limb and the intercalated cells of the collecting duct (Fig. 2 and Clarke et al., 2008).
The distribution of PtdIns5P 4-kinase IIγ in nervous tissue, the other type of tissue in which it is highly expressed (Clarke et al., 2008), shows a similar restricted and well defined expression in the brain and also the spinal cord (Clarke et al., 2009). This expression is limited to neurons, particularly the cerebellar Purkinje cells, pyramidal cells of the hippocampus, large neuronal cell-types in the cerebral cortex including pyramidal cells, and mitral cells in the olfactory bulb, but is not expressed in cerebellar, hippocampal formation or olfactory bulb granule cells (Clarke et al., 2009).
These distinct and specific expression patterns in brain and kidney will eventually tell us something of its function, though at present this is difficult to guess. However, two other properties of PtdIns5P 4-kinase IIγ point a way forward in understanding what it does. Firstly, in both kidney (Clarke et al., 2008) and brain (Clarke et al., 2009) it shows the same intracellular distribution: it appears to be associated with vesicles, which in kidney epithelial cells are concentrated towards the secreting end of the cells (Clarke et al., 2008).
We do not yet know the nature of these vesicles. In neurons PtdIns5P 4-kinase IIγ shows a partial colocalization with markers of cellular compartments of the endomembrane trafficking pathway (Clarke et al., 2009), and transfection experiments with mildly permeabilized HeLa cells showed a partial colocalization with Golgi markers such as GM130 and golgin 160, as well as the endosomal marker EEA1, but after more extensive cell permeabilization these correlations were decreased (Clarke et al., 2009). In transfected kidney cell lines (Clarke et al., 2008), PtdIns5P 4-kinase IIγ was again partially colocalized with GM130, but not with endoplasmic reticulum or ERGIC markers, and the GM130 relationship survived Brefeldin A treatment of the cells.
Together these data suggest a close association of PtdIns5P 4-kinase IIγ with vesicular cell trafficking linked with the Golgi apparatus. In the kidney this would be associated with the business of inserting and removing plasma membrane transporters and channels, and in the brain either trafficking of vesicles along neuronal processes or perhaps the delivery/insertion of channels and transporters to their correct locations. PtdIns(4,5)P 2 is of course known to be associated with the regulation of activity and trafficking of ion channels and transporters, but returning to the low levels of PtdIns5P compared to PtdIns4P in cells plus the evidence discussed above for possible roles of PtdIns5P in other cellular functions, we think it much more likely that PtdIns5P is the active molecule in whatever vesicular processing/transporting events PtdIns5P 4-kinase IIγ is regulating.
Secondly, there is another aspect of PtdIns5P 4-kinase IIγ biochemistry which may point to how it functions -actually it is two properties that suggest the same thing. PtdIns5P 4-kinase IIγ is, when bacterially expressed, catalytically close to inactive (Clarke et al., 2008); it has even less activity than PtdIns5P 4-kinase IIβ (above). But when immunoprecipitated from eukaryotic cells, PtdIns5P 4-kinase IIγ does show some catalytic activity, which can be accounted for by its association with PtdIns5P 4-kinase IIα, an association that we have demonstrated directly in vitro (Clarke et al., 2008). Thus we can see a parallel with the PtdIns5P 4-kinase IIα/IIβ heterodimerization that we have discussed above, and our current working hypothesis for PtdIns5P 4-kinase IIγ is that it is associated with a sub-population of vesicles in the cell trafficking system, and by heterodimerization with PtdIns5P 4-kinase IIα it targets the latter enzyme to the relevant vesicles, where it regulates PtdIns5P levels for a function yet to be defined (see above). We should note also that there is still the possibility that by associating with a Type I PtdIns4P 5-kinase activity (Hinchliffe et al., 2002), PtdIns5P 4kinase IIα may also target one of these enzymes to the relevant vesicles (discussed further below), so PtdIns(4,5)P 2 might yet be a relevant signal in this context too.
So our recent data are painting a different picture of Type II PtdIns5P 4-kinase function and regulation from that which we have had before. We can suggest the possibility that PtdIns5P 4-kinase IIα is the (possibly the only significantly) active enzyme of the three isoforms, and that it may have its own individual function(s) in cells. Additionally it can be targeted by the other isoforms, to the nucleus by PtdIns5P 4-kinase IIβ, or to trafficking/secretory vesicles by PtdIns5P 4-kinase IIγ. This picture paints a remarkable (and we suspect unique) relationship between three isoforms of the same lipid kinase family.
Before leaving the subject of the heterodimerization of Type II PtdIns5P 4-kinases, another dimerization (or at least, association) needs discussing in this context, which is the association of PtdIns5P 4-kinase IIα with Type I PtdIns4P 5-kinases. This binding is clearly documented both with transfected and endogenous Type I enzymes (Hinchliffe et al., 2002). We do not yet know which structural domains of either partner (Type I or II) are involved in this interaction. The regions involved in heterodimerization between PtdIns5P 4-kinases IIα, IIβ, and IIγ is fairly certain: the β isoform is a homodimer in the crystals used for structural analysis, with the interaction between the monomers being due to two opposing β-pleated sheets (Rao et al., 1998), and this part of the enzyme is identical in all three isoforms. But a crucial question is whether a hetero (or homo) dimerized Type II PtdIns5P 4-kinase is capable of also associating with a Type I PtdInd4P 5-kinase, or, put another way, does the Type I/II association involve the same region of the enzymes so that a Type II enzyme is compelled to associate either with a Type I or a Type II PtdInsP kinase, but not with both? This is not a trivial question, as one of the interesting ideas that we discussed above is that the localization of the highly active Type IIα PtdIns5P 4-kinase may be governed entirely by the relative levels of expression of the IIβ and IIγ isoforms, levels that can vary greatly between tissues (Clarke et al., 2008). If we have to factor into that idea the relative levels of the three Type I PtdIns4P 5-kinases, each of which can interact with the Type IIα PtdIns5P 4-kinase (Hinchliffe et al., 2002) (and we do not yet know if Type IIβ or IIγ can interact with Type I activities), we have a very complex picture. Indeed, such a proposal would seem to make a nonsense of the above idea that relative levels of the three Type II isoforms has physiological relevance in dictating their localization, and the whole process of PtdInsP kinases associating would have to be tightly and complicatedly regulated. We should add that, when pulling down PtdIns5P 4-kinase IIβ from DT40 cells (above) we did not detect any Type I isoforms by mass spectroscopy, so any Type I/II interaction must be of lower affinity than Type II/II interactions. Overall we feel that a more likely idea is that the domains by which Type II enzymes interact with each other (Rao et al., 1998) are different from those that PtdIns5P 4-kinase IIα uses to associate with Type I enzymes, and so a Type II PtdInsP 4-kinase homo or heterodimer can additionally associate with a (two?) Type I PtdIns4P 5-kinase enzyme molecule(s). The functional consequences of this latter association still elude us.
Our recent studies of the Type II PtdIns5P 4-kinases have revealed that the Type IIα isoform is very much more active than the IIβ or IIγ isoforms, and that it can (and does physiologically) heterodimerize with them. This suggests the idea that the Type IIα enzyme is targeted to the nucleus (by dimerization with Type IIβ), to secretory/transport vesicles (by dimerization with Type IIγ), or to the cytoplasm (as a homodimer), with the relative proportions of PtdIns5P 4kinase activity at these localizations being regulated by the relative amounts of the three Type II isoforms expressed in any cell. The targeting to vesicles by PtdIns5P 4-kinase IIγ is likely to be of particular significance in epithelial cells in specific regions of the kidney tubules and in a sub-population of neurons in the brain and the spinal cord. The relationship between this dimerization between Type II PtdIns5P 4-kinase isoforms and the known ability of Type IIα PtdIns5P 4-kinase to associate with Type I PtdIns4P 5-kinases remains to be explored. Clarke et al. Page 8 Published as: Adv Enzyme Regul. 2010 ; 50(1): 12-18. Sponsored Document Sponsored Document Sponsored Document
Published as: Adv Enzyme Regul. 2010 ; 50(1): 12-18.Sponsored DocumentSponsored Document Sponsored Document
The work from our laboratory described above was funded by a
Clarke et al.
Traditional psychometric approaches towards assessment tend to focus exclusively on quantitative properties of assessment outcomes. This may limit more meaningful educational approaches towards workplace-based assessment (WBA). Cognition-based models of WBA argue that assessment outcomes are determined by cognitive processes by raters which are very similar to reasoning, judgment and decision making in professional domains such as medicine. The present study explores cognitive processes that underlie judgment and decision making by raters when observing performance in the clinical workplace. It specifically focuses on how differences in rating experience influence information processing by raters. Verbal protocol analysis was used to investigate how experienced and non-experienced raters select and use observational data to arrive at judgments and decisions about trainees' performance in the clinical workplace. Differences between experienced and non-experienced raters were assessed with respect to time spent on information analysis and representation of trainee performance; performance scores; and information processing--using qualitative-based quantitative analysis of verbal data. Results showed expert-novice differences in time needed for representation of trainee performance, depending on complexity of the rating task. Experts paid more attention to situation-specific cues in the assessment context and they generated (significantly) more interpretations and fewer literal descriptions of observed behaviors. There were no significant differences in rating scores. Overall, our findings seemed to be consistent with other findings on expertise research, supporting theories underlying cognition-based models of assessment in the clinical workplace. Implications for WBA are discussed.
Recent developments in the continuum of medical education reveal increasing interest in performance assessment, or workplace-based assessment (WBA) of professional competence. In outcome-based or competency-based training programs, assessment of performance in the workplace is a sine qua non (Van der Vleuten and Schuwirth 2005). Furthermore, the call for excellence in professional services and the increased emphasis on life-long learning require professionals to evaluate, improve and provide evidence of dayto-day performance throughout their careers. Workplace-based assessment (WBA) is therefore likely to become an essential part of both licensure and (re)certification procedures, in health care just as in other professional domains such as aviation, the military and business (Cunnington and Southgate 2002;Norcini 2005).
Research into WBA typically takes the psychometric perspective, focusing on quality of measurement. Norcini (2005), for instance, points to threats to reliability and validity from uncontrollable variables, such as patient mix, case difficulty and patient numbers. Other studies show that the utility of assessment results is compromised by low inter-rater reliability and rater effects such as halo, leniency or range restriction (Kreiter and Ferguson 2001;Van Barneveld 2005;Gray 1996;Silber et al. 2004;Williams and Dunnington 2004;Williams et al. 2003). As a consequence, attempts to improve WBA typically focus on standardization and objectivity of measurement by adjusting rating scale formats and eliminating rater errors through rater training. Such measures have met with mixed success at best (Williams et al. 2003).
One might question, however, whether an exclusive focus on the traditional psychometric framework, which focuses on quantitative assessment outcomes, is appropriate in WBA-research. Research in industrial psychology demonstrates that assessment of performance in the workplace is a complex task which is defined by a set of interrelated processes. Workplace-based assessment relies on judgments by professionals, who typically have to perform their rating tasks in a context of time pressure, non-standardized assessment tasks and ill-defined or competing goals (Murphy and Cleveland 1995). Findings from research into performance appraisal also indicate that contextual factors affect rater behavior and thus rating outcomes (Levy and Williams 2004;Hawe 2003). Raters are thus continuously challenged to sample performance data; interpret findings; identify and define assessment criteria; and translate private judgments into sound (acceptable) decisions. Perhaps performance rating in the workplace is not so much about 'measurement' as it is about 'reasoning', 'judgment' and 'decision making' in a dynamic environment. From this perspective, our efforts to optimize WBA may benefit from a better understanding of raters' reasoning and decision making strategies. This implies that new and alternative approaches should be used to investigate assessment processes, with a shift in focus from quantitative properties of rating scores towards analysis of the cognitive processes that raters are engaged in when assessing performance.
The idea of raters as information processors is central to cognition-based models of performance assessment (Feldman 1981;De Nisi 1996). Basically, these models assume that rating outcomes vary, depending on how raters recognize and select relevant information (information acquisition); interpret and organize information in memory (cognitive representation of ratee behavior); search for additional information; and finally retrieve and integrate relevant information in judgment and decision making. These basic cognitive processes are similar to information processing as described in various professional domains, such as management, aviation, the military and medicine (Walsh 1995;Ross et al. 2006;Gruppen and Frohna 2002). Research findings from various disciplines show that large individual variations in information processing can occur, related to affect, motivation, time pressure, local practices and prior experience (Levy and Williams 2004;Gruppen and Frohna 2002).
In fact, task-specific expertise has been shown to be a key variable in understanding differences in information processing--and thus task performance (Ericsson 2006). There is ample research indicating that prolonged task experience helps novices develop into expertlike performers through the acquisition of an extensive, well-structured knowledge base as well as adaptations in cognitive processes to efficiently process large amounts of information in handling complex tasks. Research findings consistently indicate that these differences in cognitive structures and processes impact on proficiency and quality of task performance (Chi 2006). For instance, a main characteristic of expert behavior is the predominance of rapid, automatic pattern recognition in routine problems, enabling extremely fast and accurate problem solving (Klein 1993;Coderre et al. 2003). When confronted with unfamiliar or complex problems, however, experts tend to take more time to gather, analyze and evaluate information in order to better understand the problem, whereas novices are more prone to start generating a problem solution or course of action after minimal information gathering (Ross et al. 2006;Voss et al. 1983). Another robust finding in expertise studies is that, compared with non-experts, experts see things differently and see different things. In general, experts make more inferences on information, clustering sets of information into meaningful patterns and abstractions (Chi et al. 1981;Feltovich et al. 2006). Studies on expert behavior in medicine, for instance, show that experts have more coherent explanations for patient problems, make more inferences from the data and provide fewer literal interpretations of information (Van de Wiel et al. 2000). Similar findings were described in a study on teacher supervision (Kerrins and Cushing 2000). Analysis of verbal protocols showed that inexperienced supervisors mostly provided literal descriptions of what they had seen on the videotape. More than novices, experienced supervisors interpreted their observations as well as made evaluative judgments, combining various information into meaningful patterns of classroom teaching. Overall, experts' observations focused on students and student learning, whereas non-experts focused more on discrete aspects of teaching.
Research findings also indicate that experts pay attention to cues and information that novices tend to ignore. For instance, experts typically pay more attention to contextual and situation-specific cues while monitoring and gathering information, whereas novices tend to focus on literal textbook aspects of a problem. In fact, automated processing by medical experts seems to heavily rely on contextual information (e.g. Hobus et al. 1987).
Finally, experts generally have better (more accurate) self-monitoring skills and greater cognitive control over aspects of performance where control is needed. Not only are experts able to devote cognitive capacity to self-monitoring during task performance, their richer mental models also enable them to better detect errors in their reasoning. Feltovich et al. (1984), for instance, investigated flexibility of experts versus non-experts on diagnostic tasks. Results showed that novices were more rigid and tended to adhere to initial hypotheses, whereas experts were able to discover that the initial diagnosis was incorrect and adjust their reasoning accordingly. In their study on expert-novice differences in teacher supervision, Kerrins and Cushing (2000) found that experts were more cautious in over-and underinterpreting what they were seeing. Although experts made more interpretative and evaluative comments, they more often qualified their comments with respect to both their interpretation of the evidence and the limitations of their task environment.
Based on the conceptual frameworks of cognition-based performance assessment and expertise research, it is perfectly conceivable that rater behavior in WBA changes over time, due to increased task experience. Extrapolating findings from research in other domains, different levels of expertise may then be reflected in differences in task performance, which may have implications not only for utility of work-based assessments, but also for the way we select and train our raters. Given the increased significance of WBA in health professions education, the question can therefore be raised whether expertise effects as described also occur in performance assessment in the clinical domain. The present study aims to investigate cognitive processes related to judgment and decision making by raters observing performance in the clinical workplace. Verbal protocols; time spent on performance analysis and representation, and performance scores were analyzed to assess differences between experienced and non-experienced raters. More specifically, we explored 4 hypotheses that arose from the assumption that task experience determines information processing by raters. Firstly, we expected experienced raters to take less time, compared to non-experienced raters, in forming initial representations of trainee performance when observing prototypical behaviors, but more time when more complex behaviors are involved. Secondly, we expected experienced raters to pay more attention to situation-specific cues in the context of the rating task, such as patient or case specific cues; the setting of the patient encounter and ratee experience (phase of training). Thirdly, verbal protocols of experienced raters were expected to contain more inferences (interpretations) and fewer literal descriptions of behaviors. Finally, experienced raters were expected to generate more self-monitoring statements during performance assessment.
The participants in our study were GP-supervisors who were actively involved as supervisor-assessor in general practice residency training. General practice training in the Netherlands has a long tradition of systematic direct observation and assessment of trainee performance throughout the training program. GP-supervisors are all experienced general practitioners, continuously involved in supervision of trainees on a day-to-day basis. They are trained in assessment of trainee performance.
In our study, we defined the level of expertise as the number of years of task-relevant experience as a supervisor-rater. Since there is no formal equivalent of elite rater performance we adopted a relative approach to expertise. This approach assumes that novices develop into experts through extensive task experience and training (Chi 2006;Norman et al. 2006). In general, about 7 years of continuous experience in a particular domain is necessary to achieve expert performance (e.g. Arts et al. 2006). Registered GP-supervisors with different levels of supervision experience were invited to voluntarily participate in our study; a total of 34 GP-supervisors participated. GP-supervisors with at least 7 years of experience as supervisor-rater were defined as 'experts'. The 'expert group' consisted of 18 GP-supervisors (number of years of experience M = 13.4; SD = 5.9); the 'non-expert group' consisted of 16 GP-supervisors (number of years of experience M = 2.6; SD = 1.2). Levels of experience between both groups differed significantly (t(32) = 7.2, p \ .001). Participants received financial compensation for their participation.
The participants watched two DVDs, each showing a final-year medical student in a 'reallife' encounter with a patient. The DVDs were selected purposefully with respect to both patient problems and students' performance. Both DVDs presented 'straightforward' patient problems that are common in general practice: atopic eczema and angina pectoris. These cases were selected to ensure that all participants (both experienced and nonexperienced raters) were familiar with required task performance. DVD 1-atopic eczemalasted about 6 min and presented a student showing prototypical and clearly substandard behavior with respect to communication and interpersonal skills. This DVD was considered to present a non-complex rating task. DVD 2-angina pectoris-lasted about 18 min and was considered to present a complex rating task with the student showing more complex behaviors with respect to both communication and patient management. Permission had been obtained from the students and the patients to record the patient encounter and use the recording for research purposes.
The participants used two instruments to rate student performance (Figs. 1, 2): a onedimensional, overall rating of student performance on a five-point Likert scale (1 = poor to 5 = outstanding) (R1), and a list of six clinical competencies (history taking; physical examination; clinical reasoning and diagnosis; patient management; communication with the patient; and professionalism), each to be rated on a five-point Likert scale (1 = poor to 5 = outstanding) (R2). Rating scales were kept simple to allow for maximum idiosyncratic cognitive processing. The participants were not familiar with the rating instruments and had not been trained in their use.
Research procedure and data collection We followed standard procedures for verbal protocol analysis to capture cognitive performance (Chi 1997). 1 Before starting the first DVD, participants were informed about procedures and received a set of verbal instructions. Raters were specifically asked to ''think aloud'' and to verbalize all their thoughts as they emerged, as if they were alone in the room. If a participant were silent for more than a few seconds, the research assistant reminded him or her to continue. Permission to audiotape the session was obtained. For each of the DVDs the following procedure was used:
1. DVD starts. The participant signals when he or she feels able to judge the student's performance, and the time from the start of the DVD to this moment is recorded (T1). T1 represents the time needed for problem representation, i.e. initial representation of trainee performance. 2. The DVD is stopped at T1. The participant verbalizes his/her first judgment of the trainee's performance (verbal protocol (VP) 1). 3. The participant provides an overall rating of performance on the one-dimensional rating scale (R1T1), thinking aloud while filling in the rating form (VP2).
4. Viewing of the DVD is resumed from T1. When the DVD ends (T2), the participant verbalizes his/her judgment (VP3) and provides an overall rating (R1T2). 5. The participant fills in the multidimensional rating form (R2) for one of the DVDs (alternately DVD 1 or DVD 2) and verbalizes his or her thoughts while doing so (VP4).
We used a balanced design to control for order effects; the participants within each group were alternately assigned to one of two viewing conditions with a different order of the DVDs. All the audiotapes were transcribed verbatim.
The transcriptions of the verbal protocols were segmented into phrases by one of the researchers (MG). Segments were identified on the basis of semantic features (i.e. content features-as opposed to non content features such as syntax). Each segment represented a Fig. 2 6-Dimensional global rating scale clinical competencies (R2)
single thought, idea or statement (see Table 1 for some examples). Each segment was assigned to coding categories, using software for qualitative data analysis (Atlas.ti 5.2). Different coding schemes were used to specify 'the nature of the statement'; 'type of verbal protocol' and 'clinical presentation' (Table 1). The coding categories for 'nature of statement' were based on earlier studies in expert-novice information processing (Kerrins and Cushing 2000;Boshuizen 1989;Sabers et al. 1991) and included 'description', 'interpretation', 'evaluation', 'contextual cue' and 'self-monitoring'. Repetitions were coded as such.
All verbal protocols were coded by two independent coders (MG, ME). Inter-coder agreement based on five randomly selected protocols was only moderate (Cohen's kappa 0.67), and therefore the two coders coded all protocols independently and afterwards compared and discussed the results until full agreement on the coding was reached.
The data were exported from Atlas.ti to SPSS 17.0. For each participant, the numbers of statements per coding category were transformed to percentages in order to correct for between-subject variance in verbosity and elaboration of answers. Because of the small sample sizes and non-normally distributed data, non-parametric tests (Mann-Whitney U) were used to estimate the differences between the two groups in the time to initial representation of performance (T1); the nature of the statements, and performance ratings per DVD. We calculated effect sizes by using the formula ES = Z/HN as is suggested for non-parametric comparison of two independent samples, where Z is the z-score of the
Table 1 Verbal protocol coding schemes Nature of statement 1. Descriptions: (literal) descriptions of student behaviour (''he is smiling to the patient''; ''he asks if this happened before'')
2. Inferences: interpretations and abstractions of performance (''he is an authoritarian doctor''; ''he is clearly a young professional''; ''it seems that he takes no pleasure in being a doctor'') 3. Evaluations: normative judgments, referring to implicit or explicit standards (''his physical examination skills are very poor''; ''overall, his performance is satisfactory'') 4. Contextual cue: remarks referring to case-specific or context-specific cues such as patient characteristics, setting of the patient encounter, context of the assessment task (''this patient is very talkative''; ''this looks like a hospital setting, not general practice''; ''he is being videotaped'') 5. Self-monitoring: reflective remarks, nuancing (''although I am not sure if I saw this correctly''; ''on hindsight I shouldn't have…'' ''……. but on the other hand most senior residents do not know how to handle these problems either''); self-instruction and structuring of rating process (''first I am going to look at ….''; ''when evaluating performance I always look at atmosphere and balance''); explication of standards and performance theory (''one should always start with open-ended questions''; ''from a firstyear resident I expect……'') Mann-Whitney statistic and N is the total sample size (Field 2009, p. 550). Effect sizes equal to 0.1, 0.3, and 0.5, respectively, indicate a small, medium, and large effect. For within-group differences of overall ratings (R1T1 versus R1T2) the Wilcoxon signed rank test was applied.
Table 2 shows the results for the time to problem representation (T1) and the overall performance ratings for each DVD. Time to T1 was similar for experienced and nonexperienced raters when observing prototypical behavior (DVD 1). However, when observing the more complex behavioral pattern in DVD 2, experienced raters took significantly longer time for monitoring and gathering of information, whereas there was only minimal increase in time for non-experts (U = 79.00, p = .03, ES = 0.38).
Table 2 shows non-significant differences between the two groups in the rating scores. A Wilcoxon signed ranks test, however, showed significant within-group differences between the rating scores at T1 and T2. In the expert group these differences were significant for both the dermatology case (Z = -2.31, p = .02, ES = 0.40) and the cardiology case (Z = -2.95, p = .003, ES = 0.51). In the non-expert group, significant differences were found for the cardiology case only (Z = -2.49, p = .01, ES = 0.43). The impact of the differences in rating scores at T1 resp. T2 is illustrated by (significant) shifts in the percentage of ratings representing a 'fail' (R1 B 2). In the expert group, the proportion of failures for the dermatology case was 61% at T1 versus 89% at T2. For the cardiology case the proportion of failures shifted from 11% (T1) to 56% at T2 in the expert group, and from 6% (T1) to 50% at T2 in the non-expert group.
Table 3 presents the percentages (median, inter-quartile range) for the nature of the statements for each group, by verbal protocol and across all protocols (= overall, VP1 ? VP2 ? VP3 ? VP4). Overall, the experienced raters generated significantly more inferences or interpretations of student behavior (U = 62.5, p = .005, ES = 0.48), whereas non-experts provided more descriptions (U = 68.5, p = .009, ES = 0.45). The verbal protocols after viewing the entire DVD (VP3) showed similar and significant differences between experienced and non-experienced raters with respect to interpretations (U = 71.5, p = .01, ES = 0.43) and descriptions of behaviors (U = 73, p = .01, ES = 0.42). Experienced raters also generated significantly more interpretations when filling out the six-dimensional global rating scale (U = 63, p = .004, ES = 0.48). Table 3 also shows that experienced raters generated more references to contextspecific and situation-specific cues. This difference was significant at T1 (U = 83, p = .04, ES = 0.37), and similar and near-significant (U = 89, p = .06) for the overall protocols and protocol VP3.
Evaluations showed no significant differences, except for VP2, with experienced raters generating significantly more evaluations (U = 87.5, p = .05, ES = 0.34).
No significant between-group differences were found with respect to self-monitoring.
Based on expertise research in other domains, we hypothesized that experienced raters would differ from non-experienced raters with respect to cognitive processes that are related to judgment and decision making in workplace-based assessment.
As for the differences in the time taken to arrive at the initial representation of trainee performance, the results partially confirm our hypothesis. It is contrary to our expectations that the expert raters took as much time as the non-expert raters with the case presenting prototypical behavior, but our expectations are confirmed for the case with complex trainee behavior, with the experts taking significantly more time than the non-experts. This finding is consistent with other findings on expertise research (Ericsson and Lehmann 1996). Whereas non-experienced raters seem to focus on providing a correct solution (i.e. judgments or performance scores) irrespective of the complexity of the observed behavior, expert raters take more time to monitor, gather and analyze the information before arriving at a decision on complex trainee performance. Our non-significant results with respect to prototypical behavior may be explained by the rating stimulus in our study. The dermatology case may have been too short, and the succession of typical student behaviors too quick to elicit differences. Moreover, the clearly substandard performance in the stimulus may have elicited automatic information processing and pattern recognition in both groups (Eva 2004). Our results for the cardiology case, however, confirm that, with more complex behaviors, experienced raters seem to differ from non-experienced raters with respect to their interpretation of initial information -causing them to search for additional information and prolonged monitoring of trainee behavior.
As for the verbal protocols, the overall results appear to confirm the hypothesized differences between expert and non-expert raters in information processing while observing and judging performance. Compared to non-experienced raters, experienced raters generated more inferences on information and interpretations of student behaviors, whereas non-experienced raters provided more literal descriptions of the observed behavior. These findings suggest that non-experienced raters pay more attention to specific and discrete aspects of performance, whereas experienced raters compile different pieces of information to create integrated chunks and meaningful patterns of information. Again, this is consistent with other findings from expertise research (Chi 2006). Our results also suggest that expert raters have superior abilities to analyze and evaluate contextual and situation-specific cues. The raters in our study appeared to pay more attention to contextual information and to take a broader view, at least in their verbalizations of performance
Table 3 Percentages of statements in verbal protocols for experienced raters (Exp) and non-experienced raters
Descriptives 19.8 (13.2) 25.3 (11.0) a 19.8 (20.1) 26.1 (12.8) 10.8 (18.2) 10.7 (18,6) 16.9 (16.3) 29.4 (13.3) a 18.2 (16.5) 20.4 (10.6) Inferences 19.0 (7.9) 14.7 (5.1) a 38.9 (22.9) 37.5 (21.2) 14.8 (25.6) 20.0 (25,1) 14.5 (9.8) 6.8 (9.2) a 13.5 (14.0) 5.6 (4.6) a Evaluations 24.4 (10.4) 24.9 (4.7) 7.6 (16.0) 5.4 (10.1) 33.3 (15.1) 18.2 (20,0) a 24.9 (16.3) 25.9 (12.6) 35.9 (17.3) 41.3 (17.3) Contextual cues 12.9 (7.7) 10.4 (7.3) 13.2 (7.7) 6.1 (14.6) a 8.1 (13.5) .0 (13.7) judgments. They integrate relevant background information and observed behaviors into comprehensive performance assessments. The differences between experts and non-experts were most marked at the initial stage of information gathering and assessment of performance (VP1). The setting of the patient encounter, patient characteristics and the context of the assessment task all seem to be taken into account in the experts' initial judgments.
These findings suggest that expert raters possess more elaborate and coherent mental models of performance and performance assessment in the clinical workplace. Similar expert-novice differences have been reported in other domains. Cardy et al. (1987), for instance, found that experienced raters in personnel management use more and more sophisticated categories for describing job performance. Our findings are in line with many other studies in expertise development, which consistently demonstrate that compared with novices, experts have more elaborate and well-structured mental models, replete with contextual information.
The results of our study showed that, within groups, the initial ratings at T1 differed significantly from the ratings after viewing the entire DVD (T2). Thus our findings suggest that both expert and non-expert raters continuously seek and use additional information, readjusting judgments while observing trainee performance. Moreover, this finding points to the possibility that rating scores, provided after brief observation, may not accurately reflect overall performance. This could have consequences for guidelines for minimal observation time and sampling of performance in WBA. Our results did not reveal significant differences in rating scores between experts and non-experts. We were therefore not able to confirm previous research findings in industrial psychology demonstrating that expert raters provide more accurate ratings of performance compared with non-experts (e.g. Lievens 2001). Possible explanations are that, as a result of previous training and experience in general practice, both groups may have common notions of what constitutes substandard versus acceptable performance in general practice. Shared frames of reference, a rating scale that precludes large variations in performance scores and the small sample size may have caused the equivalent ratings in both groups.
Contrary to our expectations, the experts in our study do not appear to demonstrate more self-monitoring behavior while assessing performance. An explanation might be that our experimental setting, in which participants were asked to think aloud while providing judgments about others, induced more self-explanations. The task of verbalizing thoughts while filling out a rating scale and providing a performance score may have introduced an aspect of accountability into the rating task, with both experienced and non-experienced raters feeling compelled to explain and justify their actions despite being instructed otherwise. These self-explanations and justifications of performance ratings may also explain the absence of any significant differences in rating scores between the groups. Several studies have shown that explaining improves subjects' performance (e.g. Chi et al. 1994). And research into performance appraisal in industrial organizations has demonstrated that raters who are being held accountable provide more accurate rating scores (e.g. Mero et al. 2003). The think aloud procedure may therefore have resulted in fairly accurate rating scores in both groups. This explanation is substantiated by the comments of several raters on effects of verbalization [e.g. ''If I had not been forced to think aloud, I would have given a 3 (satisfactory), but if I now reconsider what I said before, I want to give a 2 (borderline) ''].
What do our findings mean and what are the implications for WBA?
Our findings offer indications that in workplace-based assessment of clinical performance expertise effects occur that are similar to those reported in other domains, providing support for cognition-based models of assessment as proposed by Feldman (1981) and others.
There are several limitations to our study. Participants in our study were all volunteers and therefore may have been more motivated to carefully assess trainees' performance. Together with the experimental setting of our study, this may limit generalization of our findings to raters in 'real life' general practice. Real life settings are most often characterized by time constraints, conflicting tasks and varying rater commitment, which may all impact on rater information processing. Another limitation of our study is the small sample size, although the sample used in not uncommon in qualitative research of this type. Also, statistical significant differences emerge despite the relatively small sample size and the use of less powerful, but more robust non-parametric tests. Finally, we used only years of experience as a measure of expertise; other variables such as intelligence, actual supervisor performance, commitment to teaching and assessment, or reflectiveness were not measured or controlled for. Time and experience are clearly important variables in acquiring expertise, though. The purpose of our study was not to identify and elicit superior performance of experts. Rather, we investigated whether task-specific experience affects the way in which raters process information when assessing performance. In this respect, our relative approach to expertise is very similar to approaches in expertise research in the domain of clinical reasoning in medicine (Norman et al. 2006).
If our findings reflect research findings from expertise studies in other domains, this may have important implications for WBA. Our study appears to confirm the existence of differences in raters' knowledge structures and reasoning processes resulting from training as well as personal experience. Such expert-novice differences may impact the feedback that is given to trainees in the assessment process.
Firstly, more enriched processing and better incorporation of contextual cues by experienced raters can result in qualitatively different, more holistic feedback to trainees, focusing on a variety of issues. Expert raters seem to take a broader view, interpreting trainee behavior in the context of the assessment task and integrating different aspects of performance. This enables them to give meaning to what is happening in the patient encounter. Non-experienced raters on the other hand may focus more on discrete 'checklist' aspects of performance. Similar findings have been reported by Kerrins and Cushing (2000) in their study on supervision of teachers.
Secondly, thanks to more elaborate performance scripts, expert raters may rely more often on top-down information processing or pattern recognition when observing and judging performance -especially when time constraints and/or competing responsibilities play a role. As a consequence, expert judgments may be driven by general, holistic impressions of performance neglecting behavioral detail (Murphy and Balzer 1986;Lievens 2001), whereas non-experienced raters may be more accurate at the behavioral level. However, research in other domains has shown that, despite being likely to chunk information under normal conditions, experts do not lose their ability to use and recall 'basic' knowledge underlying reasoning and decision making (Schmidt and Boshuizen 1993). Moreover, research findings indicate that experts demonstrate excellent recall of relevant data when asked to process a case deliberately and elaborately (Norman et al. 1989;Wimmers et al. 2005). Similarly, when obliged to process information elaborately and deliberately, experienced raters may be as good as non-experienced raters in their recall of specific behaviors and aspects of performance. Optimization of WBA may therefore require rating procedures and formats that force raters to elaborate on their judgments and substantiate their ratings with concrete and specific examples of observed behaviors.
Finally, our findings may have consequences for rater training, not only for novice raters, but for more experienced raters as well. Clearly, there is a limit to what formal training can achieve and rater expertise seems to develop through real world experience. Idiosyncratic performance schemata are bound to develop as a result of personal experiences, beliefs and attitudes. Development of shared mental models and becoming a true expert, however, may require deliberate practice with regular feedback and continuous reflection on strategies used in judging (complex) performance in different (ill-defined) contexts (Ericsson 2004).
Further research should examine whether our findings can be reproduced in other settings. Important areas for study are the effects of rater expertise on feedback and rating accuracy. Is there a different role for junior and senior judges in WBA? Our findings also call for research into the relationship between features of the assessment system, such as rating scale formats, and rater performance. Rating scale formats affect cognitive processing in performance appraisal to the extent that the format is more or less in alignment with raters' ''natural'' cognitive processes. Assuming that raters' cognitive processes vary with experience, it is to be expected that different formats will generate differential effects on information processing in raters with different levels of experience. For instance, assessment procedures which focus on detailed and complete registration of ratee behaviors may disrupt the automatic, top-down processing of expert raters, resulting in inaccurate ratings. We need to understand which rating formats facilitate or hinder rating accuracy and provision of useful feedback at different levels of rater expertise. There is also increasing evidence that rater behavior is influenced by factors like trust in the assessment system, rewards and threats (consequences of providing low or high ratings), organizational norms, values, etc. (Murphy and Cleveland 1995). These contextual factors may lead to purposeful 'distortion' of ratings. Future research should therefore include field research to investigate possible effects of these contextual factors on decision making. Finally we wish to emphasize the need for more in-depth and qualitative analysis of raters' reasoning processes in performance assessment. How do performance schemata of experienced raters differ from those of non-experienced raters? How do raters combine and weigh different pieces of information when judging performance? How are performance schemata and theories linked to personal beliefs and attitudes?
In devising measures to optimize WBA we should first and foremost take into account that raters are not interchangeable measurement instruments, as is generally assumed in the psychometric assessment framework. In fact, a built-in characteristic of cognitive approaches to performance assessment is that raters' information processing is guided by their 'mental models' of performance and performance assessment. Our study shows that raters' judgment and decision making processes change over time due to task experience, supporting the need for research as described above.
b
Presented are the median and the inter-quartile range (in parentheses). Experts take significantly (U = 79.00, p = .03, ES = 0.38) more time for monitoring and gathering of information than novices when observing performance on DVD 2 (cardio case). Rating scores are based on a 5-point scale (1 = poor, 5 = outstanding) Values in the same column (DVD 1 and DVD 2 resp.) with different superscripts differ significantly (Wilcoxon Signed Ranks test, p \ .05)
Verbal protocols refer to the collection of participants' verbalizations of their thoughts and behaviors, during or immediately after performance of cognitive tasks. Typically, participants are asked to ''think aloud'' and to verbalize all their thoughts as they emerge, without trying to explain or analyze those thoughts(Ericsson and Simon 1993). Verbal analysis is a methodology for quantifying the subjective or qualitative coding of the contents of these verbal utterances(Chi 1997).Chi (1997) describes the specific technique for analyzing verbal data as consisting of several steps, excluding collection and transcription of verbal protocols. These steps, as followed in our research, are: defining the content of the protocols; segmentation of protocols; development of a coding scheme; coding the data and refining coding scheme if needed; resolving ambiguities of interpretation; and analysis of coding patterns.
The authors would like to thank all the general practitioners who participated in this study. The authors would also like to thank
Open Access This article is distributed under the terms of the Creative Commons Attribution Noncommercial License which permits any noncommercial use, distribution, and reproduction in any medium, provided the original author(s) and source are credited.
Hodgkin lymphoma (HL) represents one of the most common non-AIDS-defining cancers with an increasing incidence overtime. Clinically, patients present advanced stages of disease with extranodal involvement in the majority of cases. In the last years, significant improvements in the treatment of patients with HL and HIV infection have been achieved. In the lack of randomized trials, several phase II studies have showed that in the era of highly active antiretroviral therapy (HAART) the same regimens employed in HIV-negative patients with HL can be used in HIV setting with similar results. Moreover, in the last years the feasibility of high dose chemotherapy and peripheral stem cell rescue has allowed to save those patients who failed the upfront treatment. Finally, in the near future, a better integration of diagnostic tools (including PET scan), chemotherapy (including new drugs), radiotherapy, HAART, and supportive care will significantly improve the outcome of these patients.
The availability of highly active antiretroviral therapy (HAART) has led to improvements in immune status among HIV-infected persons, reducing AIDS-related morbidity and prolonging survival. However, despite the impact of HAART on HIV-related mortality, malignancies remain an important cause of death in the current era [1,2]. The use of HAART was also associated with reduced incidence of the two major AIDS-associated malignancies-Kaposi's sarcoma (KS) and high-grade non-Hodgkin lymphoma (NHL) [3]. However, among non-AIDS-defining cancers, an increased risk of Hodgkin lymphoma (HL), anal cancer, lung cancer, and hepatocarcinoma has been recently observed [4].
Although HL is included in the World Health Organization's categorisation of HIV-associated lymphomas [5,6], the relation between HIV infection, AIDS, and HL is unclear.
HIV-associated HL (HIV-HL) displays several peculiarities when compared with HL of the general population. First, HIV-HL exhibits an unusually aggressive clinical behavior, which mandates the use of specific therapeutic strategies and is associated with a poor prognosis. Second, the pathologic spectrum of HIV-HL differs markedly from that of HL in the general population [7,8]. In particular, the aggressive histological subtypes of classic HL (cHL), namely, mixed cellularity (MC) and lymphocyte depletion (LD), predominate among HIV-HL and the tumor tissue are characterized by an unusually large proportion of neoplastic cells, termed Reed-Sternberg (RS) cells [7]. Finally, despite the great improvement in chemotherapy and supportive care, optimal staging and treatment is still a matter of controversy.
In HIV-negative population of the western countries, HL is one of the commonest malignancies diagnosed in young adults with 6 cases per 100.000 inhabitants under 45 years of age occur each year [9], even if an increase in incidence rate in the last decade has been observed [10]. The epidemiology of HL is characterized by a peculiar age distribution pattern-a bimodal incidence curve with a first peak around the age of 30 and the second peak around the age of 50 years-that has been taken as suggestive of an infectious etiology.
In immune suppressed patients, HL occurs more frequently than in the general population of the same age and gender. Given the relative high frequency of HL in the population groups at high risk for HIV infection, epidemiological studies conducted during the first years of the HIV epidemic in North America and in Europe had difficulties in including HL in the spectrum of HIVassociated cancers. However, with the spread of the epidemic and longer survival of infected people, the impact of HL could be better recognized. All studies [4,[11][12][13][14][15][16][17][18][19][20][21][22] strongly support the evidence that HIV-infected persons have, overall, a 10-fold higher risk of developing HL than HIV-negative persons. HL in HIV-infected individuals is more frequent in patients with moderate immune suppression and this is in sharp contrast with KS and diffuse large B-cell NHL that typically arise in a more pronounced immunosuppression setting [4]. Thus, the epidemiological pattern of HL in HAART era substantially differs from those observed for KS or NHL-two neoplasms which drastically decreased after the introduction of HAART-and pose several new questions with regard of the relationship between degree of immunodeficiency, persistent viral infections and cancer. Of some interest is the recent observation of Powles et al. who have investigated the occurrence of cancers in a prospective cohort of 11,112 HIV-positive individuals, with 71,687 patient-years of followup [23].
Standardized incidence ratios (SIRs) were calculated using general population incidence data. The incidence of HL in the HIV cohort was higher than in the general population (SIR 13.85; 95% CI, 9.64 to 19.26). There was a significant increase in the SIRs across the three study periods (1983 to 1995: 4.5; 1996 to 2001: 11.1, and 2002 to 2007: 32.4). Multivariate analysis demonstrated that HAART was associated with an increased risk of disease (SIR 2.67; 95% CI, 1.19 to 6.02). Further multivariate modeling by class of antiretroviral agents showed that of the three classes of antiretroviral therapy, only the Non-Nucleoside Reverse Transcriptase Inhibitors were associated with a significant increase in the incidence of HIV-HL (HR 2.20; 95% CI, 1.03 to 4.69). As the overall effect of HAART is to increase the CD4 count level, it paradoxically could increase HL incidence. A potential mechanism emphasizes the role of the RS cells producing several growth factors that increased the influx of CD4 cell and inflammatory cells, which, in turn, provide proliferation signals for the RS neoplastic cells. One can imagine that in the case of severe immune suppression, leading to an unfavorable milieu, the progression of the RS neoplastic cells can be compromised [24][25][26]. In addition, HIV-HL is EBV-associated in almost all cases, in contrast to what was observed in the general population, in which this association is only observed in 20% to 50% according to histological type and age at diagnosis [27]. Usurpation of physiologically relevant pathways by EBV-encoded latent membrane protein 1 (LMP1) may lead to the simultaneous or sequential activation of signaling pathways involved in the promotion of cell activation, growth, and survival, contributing thus to most of the features of HIV-HL.
Whether this change affects its categorization as HL or whether it delays HL development is unknown.
In summary, the use of HAART has improved the immunity of HIV-infected persons, diminishing the risks of developing other cancers or other opportunistic infections and paradoxically increasing the risk of HL.
In the pre HAART era, HIV-HL displayed different pathological features in comparison with those of HL in HIV-negative patients [7]. In fact, HIV-HL was characterized by the high incidence of unfavorable histological subtypes (i.e., MC and LD) [7,8]. Among HIV-infected persons, MC was the most frequent HL subtype whereas nodular sclerosis (NS) was less frequent than in HIV-uninfected persons. For each HL subtype, incidence decreased with declining CD4 counts, but NS subtype decreased more precipitously than MC subtype, thereby increasing the proportion of MC subtype of HL seen in persons with HIV/AIDS. Thus, the greater proportion of MC and LD subtypes appeared specifically related to severe immune compromise in HIV, while in HAART era HIV-infected patients with modest immune compromise are more at risk for the development of the NS subtype [4]. CHL is a monoclonal lymphoid neoplasm, derived from B cells, composed of mononuclear Hodgkin cells and multinucleated Reed Stenberg (HRS) cells residing in a cellular microenvironment. HRS cells consistently express CD30 and CD40 and express CD15 in the majority of cases. CD20 is positive in a minority of neoplastic cells in 30-40% of cases, while the plasma cell-specific transcription factor Interferon Regulatory Factor 4 (IRF4) is consistently positive in HRS cells [5,8].
On the other hand, HIV-HL may exhibit special features related to the cellular background (presence of fibrohistiocytoid stromal cell proliferation) and the high number of the neoplastic cell. Both these features may pose relevant difficulties in diagnosing and classifying the disease precisely. In these cases, a large cell lymphoma of the anaplastic cell type must be excluded on immunohistologic ground [7,28].
A high frequency of EBV association has been shown in HL (80-100%) tissues from HIV-HL [29,30]. The EBV genomes in such cases have been reported to be episomal and clonal, even when detected in multiple independent lesions. The elevated frequency of EBV association with HIV-HL indicates that EBV probably does represent a relevant factor involved in the pathogenesis of HIV-HL. An etiologic role of EBV in the pathogenesis of HIV-HL is further supported by data showing that LMP-1 is expressed in virtually all HIV-HL cases [5,[28][29][30][31]. On these bases, HL in HIV-infected persons appears to be an EBV-related lymphoma expressing LMP1.
Finally, RS cells of classical HL-of HIV-negative patients represent transformed B-cells that originate from preapoptotic germinal center (GC) B-cells [32]. Most HIV-related HL cases express LMP1 and display the BCL6-/CD138+/MUM1 IRF4+ (for Multiple Myeloma-1 Interferon Regulatory Factor-4), phenotype, thus reflecting post-GC B cells [29,32]. The possible contribution of LMP1 to the loss of BCL6 expression seems plausible given that LMP1 can downregulate many B-cell specific genes [33]. Loss of B-cell identity occurs during the normal differentiation of a GC Bcell into plasma cell or memory B-cell.
Similar to that observed in HIV-NHL, one of the most peculiar features of HIV-HL is the widespread extent of the disease at presentation and the frequency of systemic "B" symptoms, including fever, night sweats, and/or weight loss >10% of the normal body weight. At the time of diagnosis, 70-96% of the patients have "B" symptoms and 74-92% have advanced stages of disease with frequent involvement of extranodal sites, with the most common being bone marrow (40-50%), and liver (15-40%) and spleen (around 20%) [8,[34][35][36]. HIV-HL tends to develop as an earlier manifestation of HIV infection with higher median CD4+ cell count. [8,[34][35][36]. The widespread use of HAART has resulted in substantial improvement in the survival of patients with HIV infection and lymphomas, due to the reduction of the incidence of opportunistic infections, the opportunity to allow more aggressive chemotherapy, and the less aggressive presentation of lymphoma in patients in HAART in comparison with those lymphomas diagnosed in patients who never received HAART [8,[34][35][36][37].
Within the Italian Cooperative Group on AIDS and Tumors (GICAT), we have collected data on 290 patients with HIV-HL. Two hundred and eighty-one patients (87%) were males and the median age was 34 years (range 19-72 years) and 69% of patients were intravenous drug users. The median CD4 cell count was 240/µL (range 4-1100/µL) and 57% of patients had a detectable HIV viral load.
MC was diagnosed in 53% of cases, followed by NS in 24% and LD in 14%. Advanced stages of disease were observed in 79% of patients and 76% had B symptoms. The overall extranodal involvement was 59% with bone marrow, spleen and liver involved in 38, 30, and 17, respectively. With the aim to evaluate the impact of HAART on clinical presentation and outcome of our patients, we split the series into two subgroups: in the first group we included those patients who received HAART since 6 months before the onset of HL (84 patients); in the second group we included those patients who never received HAART before the diagnosis of HL or less than 6 months (206 patients). Briefly, in comparison to never experienced HAART, patients in HAART before the onset of HL are older, have less B symptoms, a higher leukocyte, neutrophils count, and hemoglobin level. The following parameters were associated with a better overall survival (OS): MC subtype, the absence of extranodal involvement, the absence of B symptoms, and a prior use of HAART. Interestingly, three parameters were associated with a better time to treatment failure: a normal value of alkaline phosphatase, a prior exposure to HAART, and an international prognostic score less than 3 [38]. A similar study was carried out within the Spanish group GESIDA where the authors compared the clinical characteristics and outcome of 104 patients with HIV-HL, treated (83 patients) or not (21 patients) with HAART. No differences were found between the groups at baseline, but the complete remission (CR) rate was significantly higher in HAART group (91% versus 70%, P = .023). The median overall survival was not reached in HAART group and was 39 months in no-HAART group (P = .0089); the median disease-free survival (DFS) was not reached in HAART group and was 85 months in no-HAART group (P = .129). Factors independently associated with CR were a CD4 cell count >100 cells/µL and the use of HAART; CR was the only factor independently associated with OS [39].
The optimal therapy for HIV-HL has not been defined yet. Because most patients have advanced stages of disease, they have been treated with combination chemotherapy regimens but the CR rate remains lower than that of HL of the general population with the OS being approximately 1.5 years [8,[34][35][36]. Due to the low incidence of the disease, no randomized controlled trials have been conducted in this setting. However, several phase II studies have evaluated the feasibility and activity of different regimens. In a prospective trial, conducted within the GICAT between March 1989 and March 1992, 17 previously untreated patients with HIV-HL were treated with epirubicin, vinblastine, and bleomycin (EVB). Overall, CR was achieved in 53% of the total group, lasting a median of 20 months. The median OS for the group as a whole was 11 months and the 2-year DFS was 55% [40]. In an attempt to improve upon these results, from 1993 to 1997, a second prospective trial consisting of full dose EVB plus prednisone (EVBP regimen) and concomitant antiretroviral therapy (zidovudine or didanosine) was conducted. The results of this trial in which 35 patients were enrolled, showed a CR rate of 74% and a 3-year OS and DFS of 32% and 53%, respectively [41]. The AIDS Clinical Trials Group (ACTG) reported the results of a phase II study in 21 patients treated with ABVD chemotherapy for 4-6 cycles and primary use of G-CSF. Antiretroviral therapy was not used. The CR rate, on an intent to treat analysis, was 43% with an overall objective response rate of 62%. Median survival for all patients was 18 months [42]. Similar data have been reported in a small trial with only 8 patients enrolled [43]. The widespread use of HAART allows the use of more aggressive chemotherapeutic regimens. We used the Stanford V regimen, consisting of short-term chemotherapy (12 weeks) with adjuvant radiotherapy. From May 1997 to October 2001, 59 consecutive patients were treated in the framework of this prospective phase II study within the European Intergroup Study HL-HIV. Stanford V was well tolerated and 69% of the patients completed treatment with no dose reduction or delayed chemotherapy administration. The most important dose-limiting side effects were bone marrow toxicity and neurotoxicity. Eighty-one percent of the patients achieved CR and after a median followup of 17 months 33/59 (56%), patients are alive and disease-free. The estimated 5-year OS, DFS, and freedom from progression (FFP) were 59%, 68%, and 60%, respectively. Probability of FFP was significantly higher (P = .002) among patients with an international prognostic score (IPS) of <2 than in those with IPS >2, and the percentages of FFP at two years were 83% and 41%, respectively. Similarly, probability of OS was significantly different (P = .0004), and the percentage of survival at three years was 76% and 33%, respectively for IPS <2 and IPS >2 [44]. Within the German group, the very intensive BEACOPP regimen has been tested in 12 untreated patients with a 100% of CR rate but a high incidence of opportunistic infections [45]. Recently, the results of a large prospective phase II study with ABVD have been published. Within a cooperative network in Spain, 62 patients with HIV-HL received the standard ABVD plus HAART. The scheduled six to eight ABVD cycles were completed in 82% of cases. Six patients died during induction, 54 (87%) achieved a CR, and two were resistant. The 5-year OS and event-free survival (EFS) probabilities were 76% and 71%, respectively. The immunological response to HAART had a positive impact on OS (P = .002) and EFS (P = .001) [46]. Interestingly, there are some anecdotic cases of use of HAART with antineoplastic intent in HIV-NHL, especially in primary effusion lymphoma. Recently a case of longlasting response to HAART as the only therapy for HIV-HL has been reported suggesting the possibility to use this approach in selected cases [47]. Finally, within the GICAT we have recently concluded the accrual of 71 patients in a prospective phase II study aiming to evaluate feasibility and activity of a novel regimen including epirubicin, bleomycin, vinorelbine, cyclophosphamide, and prednisone (VEBEP regimen). Seventy percent of patients had advanced stages of disease and 45% had an IPS > 2. The CR was 67% and 2-year OS, DFS, TTF, and EFS were 69%, 86%, 59%, and 52%, respectively [48]. The results of the largest prospective studies are showed in Table 1.
Because HIV-HL like other HL may progress or relapse, the use of high dose chemotherapy and autologous stem cell transplantation (ASCT) has been tested in this setting. Several data from different groups, including the GICAT, have demonstrated the feasibility of this approach that can be considered the gold standard in the salvage setting [49][50][51][52][53]. Different conditioning regimens, including or not total body irradiation, have been tested. Recently, the AIDS Malignancy Consortium demonstrated in a multinstitutional trial that a regimen of a dose-reduced high-dose chemotherapy, including cyclophosphamide and busulfan and ASCT, is well tolerated and is associated with favorable DFS and OS probabilities for selected patients with HIV-associated NHL and HL [54].
Positron Emission Tomography using [18F]-Fluoro-2-Deoxy-D-Glucose (FDG-PET) was first introduced in the management of lymphomas in the early 1990s. It is now recognised as an important tool for staging and treatment response assessment in Hodgkin and non-Hodgkin lymphomas [55,56]. Turning to predicting outcome, in the HIVnegative patients, residual FDG PET avidity after 2 cycles of ABVD has been shown to confer poor prognosis and therefore, has been proposed to guide future therapy [57,58]. A negative PET scan after two cycles of ABVD predicted a 96% 2-year progression-free survival (PFS). Nearly 80% of the HL patients show a complete normalization of PET scan after two courses of ABVD [56]. This phenomenon, called "metabolic CR", can be explained by the peculiar architecture and organization of the neoplastic tissue, where only few, scattered neoplastic cells (accounting for less than 1% of the total cellular population) are surrounded by a population of nonneoplastic mononuclear bystander cells. The latter cells are probably responsible for the immortalisation of Hodgkin and RS cells by stimulating cytokine production by other CD4+ lymphoid cells (paracrine loop) or by inducing cytokine production by the RS cells (autocrine loop). In cases presenting with bulky lesions at diagnosis, a negative early PET is often associated with a persisting bulky lesion of more or less unchanged size. The explanation might be that chemotherapy switches off the production of chemokines by the activated lymphoid cells, as described for TARC (Thymus and activation-related chemokine). The latter can be measured in the serum of HL patients and its level is correlated to the quality of treatment response: for patients in CR the levels are much lower than in patients with stable or progressing disease [59]. In contrast, PET scanning within the HIV framework can be problematic. Some preliminary reports suggested that FDG activity may correlate with detectable lymphoma [60,61]. Although initial staging may not alter the treatment plan, it can provide additional information, assess possible involvement of critical location, and help foresee and possibly avoid further complications. However, the experience with PET scanning in the HIV-HL needs to be further studied. A baseline study is strongly recommended, since early PET interpretation is based on a site-to-site comparison of FDG uptake both before and after chemotherapy. Pitfalls are numerous and bring a particular challenge in these patients whom HIV-associated immunodeficiency predisposes to infection, as does the use of aggressive immunosuppressive chemotherapy regimens [62]. PET imaging requires cautious reading and pertinent clinical correlation to avoid diagnosing benign disease as malignant, such as, hypermetabolic foci seen in lung or oesophagus, which are common sites of HIV-and/or chemotherapypromoted infections. Nodal FDG uptake can be observed in lymphoma, various infections (e.g., Mycobacterium avium intracellular, Mycobacterium tuberculosis, Herpes simplex virus, among others), and AIDS-related malignancies such as Kaposi sarcoma. In addition, stimulation of bone marrow following treatment with granulocyte colony stimulating factors induces a striking increase in FDG uptake in bone marrow. To take into account the possibility of minimal residual uptake, a semiquantitative approach has recently been proposed for interim PET interpretation in the context of an international protocol for advanced-stage HL [63,64].
Finally, PET is useful for an accurate initial staging and it should be recommend to monitor treatment response, because PET appears to have a prognostic value, since a negative scan seems always associated with a favourable outcome. Significance of residual uptake at sites of disease needs further evaluation (e.g., biopsy). However, the use of FDG in the followup of HIV-HL patients who achieved CR cannot be routinely recommended and further studies are warranted prior to any definite conclusion.
The outcome of patients with HIV-HL has improved with better combined antineoplastic and antiretroviral approaches. The main important challenges for the next years are (a) to demonstrate in a randomized trial that ABVD is the standard regimen in HIV setting, (b) to validate the role of PET scan both in the staging and in the evaluation of response, (c) to better understand the interactions between chemotherapy and antiretroviral therapy, in order to reduce the toxicity of both approaches, (d) to evaluate the use of new drugs (i.e., bortezomib) in this setting, and (e) to evaluate the long-term toxicity of the treatment in cured patients.
This paper is supported in part by a grant from the
In recent years, organic thin-fi lm transistors (OTFTs) have attracted a great deal of attention due to their potential applications in low cost sensors, [ 1 ] memory cards, [ 2 ] and integrated circuits. [ 3 ] Great efforts are under way to design OTFTs with high performance, high stability, high reproducibility, and low cost. [ 4 ] Two of the most crucial device parameters are the charge carrier mobility and the threshold voltage ( V Th ). Concerning the mobility, the main goals for most applications is its maximization. [ 5 ] For V Th , the situation is more complex: for example, for integrated circuits it would be desirable to tune V Th over a broad range, [ 6 ] e.g., for inverter applications. In silicon technology, complementary circuits that consist of p-channel and n-channel transistors are typically used. [ 7 ] There have been many attempts to adapt this technology to OTFTs and fabricate organic complementary inverters. [ 2 , 8 , 9 ] They, however, suffer from poor n-type transistor performance and/or air instability of n-type semiconductor materials. An alternative approach is the adaptation of unipolar depletion-load inverters enabling simplifi ed processing, even if they do not provide the low power consumption and the simple circuit design intrinsic to complementary logic. [ 10 , 11 ] Depletion-load inverters consist of an enhancement-mode driver transistor and a depletionmode load transistor and can be realized using only p-type OTFTs. So far there have been attempts to achieve this target by using a level shifter [ 12 , 13 ] or a dual gate structure. [ 14 ] The main objective is to fi nd a reproducible method to realize driver and load transistors with equivalent device characteristics (in particular mobilities), but different V Th values.
Over the years, a number of methods have been developed to tune threshold voltages. For example, oxygen plasma treatment of an organic gate dielectric (parylene) results in a large shift of V Th to more positive values. [ 15 ] This shift is related to the generation of charged surface states at the dielectric-semiconductor interface but the exact mechanism is only poorly understood. The same holds for UV-ozone treatments [ 16 ] of the parylene and for the UV treatment of a poly-4-vinylphenol layer, [ 17 ] where the latter has been used to realize and eventually to optimize depletion-load inverters. V Th can also be tuned by changing the capacitance of the dielectric, [ 18 ] by inserting a polarizable layer into the dielectric, [ 19 , 20 , 21 ] or by using this polarizable layer as an encapsulation. [ 22 ] With the latter methods, V Th can be controllably and reversibly shifted over a wide range. The drawback of such approaches is, however, that a high "programming" voltage is needed to tune V Th . Another possibility to tune V Th is by insertion of self-assembled monolayers [ 23 , 24 ] or chemically reactive thin layers. [ 25 , 26 ] The latter have been shown to tune V Th by a local channel doping using acid groups that can be combined with de-doping reactions using bases. The disadvantage of the latter methods is that a local patterning of threshold voltages is not straightforward. Such a patterning is, however, a prerequisite for the realization of integrated electronic circuits.
Here, we signifi cantly refi ne the concept of chemical channel doping by replacing the covalently bonded silane layers bearing sulfonic acid groups used in Refs. [ 25 , 26 ] with photoacid generator polymers. [ 27 ] The goal hereby is to use them as an interface-modifi cation layer (see Figure 1a ), whose properties can be patterned photochemically, because, in contrast to the molecules used previously [ 25 , 26 ] , the acid group is formed only upon illumination. This paves the way for a photolithographic patterning of the interfacial doping and, thus, for controlling which transistors in a circuit operate in depletion or in enhancement mode. In other words, it enables an accurate local control of V Th through the illumination dose by a method that is fully compatible with lithographic techniques omnipresent in conventional semiconductor industry.
To show the versatility of our approach, we have synthesized two qualitatively different polymers, namely, poly(endo,exobicyclo[2.2.1]hept-5-ene-2,3-(2-nitrobenzyl) dicarboxylate) (PBHND) and poly-[(endo,exo-N-hydroxy bicyclo[2.2.1]hept-5ene-2,3-dicarboximide perfl uoro-1-butanesulfonate)-co-(endo,exobicyclo[2.2.1]hept-5-ene-2,3-dicarboxylic acid, dimethyl-ester)] (PHDBD). The chemical structures and the main reaction occurrence of stable protonated pentacene molecules appears unlikely for the present devices and the actual fate of the protons needs to be investigated further.
The results in Figure 1d show that varying the illumination time allows tuning the threshold voltage over several volts by increasing the UV illumination time from 0 to 5 seconds in a PBHND-containing OTFT. The shape of the curves is essentially the same for the different illumination times; in particular the slopes (and thus the mobilities) of the OTFTs remain constant. Also the drain current in the output characteristics (Figure 1e ) increases with illumination, consistent with an increased channel doping. The hysteresis of the transistors becomes somewhat larger but remains small even after 5 s of illumination. The gate voltage at a drain current of I D = 0.10 m A changes by Δ V G = 2 V. Also the off-current remains in the range of nA and is barely altered by the UV treatment. The devices, for which the data are shown in Figure 1e , were kept in an argon glove box throughout all process steps so no detrimental effects due to contact with O 2 or H 2 O occured. products upon UV illumination are shown in Figure 1b,c . From PBHND, a carboxylic acid is formed that remains chemically linked to the polymer backbone, while in PHDBD a sulfonic acid is part of the leaving group. Both photochemical reactions can be followed by IR spectroscopy (see Supporting Information).
In the device, the deprotonation of the acidic groups due to the reaction with the organic semiconductor results in the formation of a space-charge region at the interface. This gives rise to a threshold voltage shift, since mobile holes have to compensate the conjugated bases. This has been shown by driftdiffusion based modelling. [ 28 , 29 ] From the data in [ 29 ] it can be estimated that the deprotonation of only a few percent of the acid groups at the interface is suffi cient to shift V Th by several tens of volts. An alternative way to understand the underlying process suggested for a polythiophene active layer [ 25 ] is to view it as acid doping, in analogy with the situation in PEDOT/PSS. [ 30 ] Indeed, a photoinduced protonation of pentacene by acids has already been shown, albeit only at temperatures around 3.5 K with the product being thermally unstable. [ 31 ] Consequently, the PBHND the acid is linked directly to the polymer backbone, while it is on the leaving group in PHDBD.
An observed disadvantage of extended illuminations is that, in addition to the V Th shift, it results in some deterioration of the device characteristics. These effects, however, play a role only for illumination times far beyond the 3 s needed for the optimum inverter operation. They include an increase of the off current, a larger hysteresis, and a deterioration of the drain current at large negative gate bias. A detailed discussion can be found in the Supporting Information. All measurements (threshold voltage shift and stability) for the PHDBD devices were reproduced for several samples. The PHBND measurements were reproduced for two batches of freshly synthesized polymer and for various samples within one batch.
To verify that the observed effects are indeed a consequence of interfacial acid doping, two test experiments were performed. First, the devices were exposed to a strong base (NH 3 as in Refs. [ 25 , 26 ] ). As expected, this treatment results in a neutralization of the acid, i.e., a de-doping of the channel, and after 10 min of exposure to ammonia the devices were switched back to a situation similar to that prior to UV exposure. As a second test, devices containing no interfacial layer were also illuminated; here, even after 20 min of UV exposure no changes in the transfer and output characteristics were observed. More details about these experiments are included in the Supporting Information.
This provides us with all the tools necessary to perform the next step in the direction of integrated p-type organic electronic devices, which is the fabrication of a (tuneable) depletion-load inverter. The wiring diagram of such an inverter is plotted as an inset in Figure 3 . The switch transistor is realized by a non-illuminated PBHND-containing device that has a negative threshold voltage of around V Th = -10 V. The load transistor (also a device containing a PBHND interfacial layer) was illuminated in steps of 1 s in an argon glove box after the inverter wiring was realized. As seen in Figure 3 , the inverter characteristic is very poor prior to illumination, the gain of the inverter is negligible, and the achievable output voltage at the low level While the V Th shifts shown in Figure 1d are suffi cient for the realization of a depletion load inverter (vide infra), the impact of longer illumination times (see left part of Figure 2 ) was also explored. The time scales for illumination cannot be directly compared, as different lamps with different emission spectra were used for the two different polymers (see Experimental Section). The observed threshold voltages changes reach close to 50 V both for PBHND and PHDBD interfacial layers. However, the threshold values without illumination differ signifi cantly between the two photoacid generator polymers (only the PHBDB devices are actually switched from enhancement to depletion mode operation). Moreover, the V Th shifts for PHDBD interfacial layers are not stable for repeated measurement cycles, as shown in the right plot in Figure 2 . [ 32 ] For the PBHND-based OTFTs, even after more than 200 measurements the threshold voltage did not shift back considerably. The small decrease during the fi rst measurements was mostly due to a decrease of hysteresis in these devices. The hysteresis is at least in part a result of the need to briefl y transport the devices through air before the illumination (which, again, happened under inert atmosphere). One also has to keep in mind that the chosen measurement range for the device characteristics (from -60 V to 60 V) induces a signifi cantly larger amount of bias stress than in any inverter-based application. The small offset between the fi nal data point in the left plot and the fi rst data point in the right plot is a consequence of using devices from different samples. The difference in device stability for the illuminated PBHND-and PHDBD-containing devices can be attributed to the different positions of the acids formed in the photochemical process. As mentioned above, in growth of the pentacene could be reproducibly tuned by the illumination time. [ 34 ] This can, for example, be used for tuning the fi lm mobilities and is possibly a consequence of the interaction of the acid groups with ambient humidity. Such processes were avoided in the present experiments.
Here, for the post-growth illumination, the PBHND TFTs were illuminated in a range from 1 s to 300 s with an intensity of I UV = 3 W cm -2 with an EFOS (now EXFO, Mississauga, Canada) Novacure lamp. PHDBD TFTs were illuminated for 1 to 60 s with a 100 W polychromatic medium pressure mercury lamp from Heraeus (Hanau, Germany) in a Newport (Irvine, USA) model 66990 housing. For the latter experiments, the light intensity (power density) at the sample surface was measured with a spectroradiometer (Solatell (Stroud, UK), Sola Scope 2000TM, measuring range from 230 to 470 nm). The integrated power density for the spectral range 230nm-400 nm was 20.9 mW cm -2 .
The devices were characterized immediately after the illumination in an argon glove box with a Parametric Analyzer (Agilent Technologies (Santa Clara, USA) -E5262A). [ 35 ] Threshold voltages are extracted from the low-voltage (i.e., saturation) regime of the transfer characteristics as this is the most relevant voltage region for the inverter characteristics. [ 36 ] This is useful here in spite of the issues arising from carrier-densitydependent mobilities. For the fabrication of the inverters, the devices were illuminated directly in the argon glove box. The wiring was performed externally through two Parametric Analyzers (HP4145, mb technologies (Grosswilfersdorf, Austria)). This was necessary, as substrates with common gate structures were used, due to which the switch and load transistor had to be fabricated on different substrates. There is, however, no fundamental reason, why the two transistors cannot be fabricated next to each other on the same substrate, with a phototuning of V Th of the load transistor through a mask, when patterned gate structures are available.
is only -4 V. This is because the non-illuminated load transistor is normally off and works in enhancement mode. By increasing the illumination time of the load transistor and shifting its threshold towards the positive voltage regime, the inverter characteristic improves signifi cantly. The turn-on voltage of the inverter is shifted to more negative values and the gain of the inverter increases. After illuminating the load transistor for 3 s, an optimum value of V Th,load with respect to the threshold voltage of the switch-transistor is reached, resulting in a steep inverter transition with a maximum gain of about 40 (see bottom graph in Figure 3 ). Further increasing the illumination time results in a deterioration of the inverter performance. At this point it should be noted that no attempts for optimizing the inverter characteristics other than tuning V Th,load were made. Therefore, further signifi cant improvements can be expected by adapting the width-to-length ratio of the channel between the load and switch and by optimizing the performance of individual transistors with respect to mobility, gate leakage, etc.
In conclusion, we have demonstrated an easy and reproducible way to switch OTFTs from enhancement to depletion mode by a photochemical reaction using photoacid generators as interfacial layers and demonstrated that this allows the fabrication of good quality depletion-load inverters with tunable characteristics. The suggested fabrication method can easily be adapted to monolithically integrated circuits by illuminating individual transistors through shadow-or photomasks when patterned gate electrodes are used (in this way simultaneously reducing leakage and parasitic effects).
The two polymeric photoacid generators poly(endo,exo-bicyclo[2.2.1] hept-5-ene-2,3-(2-nitrobenzyl)dicarboxylate) (PBHND) and poly-[(endo,exo-N-hydroxy bicyclo[2.2.1]hept-5-ene-2,3-dicarboximide perfl uoro-1-butanesulfonate)-co-(endo,exo-bicyclo[2.2.1]hept-5-ene-2,3dicarboxylic acid, dimethyl-ester)] (PHDBD) were synthesized following the previously described procedure, [ 33 ] using nitrobenzylalcohol for monomer synthesis and a Grubbs initator for polymerization. Further details are given in the Supporting Information.
Highly p-doped silicon wafers with a 155-nm-thick layer of thermally grown SiO 2 were purchased from Siegert Consulting e.K. (Aachen, Germany) All wafers were O 2 -plasma etched for 30 s and then rinsed in deionized H 2 O in an ultrasonic bath for 120 s. Subsequently, the photoreactive polymers PBHND and PHDBD were spin-cast from a solution in tetrahydrofuran (4mg mL -1 ) at 1000 rotations per minute (rpm) for the fi rst 9 s and at 2000 rpm for the following 40 s. This yielded fi lms of 35 nm thickness for PBHND and 10 nm thickness for PHDBD, as determined by X-ray refl ectivity measurements. Afterwards, pentacene layers with an average thickness of 35 nm (measured with a quartzmicrobalance) were evaporated at a base pressure of 1 × 10 -5 mbar with the substrates held at room temperature. The fi rst 5 nm were evaporated at a rate of 0.02 A s -1 and the subsequent 30 nm at a rate of 0.1A s -1 . 50-nm-thick Au source and drain electrodes were deposited through a shadow mask at a base pressure of 4 × 10 -6 mbar. The resulting channel length and width were 50 μ m and 7 mm, respectively. On the left side of Figure 1 the whole device setup of the OTFTs is displayed.
The as-fabricated and fully functional OTFTs were exposed to UV light in helium atmosphere to avoid photo-oxidation of the polymers. In this context it needs to be mentioned that in another series of experiments it was found that when illuminating PBHND prior to pentacene deposition and subsequently exposing the devices to air before growing the semiconductor, the observed V Th shifts were comparably small but the
We acknowledge fi nancial support by the
Supporting Information is available from the Wiley Online Library or from the author.
Kava is traditionally consumed by South Pacific islanders as a drink and became popular in Western society as a supplement for anxiety and insomnia. Kava extracts are generally well tolerated, but reports of hepatotoxicity necessitated an international reappraisal of its safety. Hepatotoxicity can occur as an acute, severe form or a chronic, mild form. Inflammation appears to be involved in both forms and may result from activation of liver macrophages (Kupffer cells), either directly or via kava metabolites. Pharmacogenomics may influence the severity of this inflammatory response.
Kava (Piper methysticum Frost F.) is a perennial plant which has been used for centuries by South Pacific communities for medicinal, social, and cultural purposes. Traditionally, the rhizome of the plant is macerated with water or coconut milk to produce a beverage with relaxant and psychoactive properties [1]. In the 20th century, it became popular in the Western world as a herbal supplement for anxiety and insomnia.
More than 40 compounds have been isolated from kava, with the active components present in the lipid-soluble resin containing three chemical classes (i) arylethylene-α-pyrones, (ii) chalcones and other flavones, and (iii) conjugated diene ketones. It is the substituted 4-methoxy-5, 6-dihydro-αpyrones or kavapyrones commonly known as kavalactones that possess the highest anxiolytic effects [2].
Total kavalactone accounts for 3%-20% dry weight, with the highest concentration in the lateral roots, decreasing gradually towards the aerial plant structures. To date, eighteen kavalactones have been identified from the root, six of which account for approximately 95% of the organic extract namely, kavain, dihydrokavain, methysticin, dihydromethysticin, yangonin, and desmethoxyyangonin (Figure 1). When these kavalactones are eluted from a sample of kava by HPLC and sorted by decreasing order of quantity, the chemical signature obtained distinguishes individual strains (cultivars). Such chemotyping has identified over 200 variant strains of kava, but the chemical signature can vary between roots, rhizomes, and basal stems [3]. Root and leaf extracts from four Hawaiian cultivars (Mahakea, PNG, Purple Moi, and Nene) found that dihydrokavain and dihydromethysticin comprised more than 70% of the total kavalactones in leaves, whereas each of the six major kavalactones represented around 10% to 20% of the total kavalactones in root extracts [4].
Whilst kava extracts are generally well tolerated, toxic doses were determined in vivo using animal studies and in vitro. The LD 50 (mg/kg) for kavain, methysticin, dihydrokavain, and dihydromethysticin in mice ranges from 41-69 (intravenous), 325-530 (intraperitoneal), and 920-1130 (oral) [2]. F344 rats treated with 2 g/kg/day kava extracts in corn oil by oral gavage five days a week for fourteen weeks exhibited elevations in γ-glutamyl transferase (GGT), serum cholesterol, protein, and albumin levels, along with hypoglycaemia, within days of treatment [5]. Cytotoxicity associated with kavalactones was demonstrated for human hepatocytes in vitro (EC 50 values approximately 50 μM) [6] and neurones in vivo at ≥300 μM [7], and apoptosis is the mechanism of cell death [8]. Since the highest serum kavain concentration recorded in a human is 17.4 μM [9], these data suggest that for most individuals, kavalactones have a wide therapeutic index. Pipermethysticine (PM) is a toxic alkaloid present in kava leaves and stem peelings, which can contaminate kava products during high production and/or poor quality control [10]. HepG2 human hepatoma cells exposed to 100 μM PM for 24 hours showed 90% loss of cell viability, with 65% loss of viability with 50 μM PM due to disruption of mitochondrial function and consequent apoptosis [11]. Fischer-334 rats given 10 mg/kg PM daily by oral gavage for two weeks demonstrated adaptive changes to oxidative stress such as a significant increase in hepatic glutathione and cytosolic superoxide dismutase (Cu/ZnSOD) [12]. Thus, PM is a potential cause of kava hepatotoxicity although it was not detected in any samples in a recent study using products available on the German market [13].
Flavokavain B is a cytotoxic component of kava root and is present in aqueous and organic extracts [14]. The highest yields are in chloroform extracts, followed by acetone then hexane fractions [15]. Hepatocellular toxicity results from mitogen-activated protein kinase (MAPK) signalling leading to oxidative stress and apoptosis [16]. Due to its location in the plant root, flavokavain B is likely to be present in both traditional and pharmaceutical preparations of kava.
The chemical content of kava products and hence potential for adverse effects vary according to plant age, part used, cultivar, geographical location, and growth conditions. In the Kava Act 2002 [17] the government of Vanuatu classified the different cultivars of kava as noble (traditional social beverage with long history of safe use), medicinal (used for specific therapeutic effects), two days (can cause strong side effects such as nausea due to high levels of the kavalactone dihydromethysticin and banned for export), and wichmannii (which also induce strong side effects and are banned from export) [18].
Emphasis on the links between method of preparation and toxicity has declined recently. This is due to reports identifying adverse effects in patients using either ethanolic/acetonic kava extracts or traditional aqueous extracts, suggesting that toxicity is more dependent on the kava plant itself than the extraction solvent(s) used [19].
Meta-analyses of placebo-controlled studies showed a significant reduction in anxiety for patients receiving kava extract compared with patients receiving placebo [20,21]. The most common treatment regime used was 300 mg/day for four weeks. A study using a similar dose (280 mg/day for four weeks) found no difference between kava extract and placebo in terms of occurrence of adverse events, withdrawal symptoms, effect on heart rate, blood pressure, laboratory assessments, and sexual function, which supports the safety of this protocol [22].
However, two post-marketing trials in Germany (n > 3000) showed a small dose-dependent increase in adverse effects caused by kava extracts [23]. Risk factors for adverse reactions include chronic, heavy use and concurrent use with other drugs, herbs, and dietary supplements [24,25].
Side effects of kava consumption include skin reactions and central nervous system effects. Kava dermopathy has been well documented among Pacific Islanders [26] and is a reversible condition characterised by dry scaly yellow skin covering the palms of the hands, soles of the feet, and back [27]. It is speculated to have a link to cholesterol metabolism and hepatotoxicity due to the presence of jaundice [28]. Short-term kava use can produce extrapyramidal side effects resulting in oral dyskinesia and serious exacerbations of parkinsonism-like symptoms, while heavy use may predispose individuals to seizures [27]. However, a doubleblind placebo controlled trial showed kava has no effect on motor vehicle performance. In another study, no statistically significant difference between kava and placebo was present in participants performing tracking tasks [27].
The most serious adverse reaction associated with kava is hepatotoxicity. Reports of kava hepatotoxicity first emerged in Germany in 1998, and by the end of 2005, the World Health Organisation had received 91 reports of 189 adverse reactions relating to kava-only products. Fiftly-five of those reactions involved liver and biliary system disorders, including three cases of hepatic failure and two cases of hepatic comas [29]. Reported daily doses ranged from 45-1200 mg kavalactones taken for one week to twelve months [25].
The main metabolic pathways for kavalactones in humans and rats are hydroxylation of the C-12 in the aromatic ring, breaking and hydroxylation of the lactone ring with subsequent dehydration, reduction of the 7,8-double bond, and demethylation of the 4-methoxyl group [30][31][32]. Products of kavain metabolism found in human serum and urine include p-hydroxykavain, p-hydroxy-5,6-dehydrokavain, p-hydroxy-7,8-dihydrokavain, 5,6-dehydrokavain, 6-phenyl-5-hexen-2,4-dione [33], and 6-phenyl-3-hexen-2-one [34]. Human metabolites identified for other kavalactones include 11,12-dihydroxykavain-o-quinone for methysticin, 11,12dihydroxy-7,8-dihydrokavain-o-quinone for 7,8-dihydromethysticin and 12-desmethylyangonin either from demethylation of the 12-methoxyl group of yangonin or hydroxylation at C-12 of desmethoxyyangonin [30,35]. In rats, approximately 50% to 75% of administered kavalactones are excreted in the urine, mostly as glucuronide and sulphate conjugates. Approximately 15% is excreted in the bile [30,32,36]. Reactive kava metabolites may potentially alkylate DNA or disrupt enzymatic and metabolic activity, inducing hepatotoxicity.
Despite the evidence for liver damage and inflammation in animals and humans treated with kava, clinical cases of hepatotoxicity amongst indigenous users caused by traditional aqueous root extracts are limited to two cases in New Caledonia [37]. The dose of kavalactones in one patient was 18 grams per week for four to five weeks and was unknown in the other. Clinical surveillance in the Northern Territory, Australia over 20 years has not documented any cases of fulminant hepatic failure attributable to kava, despite doses estimated to be 10-50 times the recommended therapeutic doses.
Based on these observations, most hepatic effects of kava appear to be reversible or can be compensated for. Hence, pharmacogenomic effects may be involved in the most severe cases of toxicity. For example, the hydroxylation of the aromatic ring and demethylation of kavalactones is a function of CYP2D6 enzymes [30]. Four human phenotypes for CYP2D6 activity (ultrarapid, efficient, intermediate, and poor) are defined according to debrisoquine metabolism. The two Europeans with kava-related hepatotoxicity were phenotyped as poor metabolisers according to CYP2D6 activity [38]. It is known that 12%-21% of Caucasians are poor metabolisers compared to less than 1% for Asians/Pacific Islanders, which may contribute to the lower incidence of kava hepatotoxicity in the Pacific [39]. Caucasians also have a higher frequency of ultrarapid metabolisers compared to Asians/Pacific Islanders (1% to 5% versus 0% to 2%). These individuals could experience adverse reactions following a burst of reactive kava metabolites.
Acute kava hepatotoxicity involves inflammation. Histopathology results from a 50-year-old man [40] and a 33-year-old woman [38] displaying hepatic symptoms following use of kava supplements showed extensive, severe hepatocellular necrosis and infiltration with lymphocytes, eosinophils, and activated macrophages. Experimental studies with kavain-perfused rat livers examined via electron microscopy displayed a disruption of hepatic vasculature with narrowing of blood vessels, constriction of sinusoidal Advances in Pharmacological Sciences blood vessels, and retraction of the endothelium compared to controls [41]. Liver macrophages (Kupffer cells) within the sinusoids of the kavain-perfused liver also appeared swollen with large cytoplasmic vacuoles and phagocytosed material.
Subclinical liver abnormalities occurred in clinical studies with chronic, indigenous, kava users. A study with Australian aborigines in the Northern Territory revealed elevations in γ-glutamyl transferase (GGT) and alkaline phosphatase (ALP) in 61% and 50%, respectively, of participants who reported using kava at least once in the month prior to the measurement, compared to less recent users and nonusers [42]. These enzymes return to normal after one to two months of abstinence [43]. There were no differences between the groups in ALT, bilirubin, albumin, or total protein. Unlike the German clinical trials for anxiety, kava consumption was not regulated in these studies, and the median duration of kava use was twelve years, with a range from one to eighteen years. Similar results were observed in a predominantly Tongan population in Hawaii, which compared liver function tests between 31 healthy adult kava beverage drinkers and 31 healthy adult nonkava beverage drinkers [44]. GGT was significantly elevated in 65% of the kava drinkers versus 26% in the controls, and ALP was significantly elevated in 23% of kava drinkers versus 3% in the controls. There was no significant difference in ALT, AST, bilirubin, albumin, or total protein. Increases in GGT and ALP without ALT or AST elevation are suggestive of cholestasis rather than hepatocellular damage.
Cholestasis can be due to either defective bile formation in hepatocytes or disruption to bile secretion and flow within bile ducts [45]. Mechanisms of noninflammatory cholestasis include inhibition of cellular proteins and transporters. Direct or indirect activation of Kupffer cells and their subsequent release of pro-inflammatory cytokines, growth factors, and reactive oxygen species is a cause of inflammatory cholestasis, which may involve hepatocytes or bile ducts [46]. In the absence of hyperbilirubinaemia, elevations in GGT and ALP are more likely to be due to bile duct inflammation. Each of these mechanisms could be precipitated by kava extracts and are possible explanations for the cholestasis observed in chronic, indigenous, kava users.
Hence, Kupffer cells could be involved in both acute and chronic kava hepatotoxicity. These cells are also implicated in the pathogenesis of many other liver conditions, including fibrosis, viral hepatitis, steatohepatitis, alcoholic liver disease, and activation or rejection of the liver during transplantation [47,48]. In animal studies, depletion of the Kupffer cell population is hepatoprotective during ischemia repurfusion injury [49], sepsis [50,51], radiotherapy [52], and dietinduced steatosis and insulin resistance [53]. Kupffer cell depletion experiments could be a valuable tool for determining their role in kava hepatotoxicity.
Furthur studies could investigate the extent of kavalactone metabolism in humans and the proportions of urinary versus biliary excretion of kavalactones and their metabolites at different time points after ingestion. Such studies would further define normal kavalactone metabolism, suggest target cytochrome P450 enzymes for pharmacokinetic herbdrug interactions with kavalactones, and aid the identification of toxic metabolite(s) and their contribution to kava hepatotoxicity. Similar animal studies could be conducted with the alkaloid Pipermethysticine and the chalcone Flavokavain B. CYP450 profiling of individuals involved in such studies could clarify pharmacogenomic differences in kava metabolism and could be performed either genotypically or phenotypically using substrates such as debrisoquine for CYP2D6 in humans [54].
The aim of the study was to evaluate the impact of a high-intensive exercise program containing high-intensive functional exercises implemented to real-life situations together with group discussions on falls and security aspects in stroke subjects with risk of falls. This was a pre-specifi ed secondary outcome for this study. For evaluation, Short Form-36 (SF-36) health-related quality of life (HRQoL) and the Geriatric Depression Scale-15 (GDS-15) were used. This was a single-center, single-blinded, randomized, controlled trial. Consecutive Ն 55 years old stroke patients with risk of falls at 3 -6 months after fi rst or recurrent stroke were randomized to the intervention group (IG, n ϭ 15) or to the control group (CG, n ϭ 19) who received group discussion with focus on hidden dysfunctions but no physical fi tness training. The 5-week high-intensive exercise program was related to an improvement in the CG in the SF-36 Mental Component Scale and the Mental Health subscale at 3 months follow-up compared with baseline values while no improvement was seen in the IG at this time. For the SF-36 Physical Component Scale, there was an improvement in the whole study group at 3 and 6 months follow-up compared with baseline values without any signifi cant changes between the IG and CG. The GDS-15 was unchanged throughout the follow-up period for both groups. Based on these data, it is concluded that high-intensive functional exercises implemented in real-life situations should also include education on hidden dysfunctions after stroke instead of solely focus on falls and safety aspects to have a favorable impact on HRQoL.
review, the authors state that too few studies have been done to explore any effects of physical fi tness training on mood (4). Improvement in activities of daily living (ADL) is known to enhance long-term health-related QoL (HRQoL) (5). According to the most recent Cochrane review on depression and exercise, physical activity seems to decrease depressive symptoms in people with a diagnosis of depression (6).
Health and QoL, the two components of HRQoL, have been linked together in many defi nitions but there is no universal defi nition for either QoL or HRQoL. The WHO defi nition of health is: " Health is a state not merely the absence of disease or infi rmity " (7). This is a well-cited and useful defi nition and it is Background A stroke is a life-breaking event that most often hits with no warning, giving no time for preparation for a new way of living the everyday life. It has been established that stroke has direct and/or secondary effects on the major aspects of health (physical, physiological and social) (1,2).
Moderate and high levels of physical activity have been showed to be associated with reduced risk of ischemic and hemorrhagic stroke (3). Physical activity has not been proven to enhance the quality of life (QoL) after stroke according to the Cochrane review on physical fi tness training for stroke patients (4). In the same Advances in Physiotherapy, 2010;12: 125-133 (in the IG) or group discussion about hidden dysfunctions after stroke and how to cope with these diffi culties (in the CG) on the HRQoL and on the presence of depressive symptoms among individuals with stroke and risk of falls.
This randomized controlled intervention trial, designed for individuals with stroke and risk of falls, is described in the accompanying paper and is registered at www. clinicaltrials.gov as NCT00377689.
Inclusion criteria were: current stroke, age Ն 55, fall risk assessed through clinical observations by an experienced PT, the ability to walk 10 m with or without a walking device, and the ability to understand and comply with instructions in Swedish. Individuals were excluded if they had the ability to walk outdoors independently (i.e. without assistance or walking device), severe aphasia, severe vision or hearing impairment, any medical condition that a physician determined was inconsistent with study participation, and living too far away ( Ͼ 100 km) from the training facilities.
Inclusion to the study was done 3 -6 months after stroke onset. To identify all potentially eligible individuals with stroke, 391 individuals were consecutively screened during inpatient rehabilitation at the Ume å Stroke Unit (Figure 1). After this initial screening, all eligible participants were contacted. Those still eligible for participation in the study a more thorough assessment was performed at the outpatient Clinical Research Center at the Ume å University Hospital. The fi nal judgment of participation in the study was made based on the inclusion/exclusion criteria during this assessment that was considered as the individual ' s baseline assessment. Figure 1 describes the screening/inclusion process.
The patient enrolment during the inclusion/exclusion assessments was followed by the randomization procedure. The two main investigators (EH and PW) were responsible for the randomization into the IG or CG. This was conducted with a minimization software program, MiniM (18) to avoid imbalances at baseline between the two groups. Two variables were taken into account: cognition, using the Mini Mental State Examination (MMSE Յ 24/ Ն 25) (19), and fall risk, using the Fall Risk Index (value Յ 1/ Ն 2) (20). applicable for everyone. There are many different definitions of HRQoL. According to the Swedish manual of the Short Form-36 (SF-36), HRQoL means a pragmatic delimitation and concerns mainly function and well-being during illness and treatment (8). According to the WHO, depression is a common mental disorder that presents with depressed mood, loss of interest or pleasure, feelings of guilt or low self-worth, disturbed sleep or appetite, low energy, and poor concentration (9). These problems can become chronic or recurrent and lead to substantial impairments in an individual ' s ability to take care of his or her everyday responsibilities (10). Depression also adversely affect adherence to treatment for other diseases and is among the leading causes of disability worldwide (10). Post-stroke depression is a common complication and has a prevalence of up to 33% (11,12). Data from Riks-Stroke, the Swedish Stroke Register, shows that women are more likely to be depressed than men at 3 months after stroke onset (13). Jorgensen et al. (14) showed in their study from 2002 that depressive symptoms predict falls after stroke. Depression is a common and important complication after stroke but is unclear (11).
A recently published review by Blake et al. concludes (15) that exercise intervention exerts a clinically relevant effect on depressive symptoms in older people (15). However, they do point out a gap of consistent results for the medium-term (3 -12 months) effect of exercise intervention on depression or depressive symptoms (15). A task-specifi c intervention designed to improve gait speed has been showed to have secondary ben efi ts by positively impacting depression, mobility and social participation for people post-stroke (16).
One could expect that many dimensions of a poststroke individual could be affected, not only the physical part. The psychological aspects are important for the entire rehabilitation process and for the individuals ' outcome of the same process. It is therefore important to preserve both functioning and well-being of people with stroke. Thus, there is a paucity of data from exercise intervention studies, on whether exercise/rehabilitation programs also have an effect on HRQoL and depression in elderly persons post-stroke.
The accompanying paper with the fi rst report from this intervention study evaluated the impact of a high-intensive exercise program after stroke on function and activity performance (17). The program could be benefi ciary in ADL 6 months after the intervention ended and for the Falls Effi cacy Scale International (FES-I) directly after and 3 months post-intervention for the intervention group (IG) compared with the control group (CG) ( p Ͻ 0.05).
The purpose of the present study was to evaluate the impact of a
5-week high-intensive exercise program The Consort Flowchart Assessed for eligibility (n=395) Excluded in total (n=361) Excluded patients with stroke (n=231) Not meeting inclusion criteria (n=186) Refused to participate (n=27) Other reasons (n=18) Analyzed (n=15) Excluded from analysis (n=0) Intention-to-treat Lost to 3 & 6 month follow-up (n=1) 1 subject deceased Discontinued intervention (n=1) 1 subject deceased Allocated to intervention (n=15) Received allocated intervention (n=15) Lost to 3 month follow-up (n=1) 1 subject on vacation, Lost to 6 month follow-up (n=1) 1 subject, no reason given Discontinued control intervention (n=0) Allocated to control (n=19) Received allocated control intervention (n=19) Analyzed (n=19) Excluded from analysis (n=0) Intention-to-treat Allocation Analysis Follow-Up Enrollment 3 to 6 month post stroke Randomization Patients with stroke (n=265) Figure 1. Screening process, from stroke onset to fi nal inclusion in the study.
Information about the study was communicated orally and via written materials to all potentially eligible individuals during inpatient rehabilitation, screening phone calls and at the time of inclusion in the study. All individuals provided informed written consent for study participation at the baseline assessments. The study protocol was approved by the local Ethics Committee for Human Research at Ume å University (Dnr 04-022).
A protocol was completed at baseline concerning personal characteristics, living situation, medical history, medication and history of falls. All assessments, baseline, immediately after intervention and 3-month follow-up, were performed at the outpatient Clinical Research Center. The 6-month follow-up assessment was conducted by telephone interview using those instruments with the possibility to administer through telephone, e.g. the HRQoL questionnaire (SF-36) and the Geriatric Depression Scale-15 (GDS-15). A questionnaire was used to evaluate compliance with the home exercise program for the IG. The CG received a questionnaire for evaluation of the educational sessions. Both questionnaires included space for personal comments and evaluation of the intervention program by the participants, and were delivered to the assessment personnel at the 3-month follow-up.
The nurses and physiotherapist (PT) who performed the clinical test assessments were blinded to group allocation. The participants were instructed not to reveal anything from their 5 weeks in the study at the different assessment times. If the staff had any suspicion as to which group the participant belonged to they were told to fi ll out an incidence form. All participants were blinded as for the content of the two different groups before randomization. They only knew that the two groups met different number of days per week. At the time of the study, participants were open to their own group content but not the other group.
Subjects in each group participated in a 5-week intervention program at a clinic. For the IG, the program consisted of seven sessions a week divided over 3 days with individualized group training, supervised by a PT; the focus was on physical activity and functional performance. They also received one session a week for 1 h with educational group discussions about fall risk and security aspects, led by a PT and an Occupational Therapist (OT). The control subjects program consisted of one session a week for 1 h each during the 5-week period. The session was an educational group discussion session led by one OT. The discussions were about hidden dysfunctions after stroke and how to cope with these diffi culties. The different themes discussed were chosen on the basis of being relevant to the individual ' s situation and interesting for them to participate in. The themes included communication diffi culties, fatigue, depressive symptoms, mood swings, personality changes and dysphagia. There was, however, no special focus on the risks of falling in these discussions. The intervention program for the IG is described in somewhat more detail in the accompanying paper (17).
The outcome measures in this study were HRQoL as measured by the SF-36 and symptoms of depression as measured by the GDS-15. The instruments used in this study have been tested for validity and reliability in populations similar to the population in this study (21 -24).
HRQoL was assessed with the SF-36, which is a generic instrument measuring self-reported physical and psychological aspects of health (21,22,25). The SF-36 includes eight subscales: Physical Functioning (PF), Role Functioning-physical (RP), Bodily Pain (BP), General Health (GH), Vitality (VT), Social Functioning (SF), Role Functioning-emotional (RE) and Mental Health (MH). The total score in each subscale is 100, which indicates a higher degree of perceived health. Besides the eight different scales, there are two dimensions, a Physical dimension (PCS) and a Mental dimension (MCS). The dimensions are calculated by weighing the different load of the eight scales in to these dimensions. These eight subscales are considered to be universal and to represent basal human function and well-being (8). A difference of fi ve points is considered to be of a clinical signifi cance in 23). The results for the norm population of Sweden in the eight subscales of SF-36 are shown for the age group 75 ϩ , since the mean age for this entire study is 78. Regarding the two dimensions, PCS and MCS, the norm population compared is in another age group, 75 -79, since there is no norm fi gure for the group 75 ϩ consolidated, only 65 ϩ (26).
Symptoms of depression were assessed with the GDS-15, which is a basic screening measure for depression in older adults (24). The GDS-15 screens for depression using 15 questions with a yes/no answer alternative. Depending on age, education and complaints, the score 0 -4 indicates normal (no depression), 5 -8 mild depression, 9 -11 indicates moderate depression and 12 -15 a severe depression.
Power (80%) was calculated on the Berg Balance Scale (BBS) (27,28) for the original study (17), which was the primary outcome measure to determine the sample size needed. The estimation was set to detect a signifi cant difference ( p ϭ 0.05, two-tailed test) of Ն 5 points in BBS . All analyses were performed according to the intention-to-treat principle (29). Descriptive statistics are presented in frequency as means Ϯ SD in Table I. Groups were compared at baseline using the chi-squared or independent samples t -test. For chisquare test, either Pearson chi-square was used or, in applicable cases, Fisher ' s exact test. The participants ' data were used for analysis for as long as they participated in the study. Generalized estimating equations with repeated measure statistics were used to analyze the data over time and taken into consideration the fact that each individual had multiple assessments (Table II). All data analyses were performed using the SPSS software package, version 17.0.
The study included 34 participants, 15 subjects in the IG and 19 subjects in the CG. There were no signifi cant differences in the baseline characteristics of the two groups (Table I). There was a success of blinding the group allocation to the clinical test assessment staff.
All but one participant completed the 5-week intervention period. Two participants dropped out during follow-up; the reason for dropout was worsening overall medical condition in both cases. The participants in the IG participated in the home exercise program two or three times per week according to the self-reporting questionnaire.
Table II. Repeated measure analyses for all outcome measurements at all follow-up assessments; SF-36 and GDS-15. Baseline Post-Intervention 3 months post-intervention 6 months post-intervention Norms for the Swedish population aged 75ϩ Total study population, nϭ34 IG, nϭ15 CG, nϭ19 CI Total study population, nϭ33 IG, nϭ14 CG, nϭ19 CI Total study population, nϭ31 IG, nϭ13 CG, nϭ18 CI Total study population, nϭ31 IG, nϭ13 CG, nϭ18 CI SF-36 PCS 40.1Ϯ12.7 30.8Ϯ9.6 30.9Ϯ8.3 30.8Ϯ10.7 Ϫ6.9 to 6.8 32.8Ϯ11.3 32.2Ϯ10.6 33.2Ϯ12.0 Ϫ7.2 to 9.2 35.3Ϯ13.3 35.5Ϯ14.7 35.2Ϯ12.7 Ϫ10.4 to 9.8 35.3Ϯ12.8 35.3Ϯ13.3 35.4Ϯ12.9 Ϫ9.7 to 9.7 MCS 49.0Ϯ11.4 53.2Ϯ9.4 53.6Ϯ10.0 52.8Ϯ9.2 Ϫ7.5 to 5.9 54.6Ϯ10.0 54.4Ϯ10.3 54.8Ϯ10.0 Ϫ6.9 to 7.7 53.3Ϯ10.1 48.7Ϯ12.7 * 56.7Ϯ6.2 * 0.9 to 15.0 * 53.3Ϯ12.0 50.4Ϯ15.0 55.4Ϯ9.3 Ϫ3.9 to 13.9 SF-36 PF 59.0Ϯ30.1 45.8Ϯ20.8 45.7Ϯ21.7 45.8Ϯ20.6 Ϫ14.8 to 14.9 48.5Ϯ24.6 52.1Ϯ22.2 45.8Ϯ26.6 Ϫ24.2 to 11.5 52.1Ϯ26.1 56.5Ϯ25.5 48.9Ϯ26.8 Ϫ27.2 to 11.9 48.7Ϯ23.6 51.5Ϯ18.6 46.7Ϯ26.9 Ϫ22.6 to 12.9 RP 49.3Ϯ43.2 28.7Ϯ38.0 25.0Ϯ36.6 31.6Ϯ39.8 Ϫ20.5 to 33.6 33.3Ϯ37.9 21.4Ϯ27.5 42.1Ϯ42.5 Ϫ5.9 to 47.2 41.1Ϯ39.0 28.9Ϯ39.3 50.0Ϯ37.4 Ϫ7.3 to 49.6 47.6Ϯ41.5 44.2Ϯ41.0 50.0Ϯ42.9 Ϫ25.6 to 37.1 BP 63.2Ϯ30.2 64.6Ϯ26.6 62.0Ϯ19.5 66.6Ϯ31.5 Ϫ14.3 to 23.5 72.2Ϯ28.1 67.9Ϯ28.0 75.4Ϯ28.5 Ϫ12.9 to 27.8 70.7Ϯ33.2 66.2Ϯ32.4 74.0Ϯ34.4 Ϫ17.2 to 32.8 74.8Ϯ30.5 70.6Ϯ31.2 77.8Ϯ30.4 Ϫ15.7 to 30.1 GH 59.8Ϯ24.0 56.3Ϯ23.7 57.7Ϯ20.1 55.1Ϯ26.6 Ϫ19.5 to 14.2 56.3Ϯ26.8 57.8Ϯ28.8 55.3Ϯ26.1 Ϫ22.1 to 17.1 64.4Ϯ25.3 61.7Ϯ26.0 66.3Ϯ25.4 Ϫ14.5 to 23.7 62.0Ϯ25.1 60.8Ϯ24.1 62.8Ϯ26.4 Ϫ16.9 to 21.0 VT 54.2Ϯ30.0 50.6Ϯ18.7 51.0Ϯ14.7 50.3Ϯ21.8 Ϫ14.1 to 12.6 57.1Ϯ24.1 59.3Ϯ23.3 55.5Ϯ25.1 Ϫ21.3 to 13.7 53.2Ϯ23.5 46.2Ϯ19.0 58.3Ϯ25.7 Ϫ5.0 to 29.4 51.8Ϯ23.1 46.7Ϯ21.6 55.6Ϯ24.0 Ϫ8.3 to 26.0 SF 79.1Ϯ26.2 79.4Ϯ25.5 80.0Ϯ26.6 79.0Ϯ25.4 Ϫ19.3 to 17.2 86.0Ϯ23.5 84.8Ϯ29.1 86.8Ϯ19.3 Ϫ15.2 to 19.2 88.3Ϯ18.2 83.7Ϯ23.6 91.7Ϯ12.9 Ϫ5.5 to 21.5 90.3Ϯ20.3 85.6Ϯ25.9 93.8Ϯ15.0 Ϫ6.9 to 23.3 RE 64.0Ϯ41.8 81.4Ϯ35.0 73.3Ϯ42.2 87.7Ϯ27.7 Ϫ10.1 to 38.9 79.8Ϯ35.3 71.4Ϯ38.9 86.0Ϯ32.0 Ϫ10.7 to 39.7 81.7Ϯ35.3 66.7Ϯ43.0 92.6Ϯ24.4 1.1 to 50.8 83.0Ϯ34.3 71.8Ϯ40.5 90.7Ϯ27.6 Ϫ6.0 to 43.9 MH 76.1Ϯ23.1 77.2Ϯ17.3 82.1Ϯ13.6 73.3Ϯ19.2 Ϫ20.8 to 3.1 81.6Ϯ17.7 84.6Ϯ12.4 79.4Ϯ20.8 Ϫ18.0 to 7.6 78.7Ϯ18.1 74.5Ϯ21.6 * 81.8Ϯ14.9 * Ϫ6.1 to 20.7 79.0Ϯ17.8 81.2Ϯ11.9 77.3Ϯ21.2 Ϫ17.3 to 6.5 GDS-15 n/a 3.0Ϯ2.1 2.5Ϯ1.7 3.4Ϯ2.3 Ϫ0.6 to 2.4 4.2Ϯ2.7 3.1Ϯ2.1 5.0Ϯ2.8 Ϫ0.3 to 3.6 3.3Ϯ2.0 3.2Ϯ1.2 3.4Ϯ2.5 Ϫ1.4 to 1.7 3.4Ϯ2.4 3.0Ϯ1.5 3.7Ϯ2.9 Ϫ1.4 to 2.5 Results are presented as meanϮSD. GDS-15, Geriatric Depression Scale; SF-36, Short Form 36; IG, intervention group; CG, control group; CI, confi dence interval; PCS, Physical Component Scale; MCS, Mental Component Scale; PF, Physical Functioning; RP, Role Functioning-physical; BP, Bodily Pain; GH, General Health; VT, Vitality; SF, Social Functioning; RE, Role Functioning-emotional; MH, Mental Health.
Table I. Baseline characteristics of the participants. Intervention group, n ϭ 15 Control group, n ϭ 19 Sex (M/F) 9/6 12/7 Age 77.7 Ϯ 7.6 79.2 Ϯ 7.5 mRS a 2.1 Ϯ 0.6 2.1 Ϯ 0.6 Inpatient rehabilitation, days at stroke unit 12.5 Ϯ 5.0 10.9 Ϯ 5.3 Days from stroke onset to study start 139.7 Ϯ 37.3 126.8 Ϯ 28.2 Diagnosis of depression 3 2 Use of medication, SSRI or other anti-depressants 6 4 Use of sleeping pills 4 6 Home-help service 5 9 MMSE b 26.3 Ϯ 3.5 25.5 Ϯ 4.4 Fall risk index (19,37) No 14 16 Low 1 0 Medium 0 3 High 0 0 Results are presented as proportion or mean Ϯ SD. a mRS, modifi ed Rankin Scale. b MMSE, Mini Mental State Examination. SSRI, Selective Serotonin Reuptake Inhibitors.
Signifi cant difference between the two groups were detected in SF-36 MCS and MH subscale at the 3-month follow-up in favor of the CG ( p ϭ 0.02). There were no differences between the IG and the CG in presence of depressive symptoms as measured by GDS-15.
The PCS levels are lower than in the norm population for this age bracket (norm population ϭ 40.1, total study population ϭ 30.8) (Table II). There was no difference in SF-36 PCS at baseline between the IG and the CG. At the 3-and 6-month follow-up, there was partial normalization of the SF-36 PCS for the whole study group (IG ϩ CG) vs values at baseline ( p Ͻ 0.05, Table II). However, there was no signifi cant difference for SF-36 PCS between the IG and CG over time (Table II). For the individual components of the SF-36 PCS, PF, RP, BP and GH, there were no signifi cant differences between the IG and the CG at baseline or over time post-intervention (Table II).
There is a signifi cant difference in SF-36 MCS between the IG and CG at 3 months follow-up in favor for the CG ( p ϭ 0.02, Table II). The subscale MH also showed a similar signifi cant difference at the 3-month follow-up ( p ϭ 0.02). The results for the remaining individual components of the SF-36 MCS, VT, SF and RE were without signifi cant differences between IG and CG at baseline and over time after intervention (Table II).
There were no difference in SF-36 MCS at baseline between the IG and the CG. The MCS levels in whole study group (IG ϩ CG) were higher than in the norm population for this age bracket (53 vs 49 at baseline measurements, Table II). In the whole study group (IG ϩ CG), there was no difference over time during the study period (Table II).
There was no statistically signifi cant difference in GDS-15 between the IG and CG at baseline. There was no assured difference over time in the GDS-15 between the IG and CG. All levels of GDS-15 are without signs of depression.
This 5-week high-intensive exercise program did have an impact on HRQoL. The CG had a favorable outcome in the SF-36 MCS and MH subscale at 3 months post-intervention while no such improvement was observed in the IG. This was the only difference achieved between the two groups regarding HRQoL. For the SF-36 PCS, there was an improvement in the whole study group at 3 and 6 months postintervention compared with baseline values without any signifi cant changes between the IG and the CG. The presence of depressive symptoms was unchanged throughout the follow-up period for both groups.
The SF-36 MCS was higher for the whole study group compared with the norm population at baseline and directly post-intervention. This means that the stroke subjects in this study at this time point were mentally more vital compared with the general unaffected population. This could be an effect that the stroke patients were still in a stage of recovery and after all had survived. It may take longer than 3 -6 months after stroke to adapt to the situation of living with the consequences of the stroke. One year after stroke, many persons with even mild stroke still struggle to cope with these consequences, often hidden dysfunctions (30,31). There was a difference in SF-36 MCS and MH at 3-month follow between the IG and CG vs baseline measurements. The magnitude of these changes was 9 points for SF-36 MCS and 16 points for MH. The clinical signifi cance of SF-36 is set at 5 points or more (8), indicating in the present study that a clinically detectable and meaningful change has occurred.
The reason(s) for the difference between the IG and CG in SF-36 MCS and MH at 3 months post-intervention are unclear. Possible factors may include:
(i) Different content in the intervention program in the IG and the CG. The group discussions in the IG and the CG were equal in numbers and length but they had a completely different focus. The discussions in the IG concentrated solely on fall risk and security aspects, while the discussions in the CG contained group discussion on hidden dysfunctions after stroke as their complete intervention. The vast majority of the subjects in the CG reported in a questionnaire that they perceived the group discussions as interesting and meaningful with good and open-minded atmosphere within the group. (ii) Disappointment and frustration in the IG that the intervention had ended and thereby that the subjects had lost their intensive connections to the other participants in the study group as well as to attention loss of the personnel involved in the study. This is supported by the overwhelmingly positive feedback given by the majority of the subjects in the IG to the intervention staff in the study at the end of the 5-week intervention program. Also, many of these subjects expressed a desire to participate repeatedly in high-intensive exercise program rounds.
The main focus of the presently evaluated intervention study was not to investigate the impact on HRQoL and depression but rather to explore the effects on various physical outcomes. The evaluation of HRQoL and depressive symptoms was done in order to see if an " ordinary " exercise program did have an effect on these kinds of outcomes. The study was not designed for effect on these outcomes; therefore these results are in fact confi rmatory of this. If effect on these variables, HRQoL and depressive symptoms, are sought after, the study needs to be designed for that as well. It could, however, be considered as a strength in this study, that the different discussion topics in the CG seem to be of importance for HRQoL. For future interventions programs, this knowledge is valuable and needs to be added to the discussion sessions in order to have a possible effect on these outcomes.
It is known that the different invisible symptoms such as different cognitive problems are very common, but not often primary will be taken into account during the rehabilitation period. In persons 55 years and younger, cognitive problems are frequently perceived (32). This may be important to absorb and in the design of future programs add these topics to the theory sessions to demonstrate effects on HRQoL.
Regarding depressive symptoms, there was no signifi cant difference over time in the whole study group or between the IG and the CG. Five out of 34 subjects (three in the IG and two in the CG) had a previous history of depression whereas 27% in the IG and 32% in the CG were on antidepressant therapy at baseline measurements. This shows that the proportion of subjects in the present study with treated depression is in the range of previously published data on post-stroke depression (12). However, the average value of GDS-15 was 3.0 in the whole study group at baseline measurements (2.5 in the IG and 3.4 in the CG), indicating that depressive symptoms were not common. This suggests that many of the subjects in our study had a positive effect of their anti-depressive medication. Since the intervention study could have been considered demanding to the post-stroke individuals, both in time as well as physical and mental effort, it is likely that many of these individuals with more pronounced depressive symptoms declined the offer to participate in the study.
The SF-36 has been tested for validity and reliability in populations similar to the population in this study (8,21,23,24). However, there is a debate on some of its subscales, indicating imprecision and confounding when making summed scores (33). The results from the subscales of importance in this study nevertheless support the generation of summed scores from PF, RP, BP, VT, RE and MH (2). The Stroke Impact Scale (SIS) is another possible instrument that may be suitable for a post-stroke sample (34). The SIS instrument is disease-specifi c and results in values on physical symptoms, function, activity, ability/trouble in daily life, the impact of the disease, satisfaction and patient satisfaction (34). Both SIS and SF-36 are generic instruments, but in contrast to the SIS, the SF-36 is not limited to be diseasespecifi c. Generic instrument in general, evaluates the patient ' s general health and QoL. An advantage of using SF-36 is that it could be used to compare the results with other studies of other diseases using the SF-36. In our study population, the mean age was 79 years and the vast majority of the subjects had a high co-morbidity with illnesses, such as cardiovascular and other diseases. It might therefore be diffi cult to ensure the disease-specifi c origin of the symptoms. Therefore, a generic instrument rather than a diseasespecific instrument has been considered as best relevant to use in this study.
As has been discussed in the accompanying report, this sample was considered representative for individuals with post-stroke symptoms of fall risk and the length of the intervention program is realistic for implementation in clinic (17). Fatigue as a common complication after stroke was also given as the main reason not to participate by those who declined to participate in this study. Of those eligible for inclusion in the study but were not included in the study, some lived too far away from the rehabilitation facilities. A few others did not have the time to spend in the intervention program, while some others died during the time between inpatient care and study start. With these facts, it might be that our included study group is a bit more active than the average post-stroke individuals with risk of falls. However, many participants in the study group expressed a pronounced fatigue.
In the present study, the mean result of the SF-36 subscale VT (which is defi ned as either feeling tired all the time or having lots of energy all of the time ( 8)) indicates that this proportion was higher than in the norm population. VT has been strongly associated with global ratings of health satisfaction and QoL (8).
There are several limitations in this study (17). The size of the study population is small. The power calculation was based on the BBS and estimated minimum size of 34 subjects. To reach a high precision in SF-36, there is a call for more than 200 subjects (8). This suggests that there is a possibility for a Type-2 statistical error in other subscales of SF-36 in the present study, i.e. there may be a difference that is non-detectable because of our small sample size.
We consider it a strength to try to evaluate HRQoL and depression, and not only functional and activity measures, even though the intervention study is focused on these outcomes. It is established that both physical and psychosocial well-being is greatly affected in stroke survivors (1). Many rehabilitation programs are focused on the physical outcome of the program. There is a need for evaluating rehabilitation programs in terms of HRQoL and the presence of depressive symptoms. It could also be concluded, to achieve an effect on depressive symptoms, that the program should have included exercise that focused on higher level of aerobic training (33,36).
Antidepressant drugs may be useful in treating depression after stroke, but can also cause side-effects, especially such as seizures, falls and delirium (37). Adding these side-effects to an already decreased functional ability may be problematic. Therefore, an alternative method as physical activity may be a useful way to ameliorate depressive symptoms. Ideally, in order to evaluate the effect of physical fi tness training on depression after stroke, the participating stroke individuals should be medication-free regarding anti-depressants. When designing the structured intervention program, the focus of the study was on functional rehabilitation, with an attention time of 30 h during the 5-week intervention program for the IG and 5 h for the CG and actually not to enhance the HRQoL. The psychosocial part was incorporated in the 1 h/week of education received by the individuals in both groups.
Our 5-week high-intensive exercise program was related to deterioration in the IG in SF-36 MCS and the MH subscale at 3 months post-intervention compared with baseline values while the CG improved at this time. For the SF-36 PCS, there was an improvement in the whole study group at 3 and 6 months postintervention compared with baseline values without any signifi cant difference between the IG and the CG. The presence of depressive symptoms was unchanged throughout the follow-up period for both groups. Based on these data, it is concluded that a modifi ed version of a high-intensive exercise program to be tested in the future should not entirely focus on falls and safety aspects, but should also include themes on hidden dysfunctions after stroke in order to have a favorable impact on HRQoL.
This study was supported by grants from V å rdalinstitutet, the
The study is registered at www.clinicaltrials.gov (
The authors report no confl icts of interest. The authors alone are responsible for the content and writing of the paper.
Women have long strived to possess long, thick, and dark eyelashes. Prominent eyes and eyelashes are often considered a sign of beauty and can be associated with increased levels of attractiveness, confidence, and well-being. Numerous options may improve the appearance of eyelashes. Mascara aims to temporarily darken, lengthen, and thicken eyelashes using a combination of waxes, pigments, and resins. Artificial eyelashes can be adhered either to the dermal margin or to individual eyelashes. Individuals may even use eyelash transplantations to improve the appearance of their eyelashes. The unique properties of eyelashes (e.g., relatively long telogen and short anagen phases compared with scalp hairs, slow rate of growth, and a lack of influence by androgens) may allow for specific aesthetic interventions to improve the appearance of natural eyelashes. Some over-the-counter (OTC) products may contain prostaglandin analogs that can affect eyelash growth, but neither the safety nor efficacy of these OTC cosmetics has been fully studied. Originally indicated for the reduction of intraocular pressure, the synthetic prostaglandin analog bimatoprost was recently approved for the treatment of hypotrichosis of the eyelashes. In a double-blinded, randomized, vehicle-controlled trial, bimatoprost safely and effectively grew natural eyelashes, making them longer, thicker, and darker. Bimatoprost was generally safe and well tolerated and appears to provide an additional option for individuals looking to improve the appearance of their eyelashes.
Since ancient times, physical beauty has been considered an advantageous and sought-after trait [1]. Although the definition of beauty varies over time and from culture to culture, the face and the eyes in particular are recognized as important contributors to physical beauty [1,2]. Beautiful eyes are associated with social advantages [1]. Eyelashes that are long and thick are considered a sign of beauty in many cultures and often have a positive psychological effect on women [3][4][5][6]. To enhance the overall prominence of their eyelashes, women have employed a number of techniques, some dating back millennia [7].
Presently, many women rely on makeup or over-thecounter (OTC) cosmetics to enhance the appearance of their eyelashes. As of 2005, the global cosmetics and toiletries market totaled $155 billion and color cosmetics (e.g., eye makeup, facial makeup, lip products, and nail products) comprised 15% of that market or approximately $24 billion (Allergan, Inc., Irvine, CA, unpublished data). Eye makeup, which includes eye shadow, eye liner, and mascara, represents the fastest-growing category of color cosmetics in the U.S., with a growth rate of more than 6% in 2005. In the U.S. alone, the annual mascara market is estimated at $1.1 billion (Allergan, Inc., unpublished data). Mascara can provide women with increased eyelash volume, longer, darker lashes, and improved curl [7; Allergan, Inc., unpublished data].
In addition to makeup, the effects of which are temporary and subject to smudging, women have several longerlasting and sometimes permanent options for improving the appearance of their eyelashes, including eyelash extensions and eyelash transplants [8,9]. Recently, the U.S. Food and Drug Administration (FDA) approved the use of bimatoprost ophthalmic solution 0.03% (Latisse Ò , Allergan Inc., Irvine, CA), a synthetic prostaglandin analog, to increase the length, thickness, and darkness of eyelashes in people with hypotrichosis of the eyelashes (i.e., inadequate or not enough eyelashes) (Latisse package insert, Allergan, Inc., 2008).
Beyond their aesthetic and social functions, eyelashes serve a protective function by defending the eye against debris and triggering the blink reflex [10][11][12]. People generally have between 100 and 150 lashes emanating from each of their upper eyelids [11,12]. Lower eyelashes are half as numerous as upper eyelashes [11]. Upper eyelashes are arranged in two to three rows. Similar to scalp hair, eyelashes are considered terminal hairs, and as such are coarser, longer, and more pigmented than other hair types (i.e., vellus and intermediate) [13]. Eyelashes are wider than scalp hairs and, unlike other hair types, do not typically lose pigmentation and become gray with age. Eyelashes are distinct from all other hairs on the body in that they lack an accompanying arrector pili muscle, and, unlike many other hairs, are not influenced by androgens [13,14].
As is true for all hair follicles on the body, all eyelash follicles are present at birth and their numbers do not increase during life [13,15]. The hair follicles of many mammals exhibit synchronous hair cycles, but in humans the hair cycle is asynchronous such that some hair follicles are growing while others are dormant [13]. The hair cycle for all hair types is divided into the phases of anagen, catagen, and telogen, but the average length of the cycle and the individual phases varies by body location.
Though variable, the normal eyelash cycle is estimated to last from 5 to 11 months (Fig. 1a) (Latisse package insert, Allergan, Inc., 2008, and unpublished data) [13,16,17]. The growth phase of eyelash follicles, anagen, lasts approximately 1-2 months. During anagen, in addition to growth, melanogenesis and the subsequent transfer of pigment to the hair shaft also occurs [13]. The duration of anagen crucially impacts hair length [10]. Anagen is a period of rapid cell proliferation and differentiation [18]. Following anagen, eyelash follicles enter catagen, a transition phase, which lasts approximately 15 days and is the time during which epithelial elements of the follicle undergo apoptosis or programmed cell death [13]. The longest phase of the normal eyelash cycle, telogen or the resting phase, lasts approximately 4-9 months [13,16,17]. Throughout telogen, no significant cell differentiation, proliferation, or apoptosis occurs [18]. Expulsion of the previous hair (i.e., exogen) takes place during the transition between telogen and anagen [10,14].
In contrast to eyelashes, scalp follicles have a much longer cycle, lasting several years [14,15]. Anagen alone can last up to 6-7 years for scalp hairs [13][14][15]. Relative differences in the lengths of the hair cycle phases of eyelashes and scalp hair result in approximately 50% of upper eyelash follicles being in telogen at any given time compared with only 5% to 15% of scalp follicles [13][14][15][16]. Furthermore, eyelashes are typically slow-growing hairs, growing at a rate of approximately 0.15 mm/day [17] compared with 0.3-0.4 mm/day for scalp hair [16][17][18]. The unique properties of the eyelash cycle differentiate eyelashes from other body hairs and may cause drugs that affect hair growth in one location to enhance eyelash prominence.
Artificial eyelashes are synthetic or donated human eyelashes that give the visual appearance of longer eyelashes. Strips of artificial lashes can easily be adhered to the upper eyelid [8]. As intimated by the nickname ''falsies,'' such lashes may not result in a natural appearance. Alternatively, single eyelash extensions can be attached to individual eyelashes [8,19]. The procedure requires an hour and a half to complete and can cost as much as $500 [19]. Artificial eyelashes are typically held in place by methacrylate-based adhesives that can result in allergic contact dermatitis [8]. Similarly, the solvent used to remove artificial eyelashes can also cause an allergic reaction in some individuals. Depending on the type of artificial lash and the procedure used to adhere them, artificial eyelashes can remain in place from several days to several weeks [8,19]. Lash Extender (myofibril-styled eyelash extensions applied by the patient) is another option.
Successful eyelash transplantations have been reported in the medical literature, although their use in the absence of congenital eyelash defects or a need for reconstruction is controversial [9]. Eyelash transplantations transfer hair follicles from the scalp onto the margins of the eyelid. Because these transplanted lashes retain the qualities of scalp hair, patients must regularly trim and curl the implanted lashes [9,20]. Complications of eyelash transplantation include pain, bleeding, thick scar formation, numbness, eyelid ptosis, and blindness [9].
Bimatoprost ophthalmic solution 0.03% is the only FDAapproved product to safely and effectively enhance the growth of a patient's own eyelashes. Bimatoprost is a synthetic prostamide, or prostaglandin ethanolamide, analog approved in 2001 for the reduction of elevated intraocular pressure in patients with open-angle glaucoma or ocular hypertension (Lumigan package insert, Allergan, Inc., 2006). In one clinical trial, when administered as an eye drop for the treatment of glaucoma, eyelash growth was noted in 42.6% of patients treated with bimatoprost once daily for a year [21]. Although such growth was recorded as an adverse event, the potential aesthetic benefits of eyelash growth were recognized and led to the development and testing of bimatoprost as a product designed to increase eyelash prominence. When prescribed to enhance eyelash prominence, bimatoprost is accompanied by sterile, single-use-per-eye applicators and should be applied once daily to the skin of the upper eyelid margin at the base of the eyelashes (Latisse package insert, Allergan, Inc., 2008).
The safety and efficacy of once-daily bimatoprost 0.03% solution in increasing overall eyelash prominence following dermal administration to the upper-eyelid margins was evaluated in a recent multicenter, double-blinded, randomized, vehicle-controlled, parallel study of 278 adult patients [22]. Starting at 8 weeks of treatment, bimatoprost was associated with significantly greater increases in overall eyelash prominence than vehicle as measured by at least a one-grade increase on the four-grade Global Eyelash Assessment (GEA) scale. These changes were sustained throughout the remainder of the treatment period (16 weeks) and, at week 16, 78.1% of subjects treated with bimatoprost exhibited at least a one-grade increase from baseline in GEA score compared with 18.4% of patients treated with vehicle (P \ 0.0001) (Fig. 2). Improvements in GEA scores continued to favor bimatoprost 4 weeks after discontinuation of treatment (i.e., the post-treatment visit at week 20). In the same study, efficacy was assessed by digital image analysis of superior-view eyelash photographs taken with standardized equipment and after uniform subject preparation at all visits. As assessed by such analysis, at week 16 bimatoprost treatment was associated with a mean 25% increase in eyelash length vs. 2% for vehicle, a mean 106% increase in eyelash thickness vs. 12% for vehicle, and a mean 18% increase in eyelash darkening vs. 3% for vehicle (Fig. 3). Subjects treated with bimatoprost for eyelash growth also reported feeling significantly more satisfied with their eyelashes, more confident in their looks, more attractive, more professional, and more satisfied with their daily routine than those subjects receiving vehicle as assessed by patient-reported outcome questionnaires.
Fig. 2 Sample images of patient eyelashes before and after 16 weeks of once-daily bimatoprost treatment. The patient entered the study with a baseline overall eyelash prominence assessed as moderate (grade 2) on the Global Eyelash Assessment (GEA) scale. After 16 weeks of double-blind treatment with bimatoprost ophthalmic solution 0.03%, the patient's eyelashes were markedly prominent (i.e., GEA score of 3). In determining GEA scores, raters utilized a photonumeric guide and evaluated overall eyelash prominence, including length, fullness, and color of both upper eyelashes, with length considered the most important feature Although the mechanism by which bimatoprost enhances eyelash growth has not been fully elucidated, it is believed to increase the percentage of lash follicles in and the duration of anagen (Fig. 1b) (Latisse package insert, Allergan, Inc., 2008). Bimatoprost also appears capable of stimulating melanogenesis [23], which likely explains the changes in the pigmentation of eyelashes observed with use. The duration of the effect of bimatoprost on eyelashes is not fully known and has not been evaluated beyond 4 weeks post-treatment. Furthermore, the ability of bimatoprost to affect the growth of eyelashes in patients with systemic diseases or eyelash loss has not been evaluated.
The safety of bimatoprost, applied either dermally or as an eye drop, has been demonstrated in numerous trials [21,24]. When applied as an eye drop (i.e., for the treatment of ocular hypertension), bimatoprost was generally well tolerated with long-term use for up to 4 years [24]. Other than eyelash growth, the most common adverse events reported when bimatoprost is instilled into the eye include conjunctival hyperemia, eye pruritus, eye dryness, a burning sensation in the eye, eyelid pigmentation, foreign body sensation, eye pain, and visual disturbance [21, 24; Lumigan package insert, Allergan, Inc., 2006]. Although rare, bimatoprost when applied as an eye drop has also been associated with increases in iris pigmentation, which are generally thought to be permanent [25]; Lumigan package insert, Allergan, Inc., 2006]. While iris pigmentation was not seen in the controlled clinical studies with dermal application for eyelash growth, there have been a small number of post-marketing reports of iris pigmentation following use of bimatoprost for eyelash growth (data on file, Allergan Medical Affairs, June 2010). Post-marketing reports are often incomplete and difficult to evaluate on an individual basis. However, to date there have been no identified trends of risk factors reported in these small numbers of cases. Patients should be advised of the small potential for increased brown iris pigmentation which is likely to be permanent. Infection following application is also a possibility and should be limited by following the package insert instructions for application with daily single-dose applicators for each eye, which are provided with each package.
Although clinical experience regarding the safety of bimatoprost 0.03% applied dermally is limited, available data suggest that the safety profile may be even more favorable than that described above for intraocular administration. In the pivotal trial, the most common adverse events associated with bimatoprost use were eye pruritus, conjunctival hyperemia, skin hyperpigmentation, ocular irritation, dry eye symptoms, and erythema of the eyelid, all occurring in fewer than 4% of patients (Latisse package insert, Allergan, Inc., 2008). Conjunctival hyperemia was the only specific event reported at a significantly higher rate by subjects receiving bimatoprost (3.6%) than those receiving vehicle (0%) (P = 0.03) [22]. For comparison, in a 3-month trial of once-daily bimatoprost treatment for glaucoma or ocular hypertension (i.e., the drug was instilled as an eye drop), the incidence of treatment-related conjunctival hyperemia and eye pruritus was approximately 46 and 9%, respectively [26]. Furthermore, dermal application of bimatoprost was not associated with iridal pigmentation or clinically meaningful changes in intraocular pressure [22; Latisse package insert, Allergan, Inc., 2008]. Differences in safety profiles of bimatoprost by route of administration may be partially explained in that a single application of bimatoprost 0.03% to the upper-eyelid margin using the provided applicator delivers approximately 5% of the dose delivered for the treatment of glaucoma (Allergan, Inc., unpublished data). Moreover, when applied to the upper-eyelid margin, subsequent absorption of the active drug into ocular tissues is expected to be minimal due to the barrier function of the skin and the small surface area to which the dose is applied.
The ability of prostaglandin analogs to increase eyelash growth does not appear limited to bimatoprost. Published reports describe latanoprost and travoprost, both prostaglandin analogs used to treat ocular hypertension, as being associated with eyelash changes, including increases in length and darkness [3,5,13,[27][28][29]. Studies evaluating the clinical efficacy of bimatoprost versus latanoprost for the treatment of ocular hypertension suggest that the side effect of increased eyelash growth is more common with bimatoprost [28]. However, these results are not necessarily comparable to those when these drugs are applied to the upper-eyelid margin, which have not been evaluated in any head-to-head studies. In addition, studies in which Fig. 3 Changes in eyelash length, thickness, and darkness associated with bimatoprost in a double-blind, vehicle-controlled trial [22]. Eyelash qualities were assessed by digital image analysis based on superior-view digital eyelash photographs taken with standardized equipment and subject preparation. Last observation carried forward was performed on weeks 1-16. * P \ 0.0001 vs. placebo based on the Wilcoxon rank-sum test increased eyelash growth is recorded as an adverse event typically fail to quantify the changes using objective measures.
While other prostaglandin analogs appear capable of influencing eyelash growth, these products are not FDAapproved for the treatment of hypotrichosis. Their efficacy when applied to the upper-eyelid margins has not been fully studied in clinical trials. Moreover, the safety of these drugs when applied dermally has not been fully evaluated.
The Federal Food, Drug, and Cosmetic Act defines cosmetics as ''articles intended to be rubbed, poured, sprinkled, or sprayed on, introduced into, or otherwise applied to the human body or any part thereof for cleansing, beautifying, promoting attractiveness, or altering the appearance'' [30]. The FDA requires cosmetics to have an ingredient declaration but cosmetic products and ingredients are not subject to FDA premarket approval [31]. A product can be considered both a drug and a cosmetic by the FDA if it has multiple uses [32].
Historically, mascara has been used to lengthen, thicken, and darken eyelashes via the use of waxes, pigments, resins, as well as talc or fibers. In the case of liquid mascara, these ingredients adhere to lashes after the product dries, resulting in a temporary enhancement in the appearance of a person's eyelashes [7,8]. A number of nonmascara OTC cosmetic products are advertised to increase the length, fullness, and/or darkness of eyelashes. These products contain various ingredients such as ones described as ''proprietary peptides'' [33], natural extracts [34,35], vitamins [33], and prostaglandin analogs. Some of the products that contain possible prostaglandin analogs include MD Lash Factor [36], Revitalash [37], LiLash [38], neuLash [39], Lashes to Die For [40], and RapidLash [41,42].
As cosmetics, the efficacy of the above OTC products has not been critically evaluated and their safety has not been fully studied. Most are applied to the skin of the upper-eyelid margin and the mechanisms by which they may affect eyelash growth are largely unknown and unproven. These products also vary in the quality and comprehensiveness of patient/consumer education regarding proper use of the cosmetic.
In addition to the possibility of ingredient-specific safety concerns, mascara and other cosmetic eyelash enhancers carry a risk of bacterial and fungal contamination and subsequent infection in addition to a risk of conjunctival pigmentation [7]. Unlike prescription products, cosmetics do not benefit from a comprehensive pharmacovigilance program to detect any safety concerns that become evident with increased clinical use.
It is clear that eyelashes play an important role in determining beauty, and as such, prominent eyelashes are a highly sought-after attribute. Numerous options may improve the appearance of eyelashes, including mascara, artificial eyelashes, eyelash transplantation, and prostaglandin analogs. Prostaglandin analogs appear capable of influencing eyelash growth and are contained in several OTC cosmetics. Eyelash growth has also been associated with several prescription prostaglandin analogs used to treat glaucoma and ocular hypertension. The FDA has recently approved bimatoprost as the first and only option for the treatment of hypotrichosis of eyelashes [12]. In a vehicle-controlled trial, bimatoprost resulted in increased prominence, length, thickness, and darkness of patients' natural eyelashes as assessed by clinician ratings and digital image analysis [27]. As an ocular antihypertensive agent, bimatoprost has a proven history of safety, and when applied dermally it was safe and well tolerated with a low incidence of adverse events.
The author is a consultant and investigator for Allergan, Inc.
A life-long follow-up of physiological and behavioural functions was initiated in 38-month-old mouse lemurs (Microcebus murinus) to test whether caloric restriction (CR) or a potential mimetic compound, resveratrol (RSV), can delay the ageing process and the onset of age-related diseases. Based on their potential survival of 12 years, mouse lemurs were assigned to three different groups: a control (CTL) group fed ad libitum, a CR group fed 70% of the CTL caloric intake and a RSV group (200 mg/kg.day -1 ) fed ad libitum. Since this prosimian primate exhibits a marked annual rhythm in body mass gain during winter, animals were tested throughout the year to assess body composition, daily energy expenditure (DEE), resting metabolic rate (RMR), physical activity and hormonal levels. After 1 year, all mouse lemurs seemed in good health. CR animals showed a significantly decreased body mass compared with the other groups during long day period only. CR or RSV treatments did not affect body composition. CR induced a decrease in DEE without changes in RMR, whereas RSV induced a concomitant increase in DEE and RMR without any obvious modification of locomotor activity in both groups. Hormonal levels remained similar in each group. In summary, after 1 year of treatment CR and RSV induced differential metabolic responses but animals successfully acclimated to their imposed diets. The RESTRIKAL study can now be safely undertaken on a long-term basis to determine whether ageassociated alterations in mouse lemurs are delayed with
The main causes of mortality in western countries are chronic age-associated diseases such as cardiovascular diseases, cancers, type 2 diabetes, osteoporosis and kidney diseases. These diseases have a genetic back-ground but environmental factors such as nutrition play a key role (Mokdad et al. 2004). Since the pioneering study of McCay et al. (1935), we know that caloric restriction (CR) without malnutrition can delay the onset of such age-associated pathologies, delay ageing and increase lifespan in rodents. Observational data suggest that the relationship between longevity, energy balance and caloric intake could exist in humans but further investigations are needed.
Several hypotheses have been proposed for the biological mechanism underlying the life-extending and anti-ageing actions of CR, such as slowing growth, reducing body fat, altering apoptosis, decreasing body temperature and increasing physical activity, attenuating oxidative damage, reducing glycemia and insulinemia and modifying immune function (Campbell and Richardson 1988;Gredilla et al. 2001;Gardner 2005). It is likely that all these and others actions might indeed play a role and a unifying hypothesis was needed. The concept of hormesis might well provide such a hypothesis. This refers to phenomena in which the response of an organism to a chemical or physical agent is qualitatively different when the agent is of high intensity than when it is of low intensity (Rattan 2004). A 30% CR that delays the ageing processes is a low intensity stressor, which enhances the ability of rats, mice and monkeys of all ages to cope with intense stressors (Masoro 2006). The pathway by which CR enhances protective and repair processes has not yet been well elucidated. Recently, a mechanism was proposed for CR-induced effects in yeast (Saccharomyces cerevisiae). Lin et al. (2000) demonstrated that the functional Sir2 gene is required for CR to increase replicative longevity in that species. More recently, an increasing amount of data has suggested that silent mating-type information regulation 2 homologous 1 (SIRT1), one of the seven mammalian orthologues of the yeast Sir2, regulates cell survival to a number of stressors and might be central to the effects of CR (Cohen et al. 2004;Picard et al. 2004;Rodgers et al. 2005;Bordone and Guarente 2005;Baur 2010).
Until the late 1980s, CR had not been tested in any animal model living longer than 4 years. Since then, studies in non-human primates, including rhesus and squirrel monkeys, have been in progress at the National Institute of Ageing (NIA) and the University of Wisconsin (UW) (Ingram et al. 1990;Lane et al. 1992;Kemnitz et al. 1993). The proportions of both rhesus and squirrel monkeys already dead in 2002 in the CR cohorts were about half of the controls (Lane et al. 2002). The most recent result from the UW primate study has indicated that adult-onset mild CR delays the onset of age-associated pathologies and promotes survival in rhesus monkeys (Colman et al. 2009) but this study has not yet shown any effect of CR on the maximum life expectancy of primates. Recently, a similar study was initiated on human subjects. Results at 6 months of CR are consistent to what has been observed in monkeys and rodents in terms of lower metabolic rate, body temperature and insulin and Further studies are needed to test the validity of the CR paradigm in non-human primates to delineate the still debated mechanisms by which CR extends lifespan and ultimately test the molecules or interventions that will mimic the effects of CR and be relevant to humans. The present study "RESTRIKAL" aims to investigate the long-term effects of CR or supplementation with a mimetic compound (resveratrol; RSV) on the ageing process and lifespan of a non-human primate and highlight similar action pathways in both treatments. In that respect, the main originality of this study lies in the animal model, the grey mouse lemur (Microcebus murinus) whose size (60-110 g) and life expectancy (8-10 years) will allow obtaining longevity data in 5 years only. M. murinus is a nocturnal prosimian primate originating from Madagascar. The grey mouse lemur presents special characteristics that make it a unique model to investigate energy regulation and ageing processes. Firstly, it displays a strong seasonal rhythm linked to photoperiod that is associated with natural and massive winter obesity i.e. an important increase in body mass due to fat storage in only a few weeks (Génin and Perret 2000). Secondly, grey mouse lemurs exhibit very marked daily rhythms with periods of heterothermia during their diurnal resting phase (Perret and Aujard 2001). These phases of daily heterothermia, very rare in primates, are an efficient mechanism of energy saving in small mammals (Terrien et al. 2009a) and thereby a good response index to energy stresses. It is also a unique model to assess the importance of biological rhythms on the processes of ageing and longevity. For example, a weakened and fragmented locomotor activity rhythm during normal ageing has been demonstrated in this primate (Cayetanot et al. 2005). The longevity of this primate is adequate for performing longitudinal studies, 16 AGE (2011) 33:15-31 reduced DNA damage (Heilbronn et al. 2006).
and several biomarkers of ageing have already been validated in the mouse lemur colony, in particular decrease in sexual and aggressive behaviours (Aujard and Perret 1998;Nemoz-Bertholet et al. 2004), decrease in sexual hormones (Aujard and Perret 1998), melatonin (Aujard et al. 2001), DHEA-S (Perret and Aujard 2005) and insulin-like growth factor-1 (IGF-1) (Aujard et al. 2010), cognitive impairments (Picq 2007) and MRI-evaluated cerebral atrophy (Dhenain et al. 2003;Kraska et al. 2009). Moreover, it has recently been demonstrated that the energy balance of aged mouse lemurs is impaired in response to cold exposure (Terrien et al. 2009b). The first aim of this study is to determine whether CR can modify the physiological processes in the grey mouse lemur that lead to a delay in age-related diseases and an increase in lifespan.
The second aim of this study is to determine whether RSV, a natural polyphenolic compound that activates proteins implicated in energy metabolism homeostasis, could be used as a mimetic compound of CR. Indeed, CR will be very difficult to implement in humans because of social and practical constraints. In the past 5 years, numerous studies have focused on the development of "CR mimetic" compounds that would minimise many age-related diseases in humans without a reduction in caloric intake (Weindruch et al. 2001;Chen and Guarente 2007;Ingram et al. 2007;Wakeling et al. 2009). Among them, RSV seems to be a promising molecule (Howitz et al. 2003;Borra et al. 2005;Lagouge et al. 2006;Allard et al. 2009;Anderson and Prolla 2009). Trans-RSV, the most active form of RSV, has been the subject of many in vitro, ex vivo and in vivo studies to test its various and promising antiinflammatory (Donnelly et al. 2004), anti-oxidant (Iannelli et al. 2007), anti-cancer (Baur et al. 2006), metabolic and cardiovascular (Shanmuganayagam et al. 2007) properties.
Although RSV effects share many metabolic similarities with CR (Barger et al. 2008) and although it seems to be a promising molecule to delay the incidence of age-associated chronic diseases (Athar et al. 2007), its metabolic effects in primates are still unknown. Indeed, relatively low concentrations can extend yeast, worm and Drosophila lifespan in a Sir2-dependent manner by mimicking CR (Wood et al. 2004). More interestingly, RSV can activate SIRT1, the most studied mammalian orthologue of Sir2, in humans (Borra et al. 2005). Most recently, several studies have also demonstrated RSV can stimulate AMPK activity in mice (Dasgupta and Milbrandt 2007;Canto et al. 2009). These results highlight that RSV could play an important part in energy regulation processes. RSV treatment also reduces the signs of ageing in mice but does not increase the longevity of ad libitum-fed animals when started at midlife (Pearson et al. 2008). However, non-human primates and humans present metabolic responses to a long-term CR and mimetic compound that might differ slightly from what was repeatedly observed in rodents and other lower organisms (Heilbronn et al. 2006;Colman et al. 2009;Witte et al. 2009). For example, CR in mice down-regulates genes involved in oxidative stress and reduces oxidative damage, lipid peroxidation and protein carbonyls (Sohal et al. 1994;Dubey et al. 1996;Lee et al. 1999;de Oliveira et al. 2003). In non-human primates, genes involved in protection against oxidative stress are not altered by CR, although protein carbonylation is reduced (Zainal et al. 2000). More data on primates are needed. Thus, the third goal of the study is to define whether CR and RSV act through similar pathways by stimulating sirtuins in the mouse lemur. This point will be developed in future years of the study.
Since the beginning of the RESTRIKAL study, a battery of tests has been performed on adult grey mouse lemurs at regular intervals until their natural death to assess the impact of both treatments on energy metabolism modification and the ageing process. These tests were selected to assess parameters changing with chronological ageing. They are minimally invasive and appropriate for a longitudinal study. Several indices of energy balance were measured such as body mass gain, food intake, body composition, resting metabolic rate (RMR) and physical activity. Concerning endocrine systems, IGF-1 levels were measured because this hormone is a modulator of energy balance and an ageing marker that declines with time (Aujard et al. 2010). Indeed, genetic alterations in the human IGF-1 receptor that result in an altered IGF signalling pathway seem to be linked to human longevity, suggesting a role of this pathway in the modulation of human lifespan (Suh et al. 2008). Testosterone levels are also measured to highlight possible effects on sexual behaviour and, notably, test if the balance between reproduction cost and survival of grey mouse lemurs is modified. The first-year outcome of this longitudinal follow-up demonstrates the validity of CR and RSV interventions for assessing whether they can delay ageing in a non-human primate. We expect that the present study will provide novel insights into the biology of ageing by testing the effect of CR or RSV treatment on a non-human primate model. We expect that CR will increase lifespan in the grey mouse lemurs and decrease markers of several chronic diseases. Lastly, we expect that RSV stimulation of the sirtuin pathway will mimic the effect of CR. In the long run, such a study might allow us to develop nutritional strategies to delay the effect of ageing.
The 42 male grey mouse lemurs (M. murinus, Cheirogaleidae, primates) used in this study were born in the laboratory breeding colony at Brunoy, France (agreement A91-114-1) from stock originally caught more than 40 years ago on the southwest coast of Madagascar. All animals were included in a single cohort at the age of 38±1 months, which is considered an adult age in this species. The project began at the onset of the winter-like season in the lab. The general conditions of captivity were maintained with respect to ambient temperature (25°C) and relative humidity (55%). The grey mouse lemur is a nocturnal primate that shows high levels of locomotor activity and normothermic body temperature during the night. Just before the onset of the light phase, it enters a torpid state during which it decreases its body temperature and metabolic rate until 4 to 6 h after the beginning of the day when its body temperature returns to normothermic values. The mouse lemur exhibits photoperiod-dependent seasonal variations in most of its physiological functions. Another interesting particularity of this species is its body mass gain during winter. Indeed, at the beginning of winter the body mass of the grey mouse lemur increases by approximately 50% to 70% and only returns to basal levels in the summer. In the breeding colony, animals are exposed to an artificial photoperiodic regimen consisting of six months of summer-like long day length (14:10 h light-darkness, LD) and 6 months of winter-like short day length (10:14 h light-darkness, SD). These photoperiodic regimens are sufficiently discriminating to induce radically different physiologi-cal and behavioural responses, as observed in nature in Madagascar. The change of photoperiod occurred abruptly without any significant disturbance for the animals. To minimise social influences during the different experiments, animals were housed individually in 1 m 3 cages, provided with a nest and supports and separated from each other by metallic partitions. Animals were weighed once a week to monitor body mass variations. All experiments were performed in accordance with the Principles of Laboratory Animal Care (National Institutes of Health publication 86-23, revised 1985) and French national laws.
Animals were fed with the same standard diet used in the laboratory breeding colony and in recent publications using the same animal model (Génin and Perret 2003;Giroud et al. 2008b;Giroud et al. 2010). They were fed with fresh fruit (banana and apple) and a daily mixture made up of cereals, milk and egg. This diet is composed of 61% carbohydrates, 23% proteins and 16% lipids. Water was always given ad libitum.
After a 2-week habituation phase, animals were randomised to the following groups. An ad libitum control (CTL) group of 14 animals was fed with the standard diet. The daily amount of food given to the animals (15 g of mixture and 6 g of fresh fruit per day, equivalent to 105 kJ/day on average) was estimated from preliminary internal studies of the Brunoy laboratory over a year of measuring spontaneous food intake in isolated control adult animals (unpublished data). During the first 2 weeks of the SD period, animals remained undisturbed and more food (169 kJ/ day) was given to allow annual fattening. Then, the amount was progressively decreased (128 kJ/day in the third week, 116 kJ/day in the fourth week) until the fifth week when the amount of food was fixed for the rest of the year (102 kJ/day). A CR group of 14 animals was fed the same diet but received 30% less than the CTL group, which corresponded to an average of 71 kJ per day (i.e. 10 g of mixture and 4 g of fresh fruit per day). This restriction was based on several studies on different species (Lane et al. 2000;Blanc et al. 2003). It was applied throughout the year, except during the critical period of fattening. During the first 3 weeks of the SD period, CR animals received the same amount of food as the CTL group to allow them to fatten normally and, after these first 3 weeks, the 30% caloric restriction applied was the same during SD and LD periods. Finally, a third group (RSV group) of 14 animals was fed with the same quantity of food as CTL but supplemented with 200 mg of RSV per kilogram bodyweight per day (Sequoia Research Products, UK). This dosage was selected from the literature from studies in rodents, and was intermediate between the 40 mg/kg.d -1 of Baur et al. (2006) and the 400 mg/kg.d -1 of Lagouge et al. (2006). To know exactly the quantity of food really ingested by the animals, daily leftovers were measured and corrected for water evaporation.
Daily energy expenditure, body composition and water turnover Measurements were performed twice a year on 12 of the 14 animals in each group in the first weeks following each dietary shift (SD1, LD1) (Fig. 1). Daily energy expenditure (DEE) was measured over a 3-day period using the doubly labelled water (DLW) method (Blanc et al. 2000). Animals were weighed before the experiment to determine the dose of DLW to be injected. After urine collection for the determination of basal enrichments, animals were injected in the intraperitoneal cavity with 2.3 g/kg of a pre-mixed solution composed of 0.55 g/kg H 2 18 O (Rotem Industries, Israel) and 0.15 g/kg 2 H 2 O (Cambridge Isotope Laboratories, Andover, MA, USA) diluted in 10 g of NaCl 9‰ to maintain osmolarity. These quantities of isotopes ensured in vivo enrichments of deuterium in the order of 300 ppm and 2,400 ppm for oxygen-18. Isotopic equilibration in total body water (TBW) was determined from a blood sample collected using a glass capillary at 1 h post-dose from the saphenous vein. Immediately after sampling, the capillary tubes were flame-sealed to prevent isotopic exchanges. The mouse lemur was then released into its cage and urine samples were collected in cryogenically stable tubes 24, 48 and 72 h after blood sampling. Blood and urine samples were respectively stored at 5°C and -20°C until analyses by isotope ratio mass spectrometry.
Water from serum and urine samples were extracted by cryo-distillation, as previously described (Gilbert et al. 2007). A 0.1 µL of water was reduced to hydrogen and carbon monoxide by reduction on a glassy carbon reactor held at 1,400°C in an elemental analyser (Flash HT, Thermofisher, Germany). Hydrogen and carbon monoxide gases were separated by a gas-liquid chromatography column held at 104°C coupled to a continuous flow Delta-V isotope ratio mass spectrometer. Isotopic abundances of deuterium and 18-oxygen in hydrogen and carbon monoxide gases were measured in quintuplicate and repeated if SD exceeded 2‰ and 0.5‰, respectively. All enrichments were expressed as International Atomic Energy Agency standards. CO 2 production was calculated according to the single pool equation of Speakman (Blanc et al. 2000): rCO 2 = (N/2.078) × (ko-kd) -0.0062 × kd × N,
2 5 w 6 2 w 0 w w5 w10 w12 w15 w18 w20 w26 w31 w32 w34 w41 w43 w44 w46 w49 w52 DEE WTO FM FFM IGF-1 RMR Testosterone Locomotor activity Testosterone Testosterone IGF-1 RMR IGF-1 RMR DEE WTO FM FFM IGF-1 RMR w39 w13 D L D S SD1 SD2 LD1 LD2 2 5 w 6 2 w 0 w w5 w10 w12 w15 w18 w20 w26 w31 w32 w34 w41 w43 w44 w46 w49 w52 DEE WTO FM FFM IGF-1 RMR Testosterone Locomotor activity Testosterone Testosterone IGF-1 RMR IGF-1 RMR DEE WTO FM FFM IGF-1 RMR w39 w13 D L D S SD1 SD2 LD1 LD2 Fig. 1 Experimental schedule during the first year of the RESTRIKAL study. Daily energy expenditure (DEE), water turnover (WTO), fat mass (FM) and fat-free mass (FFM) were assessed twice a year. Measurement of insulin-like growth factor type 1 (IGF-1) level and resting metabolic rate (RMR) were performed four times a year. Testosterone level analysis was performed three times a year, once in SD and twice in long days (LD) period. Locomotor activity was tested once a year at the end of the LD period during which the animals remain relatively active (w week)
where N represents the average isotope dilution space of oxygen-18 calculated from Coward (1990) by the plateau method using the 1 h post-dose sample. ko and kd represent the isotope constant elimination rates calculated by linear regression of the natural logarithm of isotope enrichment as a function of elapsed time from day 1 samples. DEE was calculated by Weir's equation (Weir 1949) using a food quotient of 0.86 estimated from the animal's diet. TBW was measured from the dilution space of 18-oxygen after correction for exchange by the factor 1.007 (Racette et al. 1994).
Fat-free mass (FFM) was calculated from TBW by assuming a hydration coefficient of 73.2%, which was shown to be unchanged by chronic CR (Blanc et al. 2005). Fat mass (FM) was calculated as the difference of FFM from body mass. Water turnover was assessed by the multiplication of the average isotope dilution space of oxygen-18 (N) with the deuterium constant elimination rate (kd) and corrected for isotope fractionation (Blanc et al. 2000). FFM and FM were expressed in g, TBW was expressed in % and DEE was expressed in kJ/day.
Oxygen consumption was measured with a closed circuit respirometer. This protocol was applied four times per year, in the first and last third of each photoperiod (SD1, SD2, LD1 and LD2) (Fig. 1). All the animals were experimented with this procedure (13 CTL, 14 CR and 14 RSV). For this nocturnal species, RMR measurements on post-absorptive animals were performed during their daily resting period, 4-6 h after the beginning of the light period to avoid torpor metabolism. Animals were trained to nest in the respiratory chamber that consisted of an opaque chamber of 2.5 l with a woven floor to absorb any urine. During the experiment, the respiratory chamber was placed in a cabinet at a controlled ambient temperature of 25.0±0.5°C, a value within the thermoneutral zone defined for the mouse lemur (Aujard and Perret 1998). After a 20 min habituation phase under constant air-flow ventilation (2 l.min -1 ) drawn through the respirometry chamber from bottom to top; the chamber was closed for a 40 min period. VO 2 consumed by the animal was calculated from initial and final concentrations of O 2 in the chamber that were measured on dried gas using a Servomex 570 A paramagnetic gas analyser (accuracy 0.01% O 2 ). The analyser was routinely calibrated with N 2 and atmospheric air. O 2 consumption was expressed as ml O 2 h -1 . O 2 consumption was adjusted for the body weight of the animal (Blanc et al. 2003). Body temperatures were not measured before the experiment to avoid disturbance of the sleeping animal but males never entered deep torpor at an ambient temperature of 25°C.
Spontaneous locomotor activity was estimated once a year at the end of the LD period, during their breeding season, when the animals are more active (Aujard et al. 2007) (Fig. 1). For technical reasons, only five CTL, nine CR and seven RSV followed this protocol. Spontaneous locomotor activity was estimated using a device with presence and motion sensors adapted to the mouse lemur (homemade system developed in the laboratory). The apparatus was placed in an ambient temperature controlled room (25°C). Animals were housed individually in a cage with a capacity of 1 m 3 each provided with nest and supports. Each animal was placed in a totally closed wooden nest, except for both sides of the nest on which were placed eight presence sensors (Honeywell, transmitter: SEP8705003, receiver: SDP8405014) to know when the animal was in the nest. These presence sensors were continuously recording. Moreover, two motion sensors (Gardtec, Gardscan 'M' series infrared detectors) were placed in the corners of the cage to detect the spontaneous movements of the animal during its activity period. If the animal was in movement, the motion sensors recorded data every two seconds. Thus, data were expressed in arbitrary units (a.u.). Data were stored in a computerised system (developed in the laboratory). They were then computed to represent the time course of these movement patterns using a software filtering "Actocebe 3.0" developed in language G from National Instruments. Total movements were averaged at 5-min intervals for further analysis.
Hormonal assays Blood collections were taken via the saphenous vein of the animals, without anaesthesia, at the end of their resting phase before the food allotment became available. Blood was collected in capillary tubes containing EDTA and immediately centrifuged (7,000 rpm at 4°C for 30 min) after collection. One hundred microlitres was collected per sampling and represented less than 2% of the blood volume of the animal. Plasma was stored at -80°C. Blood sampling for IGF-1 level analysis was performed four times a year, in the first and last third of each photoperiod (SD1, SD2, LD1 and LD2) (Fig. 1). IGF-1 level was measured for 11 animals by group, using an IGF-1 IRMA kit (Immunotech SA, Marseille Cedex, France). A dilution step was performed before the assay to dissociate IGF-1 from its binding proteins. The antibodies used in this immunoassay are highly specific for IGF-1 (extremely low cross-reactivity against insulin, pro-insulin, IGF-2 and growth hormone). The intra-assay coefficient of variation was 6.3% and inter-assay coefficient of variation was 6.8%. Assay sensitivity was 2 ng/ml. Values were expressed in ng/ml.
Blood sampling for testosterone level analysis was performed three times a year, in the middle of the SD period and twice during the LD period, which is the activity period of grey mouse lemur (SD, LD1 and LD2) (Fig. 1). Testosterone was analysed for 12 CTL, 11 CR and 11 RSV. Testosterone was assayed using the testosterone ELISA kit DE1559 (Demeditec, Kiel, Germany), which measures the total testosterone in plasma (expressed in ng/ml). After two distinct periods of incubation of 60 and 15 min with the different reagents, the optical density was read with a spectrophotometer at 450 nm (Bio-Tek Instruments Inc., Winooski, USA). Assay sensitivity was 0.083 ng/ml. The intra-assay coefficient of variation was 6.7% and inter-assay coefficient of variation was 9.7%. Values were expressed in ng/ml.
Testis size variations of each animal, expressed as a scale in a.u., was assessed once a week by the same person to complete the testosterone level analysis by estimating visually their sexual status (if there was no testis, a score of zero was given, if the testis began to appear, a score of one and when the testis were well developed, a score of two).
Animals are followed until their spontaneous death. All animals that will die during the study undergo a complete autopsy by a veterinarian. Based on specific criteria (rapid body mass loss, anaemia, difficulty breathing), dying animals will be deeply anaesthe-tised. After autopsy will be completed, all organs will be harvested and kept for future analysis. Special attention will be give to the mouse lemur's brain to check it for Alzheimer-like pathology.
For calculation, we chose the increase in the mean lifespan as the primary outcome. In the breeding colony of the Brunoy laboratory, analysis of survival from 254 male mouse lemurs allowed us to determine the mean lifespan (mean ± SEM: 6.0±0.2 years), the mean lifespan of the 10% longest living animals (10.0 ± 0.2 years) and the observed maximal survival duration (12.0 years). Based on literature (Lin et al. 2000;Tissenbaum andGuarente 2001, Anderson et al. 2003;Howitz et al. 2003;Ingram et al. 2006;Valenzano et al. 2006), we expected that the mean lifespan of the restricted and the RSV-supplemented groups would increase by a minimum of 30%. With this parameter, assuming a common standard deviation of 5% and a power analysis of 80%, a sample size of 12 individuals per group was required. We chose to increase the groups size (n=14) to compensate the effect of early hazardous deaths in reducing statistical powers of metabolic and behavioural variables measured throughout the time course of the study during the project. The actual total power
Because of the large amount of data gathered during this longitudinal study, an Access database was created to comprehensively analyse all the results. For technical and setup reasons, the number of animals used for each experiment varied, especially for the locomotor activity monitoring, which was planned after the beginning of the project. Moreover, one animal of the CTL group suffering from a urinary tract infection died during the first year of this study and was removed from the analysis. All values are expressed as mean ± SEM. After checking for the normality of the distribution, ANOVA or repeated ANOVA for related samples was used to assert significant variations in all studied parameters. Data from average seasonal body mass plateaus and IGF-1 levels were log-transformed to obtain data with a normalised distribution. Adjustments of FFM, RMR will be recalculated a posteriori. and TEE for body size were performed by analysis of covariance with a fluctuating slope model. Comparisons were considered to differ significantly when p<0.05. All statistics were performed by SYSTAT for Windows (V9, SPSS Inc., USA).
Seasonal variations in the weekly average percentage of spontaneous food ingested related to the amount of food given and seasonal variations in body mass are represented in Fig. 2. The percentage of calories ingested (CI) significantly varied throughout the year (Fig. 2a). In CTL animals, a spontaneous and progressive decrease in CI was observed during the SD period to reach a minimum of 75% (76 kJ/day) in the last third of the SD period. CI was higher in LD with an average of 92% (90 kJ/day) in the second part of the LD period. CI always remained below 100% (102 kJ/day), confirming that CTL animals were always fed ad libitum. RSV animals exhibited similar seasonal variations in CI compared with CTL, except at the beginning of LD. Indeed, CI increased rapidly in the RSV group after the shift to LD, whereas the increase was more progressive in the CTL group. CR lemurs demonstrated higher percentages of CI throughout the year compared with CTL and RSV lemurs. Seasonal variations were less marked, but a spontaneous decrease in CI during SD was still
40 60 80 100 120 140 160 1 3 5 7 9 11 13 15 17 19 21 23 25 27 29 31 33 35 37 39 41 43 45 47 49 51 Time (week) Body mass (g) CTL CR RSV SD LD 60 65 70 75 80 85 90 95 100 105 1 3 5 7 9 11 13 15 17 19 21 23 25 27 29 31 33 35 37 39 41 43 45 47 49 51 Time (week) CI (%) control CR RSV SD LD 60 70 80 90 100 110 120 130 140 150 SD LD Time (photoperiod) body mass plateau (g) CTL CR RSV Tr: F=2.833, df 2/35, p=0.072 P: F=75.797, df 51/1785, p<0.001 C: F=0.782, df 102/1785, p=0.945 Tr: F=3.927, df 2/24, p<0.001 P: F=14.554, df 51/1224, p<0.001 C: F=0.858, df 102/1224, p=0.838 Tr: F=4.298, df 2/36, p=0.021 P: F=220.918, df 1/36, p<0.001 C: F=2.388, df 2/36, p=0.106 Data of the comparison of mean body mass gain during the short days (SD) and long days (LD) plateaus were logtransformed. Plateaus were calculated from weeks 13 to 19 for the SD period and weeks 38 to 44 for the LD period. ANOVA results are reported on the right side of the graphs, for treatment (Tr), photoperiod (P) or crossed effects (C) on the percentage of CI and body mass changes during the first year of CR or RSV supplementation compared to control feeding (CTL) in SD and LD animals (n=13 for CTL group, n=14 for CR group, n=14 for RSV group). p in bold means value is significant and p in italic means value is not significant but shows a trend. Values are expressed as mean ± SEM observed (p=0.046). CR animals ingested an average 95% (68 kJ/day) of their food portion during SD and 98% (70 kJ/day) during LD. The percentage of restriction really applied throughout the year was less than 30% (19±1% in SD and 24±1% in LD).
As shown in Fig. 2b, the animals presented an initial average body mass of 90±5 g. All animals fattened during the first 12 weeks of the SD period to reach a plateau of 131±8 g about 8-9 weeks before decreasing their body mass until another plateau of 83±9 g in the middle of the LD period (Fig. 2b, p<0.001). Across the first year of the study, body mass did not significantly differ between the three groups even though a trend effect of treatment was observed (p=0.072). However, when the data were split by seasons, CR animals presented a lower body mass compared with CTL and RSV groups during the second half of the LD period (74±2 vs. 87±4 and 90± 3 g, respectively, p=0.021, Fig. 2c).
Analyses of body composition, DEE and water turnover following each new photoperiod are represented in Fig. 3. These parameters measured the short-term effect of treatment (within the second month following the first shift to SD) and the effect of 6 months of treatment (within the first month following the next shift to LD). Body mass at the time of these measures varied significantly according to season, with higher values in SD than in LD, but did not vary significantly according to treatment (Fig. 3a). There was no significant effect of time or treatment on FFM values (75±2 g for CTL, 72±3 g for CR and 73±2 g for RSV in SD and 71±1 g for CTL, 68±2 g for CR and 76±2 g for RSV in LD, Fig. 3b). However, a crossed effect close to significance was observed, with a decrease in FFM in LD compared with SD for CTL and CR animals, whereas inversely RSV-treated lemurs increased their FFM in the LD season. By contrast, a clear effect of season was observed for FM values with higher levels of FM in SD than in LD, without any significant effect of treatment (Fig. 3c). TBW did not differ between the three groups whatever the photoperiod, but it varied similarly between SD and LD for each group (46±1% in SD and 57±1% in LD) (Fig. 3d). Water turnover did not differ according to photoperiod or treatment but the cross effect was almost significant (p=0.055). CR and RSV groups presented lower values in SD compared with the CTL group and maintained similar levels of water turnover in LD compared with SD, whereas CTL animals showed decreased water turnover rate in LD compared with SD (Fig. 3e). The evaluation of DEE revealed a significant effect of treatment with no global effect of season. CR animals exhibited lower values of DEE than CTL and RSV animals in both photoperiods, and RSV animals exhibited higher DEE values than the two other groups in LD only (Fig. 3f). These effects were maintained when DEE was adjusted to body mass (Fig. 3g).
RMR varied according to photoperiod and treatment (crossed effect: F=2.450, df 6/114, p=0.029) (Fig. 4). In CTL animals, RMR was higher in early SD and late LD. CR animals presented similar RMR values compared with the CTL group throughout the year (p=1.000). By contrast, the RMR values of the RSV group significantly increased throughout the year (p= 0.011). Significant differences appeared between the groups. In LD1, RMR was higher for the RSV group compared with CTL (p=0.044). In LD2, this difference between RSV and CTL was even more pronounced (p=0.003). During the late LD season, spontaneous daily locomotor activity was no different between the three groups of animals (1,987±630 a.u. for CTL, 2,644±457 a.u. for CR and 1,379±338 a.u. for RSV, p=0.157) (Fig. 5).
IGF-1 levels significantly varied according to photoperiod (p<0.001), with lower values in early SD and late LD in all groups (Fig. 6). Despite a trend for a global treatment effect, IGF levels were not significantly different between the three groups whatever the photoperiod (p=0.078).
Highly significant differences in testis size (Fig. 7a) and testosterone levels (Fig. 7b) occurred between each photoperiod (p<0.001) but treatment did not affect these parameters (p=0.726 and p=0.203, respectively).
The RESTRIKAL project aimed to compare for the first time the effects of a long-term CR or a mimetic compound supplementation on multiple ageing processes in a heterothermic primate. After 1 year, all mouse lemurs appeared in excellent health, suggesting no detrimental effect of both diets. Expected variations because of photoperiodic entrainment were present in all groups. Neither CR nor RSV supplementation affected photoperiodic variations in FFM, FM, TBW, water turnover, IGF-1 levels or testosterone levels. However, the evolution of FFM and water turnover in RSV-supplemented animals presented several differences compared with controls. Indeed, FFM and water turnover values in the RSV group varied in an opposite manner to the CTL group. Moreover, the RSV group exhibited higher values of DEE at the beginning of the LD period and a strong increase in RMR without any increase in spontaneous daily locomotor activity at the end of the LD period. Lastly, although body mass values did not differ between groups during the SD photoperiod, CR animals presented with a significant decrease of their body masses at the end of the LD period. Fig. 3 Effects of photoperiod and treatment on body mass at the time of the experiment (a), fat-free mass (b), fat mass (c), total body water (d), water turnover (e), daily energy expenditure (f) and daily energy expenditure adjusted to body mass (g). ANOVA results are reported on the right side of the graphs, for treatment (Tr), photoperiod (P) or crossed effects (C) on body mass, fat-free mass, fat mass, total body water, water turnover and daily energy expenditure adjusted or not to body mass during the first year of caloric restriction (CR) or resveratrol (RSV) supplementation compared to control feeding (CTL) in short days (SD) and long days (LD) animals (n=12 for each group). p in bold means value is significant and p in italic means value is not significant but shows a slight trend. Values of body mass, fat-free mass and fat mass are expressed in g; values of total body water are expressed in %; values of water turnover are expressed in g.day -1 .g -1 of animal; values of daily energy expenditure are expressed in kJ.day -1 . All parameters are expressed as mean ± SEM Exposed to a SD photoperiod, mouse lemurs fattened quickly and, after 2-3 months, exhibited a spontaneous reduction of their caloric intake (Génin and Perret 2000). Consequently, CR was proportionally less severe in SD than LD (21% vs. 30%, respectively). Therefore, the body mass of CR animals did not differ from those of controls during the first 6 months of this study i.e. the SD period. Accordingly, IGF-1 levels were not modified. However, at the end of the LD photoperiod, restricted animals exhibited a lower body mass than controls and, although not significantly different, their IGF-1 levels were reduced. A similar decrease in IGF-1 levels has been demonstrated in mice exposed to moderate CR after 24 weeks of The neuroendocrine system involving GH and IGF-1 mediates some of the metabolic consequences of caloric excess or restriction (Smith 1996). Obesity is associated with increases in levels of free IGF-1 (Nam et al. 1997) and free fatty acids, both of which are known to decrease GH secretion through a negative feedback mechanism (Lee et al. 1995). Endocrine changes under CR have been extensively studied in rodents (Shimokawa and Higami 2001). Weight loss, as observed with a restricted diet, is accompanied by a decrease in both free fatty acids and free IGF-1 levels, leading to a potential increase in GH secretion (Smith 1996). However, IGF-1 and GH levels were unaffected in humans after a 6 monthperiod of CR (Redman and Ravussin 2009).
Human ageing is marked by a reduction in both GH and IGF-1 concentrations in healthy adults (Veldhuis et al. 2005). Some of the anti-ageing actions of CR involve the modification of several 3.0 3.5 4.0 4.5 5.0 5.5 6.0 6.5 7.0 SD1 SD2 LD1 LD2
Time (photoperiod) Resting metabolic rate (ml d'O2.h -1 .g -0.67 ) CTL CR RSV Tr: F=5.054, df 2/38, p=0.011 P: F=7.013, df 3/114, p<0.001 C: F=2.450, df 6/114, p=0.029 Fig. 4 Effects of photoperiod and treatment on resting metabolic rate. ANOVA results are reported on the graph, for treatment (Tr), photoperiod (P) or crossed effects (C) on resting metabolic rate changes during the first year of caloric restriction (CR) or resveratrol (RSV) supplementation compared to control feeding (CTL) in short days (SD) and long days (LD) animals (n=13 for CTL group, n=14 for CR group, n=14 for RSV group). p in bold means value is significant. Values are expressed in ml O 2 .h -1 .g -0.67 as mean ± SEM Fig. 6 Effects of photoperiod and treatment on plasma insulinlike growth factor type 1 (IGF-1) level. Data were logtransformed for statistical analysis. ANOVA results are reported on the right side of the graph, for treatment (Tr), photoperiod (P) or crossed effects (C) on IGF-1 level changes during the first year of caloric restriction (CR) or resveratrol (RSV) supplementation compared to control feeding (CTL) in short days (SD) and long days (LD) animals (n=11 for each group). p in bold means value is significant and p in italic means value is not significant but shows a slight trend. Values are expressed in ng/ml as mean ± SEM treatment (Huffman et al. 2008).
neuroendocrine pathways but they remain to be studied (Berner and Stern 2004;Meites 1989). In some mammalian species, CR can reduce the agerelated IGF-1 decline (Berryman et al. 2008). However, the observed decrease in IGF-1 levels after 1 year of CR in mouse lemurs cannot be attributed to ageing alone. Indeed, in this primate age-related changes in IGF-1 levels are characterised by a progressive decrease during the SD photoperiod, whereas levels remain high during the LD photoperiod even in old age (Terrien et al. 2009a;Aujard et al. 2010). It is most likely that changes in metabolism account for the decrease in IGF-1 in mouse lemurs at Animals under moderate CR adapted quickly to the protocol. Indeed, body composition as reflected by FM, FFM and TBW was not modified by CR. However, during the SD photoperiod DEE was lower in restricted animals compared with control ones. Similar results were found in rhesus monkeys subjected to a 30% reduction in caloric intake after several years of treatment (Raman et al. 2007). The reduction in DEE in the grey mouse lemurs has to be related to its ability to enter daily torpor to conserve energy in the face of environmental constraints (Terrien et al. 2009a). Mouse lemurs exposed to moderate food shortage (Giroud et al. 2008a) exhibited an increasing frequency of daily torpor. Therefore, animals likely minimise energy expenditure and compensate for the effects of CR on body composition by entering deep torpor. Further studies should investigate the length and duration of torpor bouts in CR animals.
RMR, recorded after the daily torpor bout of each animal, was not different between control and CR animals. By contrast, the RMR of human subjects placed under a 25% CR decreases after three months compared with controls (Martin et al. 2007). In our experiments, the percentage of CR applied in the mouse lemur might not be sufficient to induce RMR variations or the duration of treatment (1 year) might not be long enough to induce significant modifications. At any rate, this lack of change in RMR also suggests that the animals do not save energy during
Tr: F=0.324, df 2/35, p=0.726 P: F=90.497, df 51/1785, p<0.001 C: F=0.937, df 102/1785, p=0.658 0 10 20 30 40 50 60 70 80 90 SD LD1 LD2 Time (photoperiod) Testosterone level (ng/ml) CTL CR RSV 0 0.5 1 1.5 2 2.5 3 1 4 7 10 13 16 19 22 25 28 31 34 37 40 43 46 49 52 Time (weeks) Testis size (a.u.) CTL CR RSV SD LD Tr: F=1.678, df 2/31, p=0.203 P: F=56.371, df 2/62, p<0.001 C: F=1.211, df 4/62, p=0.315 a b Fig. 7 Effects of photoperiod and treatment on testis size (a) and testosterone level (b). ANOVA results are reported on the right side of the graphs, for treatment (Tr), photoperiod (P) or crossed effects (C) on testis size and plasma testosterone level during the first year of caloric restriction (CR) or resveratrol (RSV) supplementation compared to control feeding (CTL) in short days (SD) and long days (LD) animals (n=12 for CTL group, n=11 for CR group, n=11 for RSV group). p in bold means value is significant. Values of testis size are expressed in a.u. and values of testosterone level are expressed in ng/ml. Both parameters are expressed as mean ± SEM
the end of the LD period.
their resting periods but probably during deeper torpor. Moreover, the CR group presented with lower values of water turnover compared with control animals, indicating greater water retention. This is probably because of the increased frequency of torpor bouts, an important mechanism in this species during SD photoperiods.
From the beginning of the LD photoperiod, all animals lost body mass. Body composition and related body mass did not differ between groups. The LD photoperiod corresponded to the breeding season with a general activation of physiological and behavioural parameters (Perret 1985). In particular, an increase of reproductive functions was expected. As reflected in this study, higher values of testosterone levels were observed for the CTL and CR groups in the LD period but we noticed no difference between the two groups. Interestingly, the expected increase in LD period-related water turnover was not observed for any group. Moreover, DEE did not differ between control and restricted animals. One possible explanation might be the period when these parameters were measured. Indeed, when the experimental design of this study was established, it was anticipated that animals would exhibit their most significant metabolic changes just after the photoperiodic shifts. However, changes in body mass and DEE as well as IGF-1 levels were more affected by dietary restriction in the second part of the LD period. Therefore, for the remaining years of the study, measurements will be made in the second part of the LD period.
During exposure to the LD photoperiod, CR impacted heavily on body mass variations of grey mouse lemurs (Génin and Perret 2003;Séguy and Perret 2005;Giroud et al. 2008a, b). In the LD photoperiod, mouse lemurs did not use daily torpor to avoid body mass loss even under a 40% CR (Giroud et al. 2008a). Moreover, this decrease of body mass could not be explained by changes in RMR or locomotor activity. Contrary to expectations, despite having less energy reserves, no increase in exploratory behaviour, as reflected by spontaneous daily locomotor activity, was observed in CR animals. Similar results were observed after several weeks of 40% CR in mouse lemurs in the LD photoperiod (Giroud et al. 2008a). However, when CR was higher, for example 80% (Génin and Perret 2003), a strong increase in locomotor activity was recorded, suggesting the active search of required food. In the rhesus monkey (Macaca mulatta), an increase in activity levels has been observed for animals under a 30% CR (Weed et al. 1997). However, these results were only valid for the older male group followed in the NIA study. In monkeys, the literature in this domain was well developed, but the results obtained in the different studies did not allow to conclude about caloric restriction effects on locomotor activity (Kemnitz et al. 1993;Ramsey et al. 1996;Delany et al. 1999;Moscrip et al. 2000). Presently, how CR affects locomotor activity is unclear. As stressed by Ingram et al. (2001), the large variability of the results could be attributed to various factors such as the age and sex of the animals and different CR intensities and durations. Further studies will be necessary to highlight the reality of CR effects on the locomotor activity of primates.
Animals supplemented with RSV ate approximately the same amount of food as controls during both photoperiods and exhibited similar photoperiodic variations in their body mass and FM. In the SD photoperiod, their body composition, energy expenditure, RMR and hormonal level were similar to those recorded in the control group. Interestingly, the RSV group did not present any difference in body mass compared with the control group, suggesting that RSV does not mimic CR for this parameter, at least after a year of treatment.
However, the FFM of the RSV group was higher in LD period compared with the control group, with a high inter-individual variability. Animals might have developed a larger muscle mass as demonstrated in RSV-supplemented mice that displayed increased muscle strength in locomotor tests after several weeks of treatment (Baur et al. 2006;Lagouge et al. 2006). Moreover, RSV can stimulate in vivo muscle cell glucose uptake through sirtuins and AMPK (Breen et al. 2008), an effect possibly involved in the muscle mass increase of RSV-supplemented mouse lemurs. DEE was also modified by RSV treatment in mouse lemurs with higher values compared with controls at the beginning of the LD period. A similar trend was observed for water turnover. Lastly, contrary to the CR group, the RMR of RSV-supplemented animals was noticeably increased in the early and late stages of the LD period. This RSV effect on the RMR was already seen in a previous study where a 4-week supplementation with RSV was sufficient to induce an increase of the grey mouse lemurs' RMR (Dal-Pan et al. 2010). Moreover, as for high-fat mice after several weeks of RSV supplementation (Lagouge et al. 2006), RSV seemed to activate significant energy metabolism of grey mouse lemurs, in particular their RMR.
In summary, a moderate CR or a mimetic compound supplementation was safely initiated in an adult heterothermic primate. First-year results of the study suggest a good adaptation of the animals to their diet. Preliminary results after 1 year of treatment suggest that RSV produces an activation of energy metabolism without body mass loss in contrast to CR, which produces a decrease of energy expenditure. Taken together, the effects of CR or RSV supplementation in grey mouse lemurs seemed more important during LD than during SD periods. Physiological and behavioural parameters are highly photoperiod-dependent in this primate, and it is likely that during LD photoperiods energy constraints are more important than during SD photoperiods, thereby leading to a more pronounced effect of the diet treatment. However, changes observed at the end of the LD period might also reflect the chronic effects of treatment after 1 year. The respective roles of this photoperiodic dependency and duration of treatment should be resolved in forthcoming years along with the evolution of validated biomarkers of ageing.
In conclusion, after only 1 year of treatment, it cannot be proposed that the hormesis theory is the most appropriate hypothesis to explain the effects observed in a non-human primate exposed to CR or RSV supplementation. CR and RSV might be two different stressors taking a different path to produce the same effects. Indeed, several differences were observed between CR and RSV supplementation, questioning the similarity of the pathways involved. If similar pathways between CR and RSV supplementation exist, as for example through sirtuins (especially SIRT1), it will be necessary to investigate the evolution of the levels of such molecules in future years of the RESTRIKAL project.
lesterolemic rabbits. Atherosclerosis 190(1):135-142
The authors acknowledge the continuing assistance provided by
J. Terrien : F. Pifferi : R. Botalla : I. Hardy : J. Marchal : I. Chery : M. Perret : F. Aujard (*) Mécanismes Adaptatifs et Evolution,
A randomized controlled efficacy trial targeting older adults with hypertension (age 60 and over) provided an e-health, tailored intervention with the "next generation" of the Personal Education Program (PEP-NG). Eleven primary care practices with advanced practice registered nurse (APRN) providers participated. Participants (N=160) were randomly assigned by the PEP-NG (accessed via a wireless touchscreen tablet computer) to either control (entailing data collection and four routine APRN visits) or tailored intervention (involving PEP-NG intervention and four focused APRN visits) group. Compared to patients in the control group, patients receiving the PEP-NG e-health intervention achieved significant increases in both selfmedication knowledge and self-efficacy measures, with large effect sizes. Among patients not at BP targets upon entry to the study, therapy intensification in controls
According to the World Health Organization, hypertension represents the greatest risk factor for premature death worldwide (Mathers et al. 2009). More than one-third of adults aged 60 and over in the United States has hypertension-a condition resulting in more health care visits than any other chronic condition (Schroeder et al. 2004). Nationwide, it is estimated that only 35% of older adults with hypertension maintain target blood pressure (BP) readings (<140/90; <130/80 for those with diabetes or chronic kidney disease) (Schroeder et al. 2004;Wong et al. 2007). Inadequately controlled hypertension, owing to poor patient adherence to antihypertensive regimens and adverse-self medication behaviors, contribute to annual estimated health care costs of $100 billion (Institute of Medicine 2006;NHLBI 2007).
Older adults with hypertension were found to have low self-efficacy in their ability to avoid serious health consequences, owing to their large knowledge deficits regarding interactions between prescription and OTC agents (Neafsey and Shellman 2002a). Rather than identify and remediate low self-efficacy in patients, which results in poor adherence and adverse self-medication behaviors, uncontrolled BP is often treated with intensified antihypertensive therapy. This therapy is regularly administered with increased doses, additional agents, and/or drug changes, thus further heightening the risk of adverse drug effects (ADEs) and patient care costs (Ho et al. 2008;Peterson 2008). Moreover, clinical trials have yet to demonstrate any long-term improvement in patient adherence to antihypertensive therapy (Haynes et al. 2008). By contrast, intensive, monthly counseling by nurses or pharmacists has been shown to improve antihypertensive adherence in older adults, but BP control typically declined when the intervention ceased (Bosworth et al. 2005;Lee et al. 2006;Roumie et al. 2006).
There are a number of other causes and/or effects related to unsuccessful control of patient blood pressure level. For instance, patient reticence about reporting symptoms from medication side effects during provider visits is significantly correlated to ameliorable and preventable adverse drug events (ADEs) (Weingart et al. 2005). The number of symptoms reported during the past month is associated with the number of self-reported ADEs (Oladimeji et al. 2008). Over-the-counter (OTC) medications, supplements, and alcohol all interact with antihypertensives and contribute to poor BP control (Gurwitz et al. 2003;Institute of Medicine 2006;Wallsten et al. 1995). For example, patients with hypertension may choose a nonsteroidal anti-inflammatory drug (NSAID) to self-medicate pain (Neafsey and Shellman 2001;Neafsey et al. 2007)-without knowing that NSAIDS (e.g. ibuprofen) can increase blood pressure and antagonize the anti-platelet effects of low-dose aspirin and the effects of anti-hypertensive agents when taken concurrently (Aw et al. 2005;MacDonald and Wei 2003;Polonia 1997). Hence, it seems logical to find that educating patients about safe medication use can help reduce the risk of potential adverse drug interactions (PADI) (Gurwitz et al. 2003).
With advances in e-health, a tailored system could provide a cost-effective patient education tool to help increase patient knowledge and self-efficacy, and thus safe self-medication practice. The Personal Education Program (PEP) is one such network-based e-health intervention system that has demonstrated its effectiveness in improving safe self-medication knowledge, efficacy, and behaviors among older adults with hypertension (Neafsey et al. 2002(Neafsey et al. , 2001;;Strickler and Neafsey 2002). Its successor, the "next generation" PEP (PEP-NG), contains significant software and educational content enhancements. For instance, the once paper-and-pencil measurement instruments (assessing such outcome measures as medication use, medication knowledge, self efficacy, and user satisfaction) became part of the software system to enable dynamic real-time assessment of patient interface outcomes. A rules engine was added to the software system to assess self-reported self-medication behaviors and deliver tailored education to the patient.
This manuscript presents the results of an efficacy trial of the PEP-NG conducted in 11 primary care settings. The tailored, touch-screen tablet-based educational intervention is one of the first designed to reduce adverse self-medication behaviors in older patients with hypertension. In conceptualizing the study, which aimed at stimulating cognitive learning and enhancing self efficacy in patients to motivate them to adopt safe selfmedication practices and modify adverse self-medication behaviors, Bandura's social cognitive theory was utilized as the general theoretical framework to guide the study design (Bandura 1997(Bandura , 2001)). The constructs of Bandura's Social Cognitive Theory (following, in quotes, Bandura 2001), as applied to the PEP-NG content design, encompass the following conceptual aspects: 1) "symbolizing capability" (animations in the PEP-NG form mental pictures) that give "meaning, form, contiguity" to patients' selfmedication experiences to "guide future behaviors;" 2) "vicarious capability" (animations and related multiple choice questions) that enables "observational learning" so that patients can envision patterns of behavior quickly, avoiding mistakes; 3) "forethought capability" (interactive questions in the tailored-education segments) that allow patients to consider "predictive function and expectations of behavioral outcomes;" 4) "selfregulatory capability" (self-efficacy instrument and feedback from interactive questions) that motivates patients to acknowledge their confidence in performing future tasks related to self-medication; and 5) "reciprocal determinism" that enables the "bi-directional interaction" with the APRN during the focused visit following PEP-NG use.
The goal of the clinical efficacy trial (conducted in primary care settings) was to reduce adverse outcomes associated with unsafe self-medication practices in older adults with hypertensionthrough improved patient-provider communication-with the aid of the PEP-NG system. Trial objectives for older-adult patients were to show that users of the PEP-NG would: 1) increase knowledge concerning potential drug interactions stemming from unsafe self-medication practices; 2) enhance their self-efficacy in learning and adopting safe self-medication practices; 3) reduce selfreported unsafe self-medication behaviors associated with potential adverse drug interactions; 4) improve their prescription medication adherence; 5) achieve and maintain target blood pressure readings; 6) express their satisfaction with the PEP-NG; and 7) enhance the patient-APRN provider relationship.
A full description of the PEP-NG instruments, interface, prior usability, pilot testing with older adults, and clinical-trial methodology can be found elsewhere (Lin et al. 2009(Lin et al. , 2010;;Neafsey et al. 2009Neafsey et al. , 2008;;Strickler et al. 2008). A brief explanation of how the patient and provider utilize the PEP-NG software to achieve the study objectives is provided below, followed by a description of the research procedures adopted for the current project.
The PEP-NG was accessed via a wireless tablet computer*; patients used a stylus to answer questions on the touchscreen interface. In terms of on-screen display, visual objects were large (3 cm high) and text size was in a 20-point size Arial Black font to facilitate ease of reading for older adults. Ergonomically adaptive, wide-scroll bars and dropdown-menus displayed in blocks of eight lines eased selection of agents for those with impaired hand mobility and/or fine tremor. The time of medication and dosage was reported with the use of an easy-to-use animated clock. Patients were asked what they took for treating common ailments or conditions (e. g. blood pressure, blood thinning, pain, cold or sinus, allergies, sleep, stomach problems such as indigestion or gas, and low thyroid). Patients were also asked, "Did you take ____ in the last month?" with respect to calcium pills, vitamins, minerals, herbs or supplements, and alcohol, wine, or liquor.
A rules engine analyzed patient-inputted information and immediately delivered individually tailored educational content on the tablet screen. Summaries of a patient's self-reported symptoms, medication use (including frequency/time), adverse self-medication behaviors (along with a thumbnail screen shot from a related animation), and corrective strategies were automatically printed for review by the APRN provider prior to the primary care visit. The APRN reinforced the corrective education information (that appeared on the printout) with the patient as part of their primary care BP visit. At the conclusion of the APRN visit, the patient took a copy of the same printout home for self-study. A Virtual-Private-Network (VPN) transferred all PEP-NG interface data to a Microsoft Access database. The VPN met HIPAA requirements (Federal Register 2000) and the European Union Directive 95/46/EC (de Meyer et al. 1998). The interface was developed in accordance with ISO 9100 international standards (ISO 2004; Kelly 2000).
The study was approved by the University Institutional Review Board (IRB) and met all HIPAA regulations prior to enrolling any provider or patient participants. All methods adopted were performed in accordance with the 1964 Declaration of Helsinki and all study participants gave consent prior to participation in the study (World Medical Association 1964).
Two practice-based research networks (PBRN) in New England cooperated in study site recruitment. APRNet is a PBRN of APRNs, funded by the Agency for Healthcare Research and Quality (AHRQ) and administered by the Yale School of Nursing. The Connecticut Center for Primary Care (CCPC) PBRN is an independent, non-profit corporation established (under CT law) by ProHealth Physicians, Inc.
Primary-care practice owners and APRNs affiliated with each PBRN were sent an illustrated brochure describing the study and inviting them to participate. A member of the research team gave an on-site demonstration of the PEP-NG software and study materials to APRNs interested in participating in the study. Once recruited, practices were offered free installation of a wireless-access node (meeting HIPAA requirements) and a free tablet computer (in addition to the unit used for the study) as incentives for participation. APRNs were offered $80 to compensate for their completion of the 2-h, on-site PEP-NG study training. They were also offered 10 continuing education units (CEUs) for reading 10 journal articles (related to potential adverse effects caused by unsafe patient self-medication behaviors) and subsequently completing the pre-and post-training instruments. APRNs (or the primary care practices) were also offered $55 for each participant enrolled (up to 24 participants). This payment was to compensate for the approximately 40 min of time needed to ascertain study eligibility, conduct the informed consent process, show the online tutorial to the patient, keep the participant gift-card receipts, and file recruitment reports for each patient participant.
Practices associated with the PBRN networks entered the study in an ongoing basis. A member of the research team who was in a post-masters' adult nurse practitioner program conducted the 2-h on-site training session with each APRN. Each APRN was given a research notebook with a step-by-step study protocol, instruments for assessing study eligibility, record sheets for documenting each visit, grocery gift cards, and study appointment cards. Illustrated participant recruitment brochures and posters with the APRNs' names and practice contact information were placed in waiting and examination rooms. Older adults self-referred for the study by calling the practice and making an appointment with the APRN. The APRN met with each prospective participant to review the consent form (written in an Arial 14 font at a grade-6 reading level). Participants were requested not to participate in another research study related to their health while enrolled in the PEP-NG study.
After attaining patient consent to participate, the APRNs used the following inclusion criteria to assess study eligibility: 1) not previously involved in a PEP study; 2) at least age 60 (by self-report); 3) a health literacy score of at least 44 (6th grade) as measured by the Rapid Estimate of Adult Literacy in Medicine (REALM) tool (Davis et al. 1993(Davis et al. , 1998)); 4) currently taking prescribed antihypertensive medication; and 5) independent-living and cognitive-functioning ability. The latter was reflected by the older adult's ability to: a) independently manage the tasks of telephone communication, shopping, travel arrangements, self-medicating, and finance activities, as assessed with the Instrumental Activities of Daily Living Scale, (Lawton and Brody 1969); b) successfully answer 6 of 10 items on the Short Portable Mental Status Questionnaire, (Pfeiffer 1974); and c) live independently. Eligible patient participants also needed to demonstrate a visual acuity of at least 20/100 (with corrective lenses, if needed).
APRNs selected a four-digit random number from a list (provided by the study) as the log-in ID for each participant. The APRN also selected a random number for the APRN log-in ID and another for the site ID. In order to minimize confounding effects due to the heterogeneity among APRNs and site-patient populations, the PEP-NG randomly assigned participants within each site to either the control or intervention groups. APRNs mailed the PI monthly monitoring reports with the numbers of patients, using the following metrics: a) screened for participation; b) met and not met study criteria; c) enrolled; d) dropped out of the study; and e) experienced adverse or unexpected effects such as anxiety or eye strain. APRNs were also asked to immediately report any adverse events to the PI.
Before the on-site training, APRNs logged on to a dedicated website to complete pre-training Rx-OTC knowledge, Rx-OTC self-efficacy, and Eldercare self-efficacy instruments. During the training session, APRNs tested the separate patient and provider interfaces of the PEP-NG. They were also given a packet of 10 articles, written by the PI, documenting the evidence that underlies the specific adverse medication behaviors addressed by the PEP-NG. After reading these articles (over the next 2 weeks), the APRNs logged on to an APRN-dedicated website to complete post-training knowledge and self-efficacy measurement instruments. The APRNs completed the post-training instruments at two different times-after successfully enrolling their sixth participant (typically 3 months later) and their twelfth participant (typically 6 months later), respectively.
Participants met individually with their APRN four times over 3 months in a private examination room at the practice site. Participants were encouraged to bring all of their medications (including supplements) to each visit. The APRN took the patient's BP at the beginning of visit 1. By attaching a keyboard to the tablet, the APRN entered the participant's year of birth (confirmed from the medical record), gender, BP, and REALM health literacy score (Davis et al. 1993(Davis et al. , 1998) ) via the tailored APRN-provider interface. The APRN also entered each patient's prescribed and provider-recommended (e.g. low-dose aspirin) medications, including, dose, timing, and any special instructions for taking the medication.
Upon completion of patient data entry, the APRN removed the tablet from the keyboard and set it on a height/angle adjustable stand to ready the tablet for patientparticipant use. The APRN read a tutorial script to the participant, while the participant practiced using a stylus to touch the interface and sample screens (including a question, a medication screen, a "clock" screen, and an interactive animation screen). When the participant expressed comfort with the patient interface, the APRN left the patient to begin the PEP-NG interface task independently.
On visit 1, participants completed demographic questions concerning living environment (with whom they live, type of residence), education, race/ethnicity, income (e.g. whether their monthly income is at, above, or below $1,500 per month), as well as health questions about current medical problems and symptoms. The patient then completed all measurement items (except the satisfaction survey) and responded to questions about the medications and OTC agents they take for treating their blood pressure and common health problems. On visits 2-4, before asking the patient to continue with the PEP-NG unassisted, the APRN reviewed patient comfort with the stylus use and PEP-NG interface as needed. Demographic questions were omitted on visits 2-4. On visit 4, participants completed the patient-satisfaction instrument, in addition to the other scales and questions measured during visits 1-3. After each PEP-NG use, the participant visited with the APRN for approximately 15 min. During the visit, the APRN took the participant's BP, based on the JNC-7 standards (Chobanian et al. 2003). The APRN then recorded the BP reading and reviewed/updated any changes in the medication regimen on the provider interface.
Participants in the intervention group received tailored education in the following manner. The PEP-NG rules engine analyzed patient-entered information and delivered educational content tailored to the three patient-reported behaviors associated with the highest risk scores. The education components included: 1) animations and "medicine facts" that illustrated and described the adverse behaviors identified; 2) "what you can do" tips which offered corrective strategies; and 3) interactive questions that allowed the user to rehearse and apply the information learned. A printout generated by the patient-reported data on the PEP-NG listed patient-reported symptoms, the three identified adverse self-medication behaviors and corrective strategies suggested by the PEP-NG, along with thumbnail screen shots from the animations. In the case of fewer than three reported adverse behaviors, the PEP delivered a set of up to three default statements dealing with medication adherence, OTC pain relievers (that can be safely taken with antihypertensives), and dangers of combining different types of pain relievers (prescription and OTC). A copy of the printout was also given to the APRN to help inform the patient visit as described above.
Like their counterparts in the intervention group, participants in the control group were asked to complete all questions via the PEP-NG. They also received a general education message, an interactive animation, and an interactive question at the end of each session, which highlights how BP medicines work and emphasizes how BP medications must be taken every day. These participants did not receive a printout at the end of each of their PEP-NG uses or APRN visits.
Participants in both the intervention and control groups were offered a $10 grocery gift card at the end of each of the first three visits and a $25 grocery gift card at the end of the fourth visit to compensate for time in the study. At the end of the fourth visit, the patient was given a card with a dedicated telephone number to callif he or she wished to schedule a 20-min qualitative follow-up interview to be conducted at the practice with a nurse researcher. The patient was given an additional $10 grocery gift card for participating in the post-trial interview. The APRNs were also invited to participate in the post-study interview and given a $25 grocery gift card as reimbursement for their time.
The Adverse Self-medication Behavior Risk, OTC-Rx Knowledge, OTC-Rx Self-Efficacy, Eldercare Self-Efficacy, and Healthcare Relationships and Satisfaction Scales were previously validated, along with a description of their individual psychometric properties (Anderson and Spencer 2002;Neafsey 1997;Neafsey and Shellman 2002a, b;Neafsey et al. 2002Neafsey et al. , 2001Neafsey et al. , 2009;;Shellman 2006). All instruments were written at a 6th-grade Flesch-Kincaid reading level (Flesch 1968). The primary-patient outcome measure was patient Adverse Self-Medication Behavior Risk Score. BP control, OTC-Rx Knowledge, OTC-Rx self-efficacy, and health care relationships and satisfaction with the PEP-NG were secondary outcome measures for patient participants.
BP measurements were taken by the APRN at each visit-at the beginning of PEP use on visit 1, and post-PEP use on subsequent visits. Inadequate BP control was defined as: 1) a SBP ≥ 140 mm Hg or a DBP ≥ 90 mm Hg for patients without diabetes; and 2) a SBP ≥ 130 mm Hg or a DBP ≥ 80 mm Hg for patients with diabetes or chronic kidney disease, per the JNC-7 guidelines (Chobanian et al. 2003).
Adverse self-medication behaviors were identified from questions that address use of medications (in the past month) to treat high blood pressure as well as use of OTC agents and alcohol for problems that were self-treated with non-prescription agents (e.g. pain, fever, colds or sinus, allergies, sleep, indigestion, gas, constipation). Participants were also asked if they drank alcoholic beverages, smoked or used nicotine, or took any vitamin or mineral supplements (including what, when and how frequently each was taken). The Adverse Self-Medication Behavior Risk Score is the weighted sum of the scores for the adverse behaviors identified (Neafsey et al. 2009).
The OTC-Rx Knowledge scale has 14 multiple-choice items and the score is the percent of the items with correct response; these items test both knowledge and application concerning potential adverse effects of self-medication with OTC agents, supplements, or alcohol in persons with hypertension. The OTC-Rx Selfefficacy scale is a 12-item instrument with statements reflecting patient confidence in selecting appropriate OTC agents and supplements, aside from avoiding adverse effects arising from self-medication behaviors. This scale has 5-point self-report response categories (ranging from 1, "Not Sure" to 5, "Totally Sure"). Responses were summed and divided by the number of items answered, so that the overall score would not be affected by omitted items and was reported based on the original 5-point metric.
The Eldercare Self-Efficacy instrument is a 7-item, 5-point Likert-type scale that assesses APRN self-efficacy in communicating with older adults about their medications (Shellman 2006;Neafsey et al. 2009). The Health Care Relationships Instrument is a 5-item instrument for patients, gauged with a 5-point Likert-type scale (ranging from "not at all easy" to "very easy") that measures patient-provider communication (two questions), trust in provider, participation in decision-making related to care, and satisfaction with care (Anderson and Spencer 2002).
The PEP-NG user Satisfaction scale is a 14-item instrument-with eight items addressing the ease of program use, program content, and suitability of program content-and another six items addressing the intent to change behavior following program use. Ratings reflected by the 5-point Likert-type scale (ranging from 1, "strongly disagree" to 5, "strongly agree") were summed and divided by the number of items answered to ensure that the overall Satisfaction scale was not affected by omitted items and was cast in the original 5-point metric.
A complete description of the methods of data analysis and statistical power considerations are published elsewhere (Neafsey et al. 2009). The study design involved three factors: PEP-NG intervention vs. control, time (evaluation at baseline and at three subsequent time points), and APRN (10 advanced practice nurses). Variation in outcome measures between APRN, while likely to occur, was of secondary interest, therefore the APRN factor was treated merely as a source of random effects in statistical modeling and hypothesis testing. The principal study hypothesis concerned the possibility of differential change in the PEP-NG and control conditions between baseline and final study assessments. Although changes could be contrasted between the PEP-NG and control groups for many outcome measures, the comparison of changes in adverse self-medication risk score was of primary interest. Statistical power analysis showed that a total sample of 164 subjects (82 per group) would be sufficient to yield 80% power to detect a 5-point difference between the intervention and control groups in mean changes from baseline to visit 4 in the adverse self-medication risk score.
Repeated measures linear-mixed model ANOVA methodology was the basis of most statistical tests and effect estimations. There are three reasons why this more complex analysis tool was used instead of the simpler traditional least-squares ANOVA technique. First, each participant was measured repeatedly over time (at each of four visits); these four repeated measures are likely to be dependent, and therefore an appropriate covariance-structure must be selected to capture this feature of the data. Unlike traditional ANOVA, with mixed-model ANOVA it is no longer necessary to assume the covariance structure adheres to compound symmetry or sphericity; instead, one can choose from a large collection of covariance-structures as appropriate to the observed data. Second, patients who visit the same APRN might tend to have similar demographics characteristics, while patients across APRNs might tend to have different characteristics. Such clustering effects need to be adjusted through the estimation of random effects associated with the APRNs. Third, traditional linear model techniques drop an entire participant from the analysis if the subject has missing data at one visit, while linear mixed models allow subjects to have missing visit values. Hence, linear mixed models provide researchers with a powerful and flexible analytic tool for these kinds of data. The SAS software package (v. 9.2; SAS Institute, Inc., Cary, NC) allows a full implementation of this analytical tool through its Proc Mixed procedure (Brown and Prescott 2006;Verbeke and Molenberghs 2000;West et al. 2006).
The linear mixed models used study group (intervention vs. control) and visit (1-4) as categorical variables and controlled for gender, age, APRN, income, education, and computer use. Adjustment for computer use was performed because this variable was significantly different between study groups at visit 1. Initially, the models considered the possibility of an interaction effect between study group and visit on outcome measures. However, whenever this interaction was not found to be significant, it was dropped from the model. The potential impact of co-linearity between the income and education covariates was considered before identifying the final model for each outcome measure.
Paired-t tests were used for post hoc analyses of knowledge, self-efficacy, and BP within groups. Non-parametric tests (Cochran-Mantel-Haenszel statistic), based on rank scores, controlling for participant code were used for the transformed behavior risk scores within groups. As age is significantly inversely correlated with DBP (Chobanian et al. 2003), correlations between outcome measures were conducted while controlling for age. Cross-tabulations comparing the frequencies of dichotomous assessments relative to study group were created and subjected to statistical testing through either the Pearson chi-square test or Fisher's exact test.
Fifteen provider practices and 20 APRNs consented to the study. Five practices and five APRNs withdrew soon after the installation of the wireless access nodes for reasons unrelated to the study (including APRN illness, APRN job change, and practice-location change).
Fifteen APRNs enrolled in the study and completed training. Three APRNS withdrew from the study after training and before patient enrollment; two of them for other jobs and one due to illness. An additional APRN withdrew after enrolling one participant (who did complete the four visits). The patient data from this APRN were removed from the analyses. The primary care practices had widely different patient demographics and practice characteristics. The participating practices were located in two urban centers, three small cities, two suburbs, and two rural areas. Eight of the APRNs were salaried, two were paid by the number of patients seen, and two were paid by the hour. All of the APRNs were Caucasian, with a mean age of 44.54 (9.71) (range 31-60 years). The mean APRN practice years was 8.4 (6.67) (range 1-23), and the mean nursing practice years was 18.3 (9.76), range 6-38.
Data were missing for four APRNs on the fourth observation (after the 12th participant was enrolled in each site). Therefore, APRN outcomes were analyzed from pre-training to post-training after the sixth participant was enrolled at each site (approximately 3 months post-training). APRN scores (N=11) on the Rx-OTC knowledge and Rx-OTC self-efficacy scales increased from baseline to after the APRN enrolled the sixth participant. Rx-OTC Knowledge increased from 67.7% (11%) pre-training to 80.9% (12%) post-training (two tailed t=2.94, p=.014). Rx-OTC Self-Efficacy increased from 3.82 (.61) pre-training to 4.13 (.47) 3 months post-training (two tailed t=2.49, p=.016). The mean score on the 5-point Eldercare Self-Efficacy Scale during pre-training (N=11) was 3.28 (.66), and it rose 3.67 (.67) after the sixth participant enrolled (approximately 3 months later); the increase was statistically significant (two tailed t=2.37, p=.039).
A total of 164 patient participants were screened for eligibility, 160 were eligible. Two patients died during the course of the study (both in the control group). Ten (6.25%) withdrew at various times during the study (five from the control group and five from the intervention group). The baseline characteristics of participating patients are shown in Table 1. There were no significant differences between the control and intervention groups for any of the demographic variables. Baseline characteristics of the entire patient sample are described as follows. Patients had a mean age of 68.59 (8.71) and were predominantly US born (89%), Caucasian (93%) and females (78%). Twenty-three percent of them had monthly household incomes at or below $1,500. Their mean REALM scores were at the high school level (grade 10-12) and 92% had at least a high school diploma or GED. BP Measurements administered during the pre-intervention period on visit 1 revealed that 31% of all participants were not at JNC-7 BP targets.
More than half of the participants (53.7%) had three or more chronic conditions. The most common co-morbidities were high cholesterol (37.5%), arthritis (35.6%) anxiety (21.8%) and diabetes (16.25%). The majority of participants (66.8%) rated their health as very good or excellent during the preceding month. Most reported (63.2%) living with someone else and 36.8% reported living by themselves. While a med box was reported as the tool used most often (52.5%) to help them remember to take their medications, 37.5% of the participants selected, "I just remember," as their response.
The most common symptoms (during the last month reported by 20% or more of participants) were pain, fatigue, difficulty sleeping, allergies, heartburn, cough, leg cramps, anxiety and a cold. Medication labels and the pharmacist were reported as the most common sources of medication information, followed by the doctor, medication insert (dispensed by the pharmacy or the medicine package), and nurse. Participants indicated that they buy OTC medicines primarily in grocery stores and discount stores with pharmacists. Fewer than 5% of these participants reported buying OTC medicines over the Internet.
A majority (73.6%) of participants took five or more Rx medications in the past month, with a mean of 6.93 (3.41). When OTC agents were included, nearly all (98.1%) of the participants reported taking five or more different medications in the past month, with a mean of 11.41 (4.31). More than 12 medication doses per day were taken by 9.5% of the participants. The most common medications reported are profiled in Table 2. Self-reported medication adherence for antihypertensives was high, with 89% or better daily adherence reported for each category of antihypertensive. None of the participants reporting less than daily adherence on all antihypertensives were at BP targets upon study entry.
The mean number of antihypertensive medication formulations prescribed to the study participants upon entry to the study at visit 1 was 1.52 (0.83), with a range of 1-5. One antihypertensive medication formulation was prescribed to 63.9% of the participants, 24.3% were prescribed two, 8.11% were prescribed three, 2.79% were prescribed four, and 0.84% were prescribed five antihypertensive medication formulations.
A total of 44 different antihypertensive formulations were prescribed to the study participants upon entry to the study at visit 1. The most common were: lisinopril (10%), hydrochlorothiazide (8.1%), Toprol XL® (8.1%), atenolol (6.25%), Diovan® (4.4%), and Coreg® (4.4%). The most common antihypertensive categories prescribed were calcium channel blockers (41.9%), angiotensin converting enzyme inhibitors (ACEIs) (38.8%), beta blockers (29.4%), angiotensin II receptor blockers (ARBs) (28.1%) and thiazides (20.0%). Treatment with a single antihypertensive agent (monotherapy) was prescribed to 65.3% of the patients upon entry to the study at visit 1.
Among participants not at BP targets upon study entry at visit 1, the mean number of antihypertensive medication formulations was 1.45 (0.82), with a range of 1-4. One Child 0 (0) 0 (0) 0 (0) Relative 1 (0.6) 1 (1.4) 0 (0) Friend 0 (0) 0 (0) 0 (0) Nurse 0 (0) 0 (0) 0 (0) Other 0 (0) 0 (0) 0 (0) Where buy OTC medicines Varies 0 (0) 0 (0) 0 (0) Internet 7 (4.4) 3 (4.1) 4 (4.6) Out of Country 0 (0) 0 (0) 0 (0)
Store with a pharmacist 129 (80.6) 59 (80.8) 70 (80.5) Store without a pharmacist 26 (16.3) 11 (15.1) 15 (17.2) Other 11 (6.9) 4 (5.5) 7 (8.1) Grocery store 76 (47.5) 36 (49.3) 40 (46.0) Discount store 54 (33.8) 19 (26.0) 35 (40.2) Health food store 20 (12.5) 8 (11.0) 12 (13.8) Convenience store 6 (3.8) 3 (4.1) 3 (3.5) Other 47 (29.4) 23 (31.5) 24 (27.6)
antihypertensive formulation (containing one or more antihypertensive agents) was prescribed to 71% of the patients, 18.4% were prescribed two formulations, 5.3% were prescribed three formulations and 5.3% were prescribed four antihypertensive medication formulations. Treatment with a single antihypertensive agent (monotherapy) was prescribed for 57.9% of patients not at BP targets upon entry to the study at visit 1. (Of those participants who were at goal upon entry at visit 1, 72.5% were on monotherapy). In terms of offline media exposure, participants from the control and intervention groups appeared to consume radio, television, newspaper and magazine content on a similar number of days per week (3.6 vs. 4.2). Even though fewer participants from the control group were PC users than the intervention group (58.9% vs. 70.11%), this discrepancy was not statistically significant (Χ 2 (1, 157)=2.52, p=.1118) and both groups used a personal computer on 6.6 days per week (or nearly 7 days a week) and about 3.3-3.4 h per day. There were significantly fewer Internet users in the control group compared to the intervention group (50.7% vs. 67.6%) (Χ 2 (1, 157)=6.89, p=.0086), but the number of days per week (5.67 vs. 6.01) and hours per day (2.24 vs. 2.56) each group of Internet users went online was not significantly different. The number of days these two Internet user groups used the email system per week (4.82 vs. 4.97) was not statistically
Normality of the behavior-risk score was tested with the Shapiro and Wilk's W statistic. The skewness value was 1.4 and the kurtosis value was 2.99. The W statistic was 0.877, and normality was rejected at the 0.05 level of significance. The SAS boxcox method determined that the square-root function (square-root risk) was the best transformation of the behavior-risk score, producing a skewness value of 0.25, a kurtosis value of -0.66, and a W statistic of 0.96. Consequently, the squareroot risk (transformed behavior-risk score) was used in all analyses. While there was no significant condition by visit interaction for the transformed behavior-risk score, main effects were significant for visit (F (3, 418, p<.0180). There was a significant linear trend (F (1, 418)=5.18, p=.0233), but there was no significant quadratic or cubic trend. Results of post hoc paired t-tests and nonparametric tests showed a significant reduction in the transformed behaviorrisk score for the intervention group from visit 1 to visit 4 (paired t (73)=2.17, p=.033; Cochran-Mantel-Haenszel statistic (1, 161)=4.44, p=0.035).
Knowledge scores were normally distributed. The interaction between visit and condition was significant (F (3, 411)=7.15, p=.0001). Main effects were significant for computer use (F (1, 139)=5.67, p=.0186), gender (F (1, 139)=24.1, p<.0001), condition (F (1, 139)=20.51, p<.0001), and visit (F (3, 411)=10.07, p<.0001). There was a significant linear trend for visit alone (F (1, 411)=26.76, p<.0001) and for visit within the intervention group (F (1, 411)=36.51, p<.0001). Post hoc analyses detected significant differences between the control and intervention group at visits 2, 3 and 4, with a large effect size on visit 4 (Cohen's d=0.877). Within the intervention group, there was a significant difference between visit 1 and 2, 2 and 3, and 3 and 4. There was also a significant increase in knowledge from visit 1 to visit 4 for the intervention group (paired t (73)=6.26, p<0.0001). Post hoc paired t-tests showed no significant change in knowledge from visit 1 to visit 4 for the control group. The main effect of being a computer user resulted in a larger knowledge score overall of 5.2 in the percent correct. By contrast, the main effect of gender resulted in males having a lower knowledge score overall of |-9.97| in the percent correct. Rx-OTC knowledge scores were significantly correlated with Transformed Behavior risk score for the control group on visit 4 (Spearman r (141)=.41, p=.0007).
Self-efficacy scores were normally distributed. The interaction between visit and condition was significant (F (3, 412)=9.33, p<.0001). Main effects were significant for condition (F (1, 139)=9.67, p=.0023) and visit (F (3, 412)=23.68, p<.0001). There was a significant linear trend for visit alone (F (1, 412)=70.52, p<.0001), in addition to visit within the both the control group (F (1, 412)=5.15, p=.0237) and the intervention group (F (1, 412)=61.70, p<.0001). Post hoc analyses detected significant differences between the control and intervention group at visits 2, 3 and 4, with a large effect size on visit 4 (Cohen's d=1.06). There were significant differences between visits 1 and 2, 2 and 3, as well as 3 and 4 in the intervention group. Post hoc paired t-tests showed a significant increase in self-efficacy from visit 1 to visit 4 for the intervention group (t (73)=10.38, p<0.0001). There was no significant change in self-efficacy from visit 1 to visit 4 for the control group.
With the exception of Rx-OTC knowledge scores and Transformed Behavior Risk scores on visit 4 in the control group as described above, correlations were weak among Transformed Behavior Risk, Rx-OTC knowledge scores, Rx-OTC selfefficacy scores, and education. These variables were not correlated with SBP or DBP at any visit for either controls or intervention patients.
Both SBP and DBP values were normally distributed. There was no significant interaction between visit and condition for either SBP or DBP. Main effects on SBP were significant for age (F (1, 139)=4.22, p=.0419), income (F (1, 139)=6.41, p=.0125), and condition (F (1, 139)=5.50, p=.0204). There was no linear, quadratic or cubic trend. The main effect of age resulted in a larger SBP overall of 0.25 mm Hg per year (i.e. an increase of 10 years in age added 2.5 mm Hg overall to SBP). The main effect of higher income (above $1,500 per month) was a lower SBP overall of -2.68 mm Hg. The main effect of being in the control condition was a higher SBP overall of 3.6 mm Hg. Post hoc paired t-tests did not reveal any significant differences in the SBP of visit 1 with visits 2, 3 or 4 within either group.
Main effects on DBP were significant for computer use (F (1, 139)=4.73, p=0.0314) and condition (F (1, 139)=5.69, p=.0184). There was no linear, quadratic or cubic trend. The main effect of being a computer user resulted in a higher DBP overall of 2.37 mmHg. Moreover, the main effect of being in the control condition was a larger DBP overall of 2.18 mmHg. Post hoc paired t-tests did not reveal any significant differences in the DBP of visit 1 with visits 2, 3 or 4 within either group.
The percentage of control group participants not at their BP target went from 32.9% on visit 1 to 31.3% by visit 3. This represents a 4.8% reduction in the percentage of control group participants not at BP targets at visit 3 compared to visit 1. The percentage of control group participants not at BP targets was at 35.3% by visit 4 (an increase of 7% compared to visit 1). The percentage of intervention participants not at their BP target went from 29.9% on visit 1 to 21.8% by visit 3. This represents a 27% reduction in the percentage of intervention participants not at goal at visit 3 compared to visit 1. The percentage of intervention participants not at goal was at 26% on visit 4 (a reduction of 13% compared to visit 1).
Satisfaction scores are shown in Table 4. Participants in both groups indicated a high level of satisfaction with aspects of the PEP-NG interface and program. Compared to the control group, the intervention group had significantly higher overall mean satisfaction scores (t (132)=2.04, p=.0431). Post hoc t-tests revealed significantly higher ratings in the intervention group on two satisfaction items: "Much of the information in the program was new for me," (t (132)=3.83, p=.0002) and "The advice in the program suited my special needs," (t (132)=3.14, p=.0016).
The intervention group had a significantly higher "mean intent-to-change" score compared to the control group (t (132)=2.16, p=.033). Post hoc t-tests revealed a significantly greater score for the intervention group on three intent-to-change items: "This program helped me to want to change how I use medicines," (t (132)=4.04, p=.0001), "After using this program I will make some changes in how I use medicines," (t (132)=3.424, p=.0008), and "After using this program I will change when I take some medicines," (t (132)=3.60, p=.0005).
Scores (SD) on the 5-point Likert type Healthcare Relationships Scale were similar for the control group (4.46 (0.83), n=69) and the intervention group (4.40 (0.68) n=82) participants upon entry to the study on visit 1. On visit 4, scores (SD) remained high with means of 4.42 (0.63) for the control group (n=65) and 4.34 (0.60) for the intervention group (n=75).
Sub analyses conducted for the study participants in both groups not at BP targets upon entry to the study are given in Table 5.
There was a significant visit effect for transformed behavior risk score for those participants not at BP targets upon entry to the study (F (3, 125)=2.75, p=.0454). There was no interaction with condition or gender. While there was a significant linear trend (F (1, 125)=4.31, p=.040), there was no quadratic or cubic trend. There was a 49.8% reduction in the Behavior Risk Score on visit 4 for the intervention patients who were not at BP targets upon entry to the study. Results of post hoc paired t-tests and nonparametric tests showed a significant reduction in transformed behavior risk score for the intervention group from visit 1 to visit 4 (paired t (21)=2.41, p=.0253; Cochran-Mantel-Haenszel statistic (1, 48)=5.76, p=.0164).
There was no significant interaction between visit and condition for SBP among participants not at BP targets on study entry. Main effects on SBP were significant for income (F (1, 34)=4.44, p=.0425) and visit (F (3, 129)=9.54, p<.0001). There was a significant linear trend for visit alone (F (1, 129)=20.26, p=<.0001), and a significant quadratic trend for visit alone F (1, 129)=7.14, p=<.0085). There was also a significant linear trend for visit within both the control (F (1, 126)=6.27, p= .0136) and intervention (F (1, 126)= 21.53, p =<.0001) groups. The main effect of higher income (above $1,500 per month) was a lower SBP overall of |-5.2| mmHg (both groups considered together). Post hoc paired t-tests showed significant reductions in SBP from visit 1 to visit 4 in both the control (-6.11 mm Hg) (t (20)=2.22, p =.0380) and the intervention (-15.51 mm Hg) (t (21)=5.95, p <.0001) groups. The intervention group had a medium effect size in reducing SBP (Cohen's d =.2548) among patients not at BP targets upon entry to the study. Post hoc paired t-tests also showed significant declines for SBP in the intervention group from visit 1 to visit 2 and 3. Main effects on DBP were significant for visit (F (3, 129)=3.48, p=.0178). There was a significant linear trend for visit alone (F (1, 129)=8.04, p=.0053). There was also a significant linear trend for visit within both the control (F (1, 126)=4.68, p=.0324) and intervention (F (1, 126)=3.92, p=.0498) groups. The decline in DBP (-2.31 mm Hg) in the control group from visit 1 to visit 4 was not statistically significant. By contrast, the intervention group had a 2.70 fold greater decline in DBP (-6.26 mm Hg) from visit 1 to visit 4 compared to the control group, which was statistically significant (paired t (21)=3.70, p=.0013).
Among all patients (both groups) at BP targets upon entry to the study who reported being Internet users, Internet Self-efficacy was significantly correlated with Rx-OTC self-efficacy on visit 4 (Spearman r (50)=.4417, p=.0015). Among the control group participants not at BP targets upon entry to the study, Transformed Behavior Risk score was significantly correlated with SBP on visit 2 (Spearman r (21)=.51298, p=.0207) and Rx-OTC Knowledge scores on visit 4 (Spearman r (22)=.64312, p=.0017).
In patients who were not at BP targets at study entry, movement to controlled BP was uneven, but not significantly different, between the intervention and control groups. Among 22 patients not at BP targets in the intervention group at visit 1, 14 (63.6%) had controlled BP at visit 4. In contrast, among 21 patients initially not at BP targets in the control group, only 8 (38.1%) had controlled BP at visit 4. The odds ratio comparing frequencies of these changes between groups was 2.84 (Χ 2 (1)=2.81, p=0.094, 95% CI: [0.83, 9.80]).
Adherence was considered to be the percentage of patients reporting taking all of their antihypertensive medications "daily." Of patients not at BP targets on visit 1 (both control and intervention groups), 93% self-reported taking all of their antihypertensives daily upon entry to the study. Self-reported daily adherence rose to 100% of patients in both groups of patients on visit 2 and remained at 100% of patients on the ensuing 2 visits for the intervention group, while 95% of control patients reported daily adherence on visits 3 and 4. Among control patients who were at BP targets upon entry to the study, self-reported adherence declined over time to 87% by visit 4. Self-reported daily adherence stayed close to 95% or greater during the subsequent three visits for intervention patients who were at BP targets on study entry.
More than half of patients not at BP targets upon entry to the study reported NSAID usage in the previous month (54.2% of control and 57.7% of intervention patients). Patients in the intervention group reduced self-reported NSAID use over time from 57.7% upon study entry to 9.09% on visit 4, while 42.8% of the control group still reported NSAID usage at visit 4. Monthly NSAID use was categorized as: 1) 1-10 days/month; 2) 11-20 days/month; 3) 21-31 days/month; and 4) other. Results of a Cochran-Armitage test for trend detected a significant declining linear trend in proportion of the intervention patients in category 3 (taking NSAIDs 21-31 days/month) (Z (35)=2.5023; Exact test one sided p=.0084).
Alcohol has a pressor effect in a dose-related manner beginning at >2 drinks/day (World Hypertension League 1991). Among patients not at BP targets upon entry to the study, 41.4% of controls and 57.7% of intervention patients reported daily alcohol consumption. None of the control-group patients and two of the intervention-group patients reported having three or more drinks daily in the month prior to study entry. Both of these intervention-group patients reduced their selfreported daily alcohol consumption to one drink/day. (One reduced to one drink per day by visit 2, the other reduced to one drink by visit 3. Both of them reported one drink per day on visit 4).
Decongestants can increase BP to a variable degree-depending on agent, dose, and duration of use (Johnson and Hricik 1993;Kollar et al. 2007;Salerno et al. 2005)-but effects in older adults have not been well studied. Five (20.8%) of the control group patients and 6 (23.1%) of the intervention group patients not at BP targets upon study entry reported taking a medication containing a decongestant on visit 1. Of these patients, (40%) in the control group and three (50%) in the intervention group reported taking the decongestant medication daily. On visit 4, 6 (28.6%) in the control group and 3 (13.6%) in the intervention group reported taking a medication containing a decongestant in the previous month. The ability of the PEP-NG to capture decongestant taking behaviors was limited for the following reasons: 1) decongestant taking behaviors may be episodic in nature due to seasonal allergies and colds; 2) OTC cold medications underwent reformulation during the trial period to replace pseudoephedrine with phenylephrine; and 3) study entry was on a rolling basis. Therefore, data on decongestant taking behaviors were not analyzed for time trends.
Table 6 shows changes made in patient antihypertensive regimens during the study period for patients not at BP targets on visit 1. For controls, 11 (45.8%) patients not at BP targets on visit 1 had either an increase in dose or an added antihypertensive Rx. Only one intervention patient (3.85%) had an added antihypertensive Rx and none had a dose increase. For controls, 2 (8.33%) of patients had either a decrease in dose or an antihypertensive Rx discontinued. For intervention patients, 5 (19.32%) had either a decrease in dose or an antihypertensive Rx discontinued. A Fisher's exact test on therapy intensification (the number of patients receiving an increased antihypertensive dose and/or an additional antihypertensive) was significant (p=.001) with an odds ratio of 21.27, 95% CI: [2. 45,200] for the control group compared to the intervention group.
Providers may recommend low-dose aspirin for patients with hypertension because of its antiplatelet effects and its contribution to lowering BP in the morning hours, if taken in the evening prior (Hermida et al. 2005). Resistance to the antiplatelet effects of low-dose aspirin may be due to low adherence rates to low-dose aspirin therapy, genetic polymorphisms, or high platelet counts (Tran et al. 2007). Another potential cause of aspirin resistance is frequent concurrent use (3 days or more per week) of NSAIDs (Gladding et al. 2008;MacDonald and Wei 2003;Tran et al. 2007). As shown in Table 2, 31.5% of control-group patients and 50% of intervention-group patients reported taking daily low-dose aspirin at visit 1 upon entry to the study. Of patients taking daily low-dose aspirin, 44.7% of control and 45.4% of intervention group patients reported taking NSAIDs on a frequent basis of 3 days a week or more (26.3% of all control and 36.4% of all intervention patients who reported taking daily low-dose aspirin also reported taking an NSAID daily). Among patients not at BP targets upon entry to the study, 71% of controls and 65% of intervention group patients reported taking daily low-dose aspirin at visit 1 upon entry to the study. Of these, 58.8% of control and 64.7% of intervention group patients also reported taking NSAIDs on a frequent basis of 3 days a week or more; 29.4% of these controls and 52.9% of these intervention patients reported taking an NSAID daily. By visit 4, 69.2% of the controls and none of the intervention participants who indicated taking daily low-dose aspirin reported frequent NSAID use of 3 days a week or more (30.8% of control patients who reported taking daily low-dose aspirin still reported taking an NSAID daily on visit 4). Categorizing monthly NSAID use as described above and applying the Cochran-Armitage test for trend detected a significant declining linear trend across the four visits, in proportions of intervention participants who took NSAIDs 21-31 days/month (Z (29)=2.6049, Exact test one sided p=.0062). There were no significant trends in NSAID use among patients in either group who were at BP targets upon entry to the study, regardless of low-dose aspirin use.
Taking NSAIDs with the combination of diuretics and either ACEIs or ARBs increases the risk of renal impairment in older adults (Juhlin et al. 2005;Loboz and Shenfield 2004). Among all participants in the study, 15.7% took either an ACEI or ARB and took a diuretic and NSAIDs concurrently in the previous month on visit 1. By visit 4, this combination was taken by 23.5% of control-group patients, compared to 12% of intervention-group patients. Among patients not at BP targets upon entry in to the study, 33.3% of the control group and none of the intervention-group participants were taking this nephrotoxic combination on visit 4. Among the controlgroup participants not at BP targets on visit 1-in five of seven cases where a Rx was added to the antihypertensive regimen-the Rx added was a diuretic; in four of these cases, the patient also reported taking an NSAID (other than low-dose aspirin) in the previous month. In two of the four cases where an antihypertensive dose was increased to control BP, the agents were diuretic/ACEI combinations and both of these participant also reported taking NSAIDs in the previous month.
Older adults' risk of potential adverse drug interactions (PADI) is greatly increased when they have three or more chronic illnesses, take five or more concurrent medications, ingest more than 12 medication doses taken per day, have a history of nonadherence, or take a drug that requires therapeutic monitoring (Isaksen et al. 1999).
In the current study, 53.7% of the participants had three or more chronic illnesses, 98% reported taking five or more different medications in the past month (combining both Rx and OTC medications), and 9.5% took 12 or more medication doses per day. This suggests that nearly all of the participants were at risk for a PADI. That 37% of patient participants selected "just remember" as the only means that helps prompt them to take their medications is disconcerting and supports past findings that older adults make infrequent use of medication management tools (Lakey et al. 2009).
Satisfaction with the PEP-NG interface and the APRN provider relationship were high in both the intervention and control groups. The intervention group had significantly higher scores on the "intent to change" subscale. Both groups in the current study had high mean scores on the "health care provider relationship" scale, indicating a high degree of trust in and satisfaction with their care, as well as feeling involved in their health care decisions and an ease and comfort in communicating with their APRN provider.
Compared to the control condition, participants in the intervention group demonstrated significant increases in both Rx-OTC self-medication knowledge and self-efficacy measures (with large effect sizes) via repeated measures. Moreover, patients in the intervention group did not need more interface time to complete the longer education/learning-outcome assessment content than those in the control group to achieve increased knowledge and self-efficacy. These results could be indicative of the user-friendly nature of the tailored PEP-NG system-which enables the users to tailor their own learning style and focus to maximize their learning outcomes-such as concentrating on remedying their riskiest self-medication behaviors (as revealed by the PEP-NG's built-in risk-score calculation metrics). These findings provide a strong testament to the advantage of an e-health intervention with tailored content and an interface design which was iteratively tested and validated for system usability and content usefulness (see Lin et al. 2009Lin et al. , 2010)).
The study evidence shows that SBP and DBP declined in both intervention and control-group patients who were not at BP targets upon study entry. However, the decrease in the intervention group was more than two-fold greater than that in the control group having both clinical and statistical significance. A report prepared for the Agency for Healthcare Research and Quality (AHRQ) documented the mean reductions in SBP and DBP as 4.5 mm Hg and 2.1 mm Hg respectively, across all studies examined and a variety of BP management strategies (Shojania et al. 2005). The present study found mean BP reductions of 15 mm Hg for SBP and 6 mm Hg for DBP among the intervention-group participants. According to unbiased estimates of efficacy from a recent meta-analysis of BP reduction and cardiovascular outcomes (Law et al. 2009), a reduction of 10 mm Hg in systolic BP or 5 mm Hg in diastolic BP reduces coronary heart disease events by 22% and stroke by 41%.
The manner in which patients moved to BP targets in the current study differed between the intervention and control groups. Providers intensified antihypertensive therapy by adding new antihypertensives and/or increased doses with an OR of 21.27 in the control group, compared to the intervention group. The intervention group changed behaviors that counteracted the efficacy of antihypertensives-by significantly decreasing NSAID use and intake frequency-and, in two cases, decreasing alcohol consumption. Self-reported adherence (taking all antihypertensives daily) reached 100% in the intervention-group participants who were not at BP targets upon entry to the study. Of particular concern is that therapy intensification in the control patients not at BP targets upon study entry resulted in placing 33 t% of those patients at added renal risk because of continuing, concurrent NSAID use. The American Geriatrics Society guidelines specifically note that this so called "triple whammy" (Loboz and Shenfield 2004) combination (ACEIs/ARBS, thiazides, and NSAIDS) should be strictly avoided (American Geriatrics Society 2009).
Upon entry to the study, 53.4% of participants reported taking NSAIDs in the previous month, a practice that should be avoided in older patients with hypertension, because NSAIDs both increase BP as well as counteract the efficacy of antihypertensives and low-dose aspirin (American Geriatrics Society 2009). The current study findings were somewhat dissimilar to the results reported by a recent cross-sectional homeinterview study. This particular study, a nationally representative probability sample of 3,500 older adults, examined the use of Rx and OTC agents by ranking specific medication ingredients (rather than categories of agents); none of the NSAIDs were mentioned among the top 20 most commonly used medications by study participants (Qato et al. 2008). Qato et al. (2008) also found lower rates of low-dose aspirin use among their study participants (28%), compared to the participants in the current study (41.9%).
Among intervention participants not at BP targets upon entry to the study, NSAID use was greatly reduced by visits 2-4, both in terms of numbers of patients reporting any NSAID use and frequency of use. This likely had a major influence on the greater reductions in both SBP and DBP in the intervention group. The findings that NSAID usage significantly declined only among intervention patients not at BP targets upon study entry, but not among control group patients or patients in either group who were at BP targets upon study entry, is an important one, suggesting that patients whose BP was under control were less likely to heed advice to avoid selfmedication with NSAIDs.
It is also of interest that none of the participants who reported less than daily adherence for all of their antihypertensives were at BP targets upon study entry. Patients who self-report nonadherence, as identified by their answer to the question-"In the last month, how often did you take your medications as your doctor prescribed?"-have as great a cardiovascular risk as patients who smoke or have diabetes (Gehi et al. 2007). While the PEP-NG may over estimate adherence and cannot provide precise adherence data, nonadherence documented on the PEP-NG printout can foster patient-provider communication about reasons the patient did not adhere to the prescribed therapy.
The study design consisted of four monthly APRN provider visits. This intervention design may be considered a study limitation, as clinical interventions involving repeated visits have resource and workflow barriers to widespread implementation in a clinical setting. Another potential limitation associated with this study is that the BP measurements were taken by the participating APRN and could be subject to observer bias. As patients self-referred to the study, the percentage of patients not at BP targets at study entry (31.2%) may not have been representative of the patient population at the practice sites and was not representative of the U.S. as a whole (Chobanian et al. 2003). In general, participants were predominantly female, Caucasians; they also had higher health literacy (REALM) scores, education attainment, and self-health ratings than the population of adults aged 60 and older with hypertension (Bennet et al. 2009). Therefore, participants may not reflect the general population of patients in the primary care practices with respect to either demographic characteristics or degree of adherence to their antihypertensive regimen. While the characteristics of the participants limit generalizability of study results, the findings do support the feasibility of a contenttailored e-health intervention with older adults in primary care practices.
Despite fairly homogeneous participant characteristics, participants did have income diversity with 23% reporting a monthly income at or below $1,500. The results suggest that income did not play a role in Rx-OTC knowledge and selfefficacy scores, adverse self-medication risk scores, or BP. Gender was found to have a main effect in Rx-OTC knowledge scores (with approximately a 10% lower knowledge score overall for males); this suggests that older men in the study were less likely to obtain information on OTC agents and their interactions with antihypertensives. Computer users had an overall 5% higher Rx-OTC knowledge scores; this indicates that prior computer-use experience may provide a small advantage in learning from computer-based education programs or seeking knowledge about medications and OTC agents from computer-based sources.
The finding that Rx-OTC knowledge was significantly correlated with transformed adverse self-medication risk scores for the control group on visit 4 suggests that knowledge alone does not necessarily transfer to safe medication-taking behaviors. This is because patient self-efficacy is a key motivating factor for behavioral adoption and change, as evidenced by the large effect size found in both knowledge and self-efficacy for the intervention group. These findings confirm Bandura's theory that knowledge and self-efficacy are separate domains and both are essential to effecting positive behavior change (Bandura 1997). Naturally, the role of patients' computer and Internet efficacy should also be considered when implementing an e-health intervention program. As suggested by past literature, while a large percentage of older adults have basic computer and Internet-use skills, they are also intrigued by and enthusiastic about becoming more computer and Internet literate (Alemagno et al. 2004;Lin et al. 2009;Nahm et al. 2004).
With the median time of a primary care visit at 14 min (Hing et al. 2006), providers lack the time to elicit patient medication-taking behaviors or conduct a comprehensive review of medications taken on a regular basis (Tarn et al. 2009). The provider support offered by the PEP-NG printouts (symptoms, Rx, and OTC agents taken, including frequency and timing) can free up time for the provider to engage in the sorely needed provider-patient communication by reinforcing the patient-tailored education outcomes derived from the PEP-NG interface. PEP-NG use implanted during patients' "waiting-room time" can identify their symptoms and those with PADIs, in addition to allowing them to initiate an e-health education experience that is tailored to their specific self-medication behaviors.
As demonstrated by the current study, an on-site e-health intervention combined with provider-patient follow-up communication has produced significant and positive health outcomes for study participants. As patients took only 25 min to interface with the PEP-NG on the two follow-up visits, which omitted demographic, media use, and satisfaction questions, a repeated e-health intervention procedure is both feasible and doable in a clinical setting. As the current PEP-NG intervention was able to generate large effect sizes on knowledge and self-efficacy, by omitting the knowledge and self-efficacy scales in follow-up visits, the concern about time efficiency should also be mitigated.
The study described herein is an efficacy trial in the "realistic setting" of primary care practice. Additional studies of provider workflow using the PEP-NG and health care utilization costs are under way. A cost-benefit analysis (CBA) using data from time-motion studies and 52-week health care utilization data (i.e. total number of provider visits, emergency room visits, and hospitalizations) following each participant's entry to the study is being conducted and will be reported separately.
Future study with the PEP-NG will involve implementation on a wider scale with both provider-site and in-home access to the PEP-NG portal. Patients who are Internet users will be allowed to complete the program at home with reports generated to their provider prior to their office visits. This could greatly reduce the numbers of patients who would need to use the PEP-NG on-site at the provider's office and permit wider implementation in a given primary care practice. The target audience of the PEP-NG will also be expanded to patients diagnosed with prehypertension (BP greater than 120/80 but lower than BP targets).
The patient gains found in the current clinical trial resulting from the e-health intervention suggest that the PEP-NG system can be instrumental in facilitating better patient care in the primary care setting. A cost-effective and time-efficient ehealth intervention system such as the PEP-NG can become a model for guiding other self-management and medication-adherence interventions directed at other major national health problems, such as diabetes, heart failure, etc., to help stem the escalating health care costs in our nation.
a 0-7 scale where 0 = never, 1 = very rarely, 7 = very often
a Post hoc two sample t tests detected significant differences (p<.05) between intervention and control groups at visits 2, 3, 4 b Post hoc paired t tests detected significant increases (p<.05) within intervention group between visit 1 and visits 2, 3 and 4 c Post hoc paired t tests detected a significant decrease (p<.05) in transformed behavior risk score between visit 1 and visit 4 differentiated
a Mean degree of agreement with statement; 1 = strongly disagree to 5 = strongly agree *Two sample t test significant difference (p<.05) between groups
* Post hoc paired t tests detected significant decrease (p<.05) compared to visit 1
The authors wish to thank
This work was supported by funding from the
Conflicts of Interest The University of Connecticut granted an exclusive license for the PEP-NG to AdhereTx Corporation on August 25, 2009. The University of Connecticut and Patricia J. Neafsey are stockholders of AdhereTx.
Open Access This article is distributed under the terms of the Creative Commons Attribution Noncommercial License which permits any noncommercial use, distribution, and reproduction in any medium, provided the original author(s) and source are credited.
Previous research on the care-giver burden experienced by adult children has typically focused on the adult child and parent dyad. This study uses information on multiple informal care-givers and examines how characteristics of the informal care-giving network affect the adult child's care-giver burden. In 2007, 602 Dutch care-givers who were assisting their older parents reported on parental and personal characteristics, care activities, experienced burden and characteristics of other informal care-givers. A path model was applied to assess the relative impact of the informal care-giving network characteristics on the care-giver burden. An adult child experienced lower care-giver burden when the informal care-giving network size was larger, when more types of tasks were shared across the network, when care was shared for a longer period, and when the adult child had no disagreements with the other members of the network. Considering that the need for care of older parents is growing, being in an informal care-giving network will be of increasing benefit for adult children involved in long-term care. More caregivers will turn into managers of care, as they increasingly have to organise the sharing of care among informal helpers and cope with disagreements among the members of the network.
Long-term dependencies and high care-giving demands cause distress for many care-givers as they are confronted with role overload and disruptions in regular daily routines (Sales 2003). Given that after spouses, adult children are important sources of care for older people, an adult child care-giver runs a high risk of becoming overburdened. At the same time, adult children seldom provide care on their own (Szinovacz and Davey 2007). They usually share care activities with others, including their spouse, siblings, other kin, friends or neighbours (Ingersoll-Dayton et al. 2003;Szinovacz and Davey 2008;Wolf, Freedman and Soldo 1997). The presence of other informal care-givers suggests that the adult child caregiver is embedded in an informal care-giving network in which that person has to interact regarding care and co-ordinate his or her own care-giving with that provided by others. It is likely that a well-functioning care-giving network reduces the adult child's care-giver burden, but if disruptions in the network or co-ordination problems occur, the network may cause the care-giver additional stress. We will examine whether and how various characteristics of the informal care-giving network affect adult children's care-giver burden.
Care-giver burden has been extensively studied in previous research as one of the negative outcomes of care-giving. The ' stress process model' developed by Pearlin et al. (1990) is one of the most cited theoretical frameworks to explain variations in care-giver burden, stress and wellbeing and its determinants. The model views care-giver burden as the outcome of a process that varies with the characteristics and resources of care-givers and the stressors to which they are exposed. Most studies seek to identify the burden determinants among the care recipient's or the caregiver's characteristics, which emphasises the core dyad in the care-giving. Care-giver burden has been shown to become aggravated by the higher severity of needs in care for the care recipient as well as by the frequency of performance of care-giving tasks or a care-giver's lower mastery and self-esteem (Chappell and Reid 2002;Dwyer, Lee and Jankowski 1994;Sherwood et al. 2005;Yates, Tennstedt and Chang 1999). Furthermore, gender is a factor in determining care-giver burden; female care-givers often experience more burden than male care-givers (Stuckey and Smyth 1997).
Research on adult children's burden often overlooks the facts that multiple informal helpers may be around and that an adult child frequently belongs to a wider care-giving network. We believe there should be more attention to networks in the care-giving burden literature. One facet of the network perspective is reflected in the care-giver stress model. Emotional support provided to a care-giver by friends and family is associated with reduced distress (Miller et al. 2001;Yates, Tennstedt and Chang 1999). The current study extends the examination of the stress process from the network angle, and we expand the stress process model by examining the impact of various aspects of care-giving networks on the adult child's care-giver burden.
Several studies have investigated care-giving networks, and many have considered their gender composition (e.g. Matthews 2002; Matthews and Rosner 1988). Coward and Dwyer (1990), for example, showed that in mixed-gender families, daughters provided more hours of daily care and engaged in more care-giving activities than sons. Others have examined the probability of network members participating in care-giving : Wolf, Freedman and Soldo (1997) reported a negative association between the hours of parent care given by a child and the number of the child's sisters, but few studies have investigated sources of support and interpersonal stress in care-giver's personal networks. Suitor and Pillemer (1993) demonstrated that for daughters caring for parents with dementia, siblings and friends were equally important sources of support, whereas siblings were overwhelmingly the greatest source of inter-personal stress. Generally, these findings imply that a care-giver's network members, either participating in care-giving or not, have the potential to provide instrumental support (such as helping with care-giving tasks) and emotional support, or to be a source of strain for the care-giver. There has been no systematic research that examines how existing care-giving networks influence the care-giver and whether being part of a network affects the care-giver's burden.
Examining care-giving networks and how they affect an adult child's burden prompts us to consider the general notion of social capital and, more particularly, the trusted ties that provide social, emotional and practical support (Gray 2009). Personal networks reflect the availability of persons with whom an individual maintains interpersonal relations and upon whom that individual may rely for support and care. In general, social contact and positive interactions make individuals feel better about themselves and their social world. These interactions provide them a sense of security and a potential support base. People who feel more supported cope better with stress and difficult situations (Antonucci 2001). We assume that the supportive mechanism of personal networks is by analogy applicable to care-giving networks. Being part of a care-giving network, caring for a parent together, and having positive interactions with other care-givers all signify support for a care-giver and may reduce her or his burden. Disruptive interactions with other care-givers probably increase a care-giver's burden. Dependent upon how supportive is the caregiving network, adult children's perceptions regarding care-giving burden may vary.
We distinguish network characteristics that are indicative of support potential for an adult child. First, we expect that an adult child who designates the availability of support and appreciation from other caregivers will experience lower levels of care-giver burden (Hypothesis 1). Second, the size of the care-giving network is likely to be important. The larger it is, the more helpers an adult child can count on and the more emotionally secure an adult child will feel because he or she does not have to respond to care-giving demands alone. Moreover, care can be divided among care-givers, which suggests a lower caring requirement on the child, which in turn might reduce care-giving burden. We hypothesise that as the number of informal care-givers involved with helping an adult child's parent increases, the lower the care-giver burden (Hypothesis 2).
Third, the composition of the care-giving network also seems important. An adult child might feel more secure about the care-giving network when it comprises family members rather than non-kin care-givers. As a result of normative solidarity, family members are less likely to give up the caring role when other responsibilities interfere. As demonstrated by Silverstein, Conroy and Gans (2008), full-time employment or having children in the household does not lessen the time adult siblings provided care to their parents. Friends, for example, may be less likely to take on care responsibilities when they conflict with other roles: working reduces the chance of caring for a friend compared to caring for a family member (Himes and Reidy 2000). Involvement of a family member in the caregiving network probably results in lower care-giver burden compared to involvement of a non-kin care-giver (Hypothesis 3).
Furthermore, the network members' actual contributions to care-giving are critical to whether a network is supportive. As participants in caregiving, network members not only support a care recipient but also the adult child responsible for a parent. Such support can be expressed in two ways: sharing tasks means decreasing the number of care-giving hours (instrumental support), and simultaneously generates an understanding of a responsibility shared -the adult child care-giver is not solely responsible for the care (emotional support). As proposed by a principle of equity and putting other factors aside, multiple care-givers will tend to share care responsibilities equally because unequal involvement is likely to create stress (Ingersoll-Dayton et al. 2003). This suggests that the longer care-givers work together, the more time they have to create, balance or maintain equity in care and to realise the support potential of the care-giving network. Also derived from the principle of equity, network members are likely to agree joint responsibility for each type of task. Sharing these responsibilities should be supportive for the care-giver. We expect that the longer an adult child shares care with other informal care-givers, the lower the burden that an adult child experiences (Hypothesis 4) ; and the more types of care-giving tasks an adult child shares with others, the lower the care-giver burden (Hypothesis 5).
The care-giving network can, however, be disruptive and create or worsen burden. We know that interaction between the adult child and other informal helpers about the provision of care can negatively affect an individual. Earlier research has demonstrated that negative social interactions harm one's wellbeing (Rook 2001). Negative interactions may be present in an informal care-giving network, as when care-givers disagree about the type and the amount of care that should be given and about how care-giving responsibilities should be divided. Strawbridge and Wallhagen (1991) showed that of 100 studied care-givers, 40 per cent had conflict with other family members because they failed to provide sufficient care. Furthermore, conflicts in a network (e.g. because others do provide enough care) can also be associated with providing more hours of care, which aggravates the care-giver's burden. Hypothesis 6 is then that having disagreements with other informal care-givers increases care-giver burden.
To test the hypotheses, we adopted the care-giver stress process model (Pearlin et al. 1990) as modified by Yates, Tennstedt and Chang (1999). Analytically, we developed a specific part of the model and linked three elements (Figure 1). Parental needs in care influence adult child's hours of informal care and adult child's care-giver burden (Chappell and Reid 2002 ;Yates, Tennstedt and Chang 1999). As a new element, we added informal care-giving network characteristics as factors that influence caregiver burden, and examined the effects of six network characteristics on an adult child's care-giver burden. We modelled the association between network characteristics and adult child's care-giver burden both directly and indirectly through hours of care, as shown in Figure 1. Two dependent variables were considered, and we controlled for the adult child's and the parent's characteristics: parent's gender, age, parent's availability of a spousal care-giver, adult child's gender and discretionary time available for care-giving. All these factors are potential predictors of how many hours the adult child provided care to the parent (Barrett and Lynch 1999 ;Cicirelli 1983).
The data were collected in two steps for the study 'Informal Care' by Statistics Netherlands and The Netherlands Institute for Social Research in 2007. At the first step, informal care-givers were identified with four screening questions included in the Statistics Netherlands' Labour Force Survey of 2007. A representative sample of 54,451 Dutch adults drawn from different areas (i.e. rural, urban and mixed), aged 18 or more years and living in a household were asked whether they had provided care for two weeks or longer during the last 12 months for a family member who was severely ill or needed assistance because of an illness, accident, hospital admission or other reasons. Of the identified 4,484 care-givers, 2,813 participated in the follow-up written questionnaire on informal care-giving. To adjust for selective non-response, the remaining sample was weighted for a number of characteristics (i.e. gender, age, marital status and level of urbanisation of the residential area). The respondents selfcompleted the information on their own characteristics and on the characteristics of their care recipients, including needs for help and various aspects of care-giving.
For the current analysis, we examined the data for 1,112 respondents who helped their older parents (including parents-in-law) who were aged from 55 to 103 years. We excluded 207 parents living in institutions, the 25 respondents who lived with the care recipient, and 10 people who provided no information about their care-giving activities. Five respondents who accomplished two or more tasks less than the other members of their informal network were also excluded. Because we focused on the informal care-giving network, the analyses relied on the cases in which the respondent identified other informal care-givers. A preliminary analysis demonstrated that respondents without an informal care-giving network did not experience a higher care-giver burden, and in these cases the care recipients had significantly less mental and physical impairment. The final analysis sample comprised 602 care-givers with informal care-giving networks (479 women and 123 men aged 21-78 years).
The main dependent variable, experienced care-giver burden, was measured using an extended version of the Self-Perceived Pressure from Informal Care Scale (Pot, Van Dyck and Deeg 1995;Timmermans et al. 2001). The scale takes into account both low and high (or intense) pressure, and measures only subjectively experienced pressure. Psychical complaints or stressors, such as the amount of help provided, are excluded. This scale was modified from the burden scale of nine items suggested by Zarit, Reever and Bach-Peterson (1980) with five additional items. The respondents were asked whether they agreed with 14 statements on perceived time and emotional pressure, such as: ' Generally speaking I felt very pressured because of the situation of my care receiver' ; ' My independence suffered' ; 'I was too tired to do anything in my free time in the period that I was providing help'. The answers were coded as dichotomies (' 0 ' for any level of disagreement, ' 1' for any level of agreement). The sum scale scores for the 14 items were computed and varied from 0 (not burdened) to 14 (highly burdened). The hierarchical order of the burden items was tested with the Mokken scale analysis (H-value 0.47) and the scale was moderately homogeneous (Molenaar and Sijtsma 2000). As the additive scale was skewed, we considered transformation by using an alternative coding of the dependent variable. Preliminary regression analysis did not reveal different results and we kept the original scale in the analysis.
The second dependent variable in the model was the average number of hours of informal care given per week during the 12 months prior to the interview. The question asked was, ' How many hours did you give care on average per week when the need for care was highest? ' Respondents who mentioned more than 112 hours per week were recoded as giving 112 hours per week, as that is the maximum possible number of hours per week allowing for eight hours of sleep per day.
A measure of the parent's cognitive impairment was obtained from the adult child's assessment of whether the parent had (early stage) dementia or other mental problems (0 or 1), and whether care was required in connection with other psychiatric problems (0 or 1). Physical limitations of the care recipient were measured on the basis of 13 items of functional limitations to perform activities of daily living (ADL), such as being able to dress and bathe, use the restroom without assistance, walk up and down stairs, do household chores and shop for groceries. The scale of physical limitations was based on the Katz et al. (1970) ADL scale but with household activities and mobility items added. The response codes were ' 1 ' ' can do without difficulty', ' 2 ' 'can do with difficulty', '3 ' 'can do only with help/no, or unable to do because of health conditions'. Mokken scale analysis was performed to test the homogeneity of the scale (H-value 0.66). The scores for all answers were aggregated and the scores ranged from 13 to 39. Last, the adult child reported whether the parent could be left alone longer than half-an-hour (0 or 1).
To measure the designation of support and appreciation of other informal care-givers, the respondents were asked whether they agreed with the following statement : 'I receive a lot of support and appreciation from other care-givers'. The answers varied from ' 1 ' ' fully agree' to ' 4' ' fully disagree ' and recoded into ' 1 ' 'agree or fully agree ' and ' 0 ' 'disagree or fully disagree '. The respondents reported on the number of other care-givers giving help to a care recipient (range 1-9). If there were more than three, the respondents were asked to identify no more than three that provided most care (excluding themselves). Most (78%) respondents did not identify more than three care-givers. For the three, information was collected regarding the relationship with the respondent (i.e. adult child's partner, adult child's own child, adult child's parent or parent-in-law, adult child's sibling or sibling-in-law, other family member, a friend of care-recipient or adult child's friend). These relationships were grouped into four categories : adult child's parent, sibling (including sibling-in-law), other kin and non-kin care-giver. We created four dummy variables indicating the presence of each of these four relationships in the network (0 or 1).
Adult children provided information about the duration of care provision (in months) during the past 12 months as well as the duration of care provision (in months) by each of the other informal care-givers. We calculated the number of months that a respondent gave care together with at least one other informal care-giver. Furthermore, the respondents provided information about whether they provided six types of care: household tasks, personal care, nursing care, emotional support, administrative help and helping with visits (yes/no). They also reported whether each of the network members performed those six task types. We calculated the proportion of task types respondents shared with at least one other care-giver of the total number of task types performed by the respondent (ranging from 0 to 1).
The respondents were asked how often they experienced disagreements with other care-givers in the informal network of the following kinds : the type of care that should be given, how often care should be given, the division of the care-giving tasks, and placing an older adult in an institution. The response categories were 'seldom to never', ' regularly ' and ' often'. Because the answers had a skewed distribution, we created a variable that indicated whether respondents had disagreements (regularly or often) with other care-givers on at least one item (0 or 1).
We used the information on care-giver and care recipient characteristics as control variables. Respondents reported on their age (age difference between adult child and parent was used in the analysis to avoid multi-colinearity with the parent's age), gender (0 male, 1 female), partner status (0 no partner, 1 has a partner), the number of children in the household (range 0-6), paid work in the last 12 months (0 no work, 1 having paid work), travel time to parent's place of residence in minutes (range 0-240) and the number of task types an adult child accomplished (range 0-6). Care-givers also reported on the characteristics of care recipients, such as age, gender (0 male, 1 female) and partner status (0 no partner, 1 has a partner).
Parameters for the associations between parental needs in care, hours of informal care provided by an adult child, adult child's care-giver burden, characteristics of the care-giving network and control variables (i.e. adult child and parental characteristics) were estimated in a path model using structural equation modelling with AMOS (Kline 1998). The hours of informal care and adult child's care-giver burden were modelled as dependent variables. The independent variables were parent's cognitive and physical impairments, six network characteristics (support and appreciation from other care-givers, the number of other care-givers, type of relationship with other care-givers, the duration of sharing care, the proportion of task types shared with others and disagreements among care-givers) and control variables (parent and adult child characteristics). We estimated the final trimmed model after eliminating all insignificant covariance among the independent variables (Kline 1998). To assess the fit between the model and the data we calculated the Comparative Fit Index (CFI) and the Root Mean-Squared Error of Approximation (RMSEA). CFI values greater than 0.95 are considered acceptable. A RMSEA value less than 0.05 is acceptable (Hu and Bentler 1999). The six hypotheses were tested by regressing the adult child's care-giver burden on six caregiving network variables. Hours of an adult child's care-giving were also regressed on six network variables, on parental care needs, and on the control variables (Figure 1).
The average burden score was 4.3 (standard deviation (SD) 3.7), and scores ranged from 0 to 14. About 17 per cent of the care-givers were not burdened at all and about 8 per cent were heavily burdened, meaning they scored at least 10 on the scale. As shown in Table 1, the care-givers' ages varied greatly with a mean of about 48 years. The sample of care-givers was mostly female (80 %). Most (73%) had partners and carried out paid work (71 %). They had, on average, about one child in the household and needed on average about half-an-hour to get to the parent's residence, and they performed on average almost four tasks out of six. The care recipients were aged between 55 and 103 years but the majority were elderly (mean 79.6 years, SD 9.0). The sample of care recipients was composed mostly of mothers (or mothers-in-law); about one-third of the care recipients had partners. Almost one-third of care recipients had (early stage) dementia, and 9 per cent had psychiatric problems ; 14 per cent of parents could not be left alone longer than half-an-hour. The estimates of parental physical limitations were relatively high (mean 30.8, SD 6.4, range 13-39).
Most (89%) of the adult children reported that they perceived support and appreciation from other care-givers. The average size of the
T A B L E 1. Characteristics of the sample of adult care-givers with informal care-giver networks, The Netherlands, 2007 Variables and categories % Mean SD Range Dependent variables: Hours of informal care per week 15.39 18.69 1-112 Caregiver burden 4.29 3.67 0-14 Care-giver attributes : Age 48.62 0.38 21-78 Gender (female) 80 0 or 1 Have a partner (yes) 73 0 or 1 Number of own children in household 0.90 0.04 0-6 Paid work in the last 12 months (yes) 71 0 or 1 Travel time to parent's residence (in minutes) 28.44 35.80 0-240 Number of task types 3.89 0.05 0-6 Care recipient attributes: Age (years) 79.61 9.02 55-103 Gender (female) 70 0 or 1 Have a partner (yes) 33 0 or 1 Needs in care: Having dementia (yes) 28 0 or 1 Having psychiatric problems (yes) 9 0 or 1 Physical limitations 30.82 6.39 13-39 Cannot be left alone 14 0 or 1 Network characteristics: Support and appreciation from network members (yes) 89 0 or 1 Care-giving network size 2.76 1.92 1-9 Having a parent within a network (yes) 13 0 or 1 Having a sibling within a network (yes) 75 0 or 1 Having other family within a network (yes) 30 0 or 1 Having non-kin within a network (yes) 21 0 or 1 Duration of shared care provision (in months) 7.77 4.16 1-12 Proportion of shared task types 0.74 0.29 0-1 Disagreement between child care-giver and network members (yes) 22 0 or 1 Sample size 602 Note : SD : standard deviation.
care-giving network was 2.8 and ranged from one to nine. Of all adult children, 13 per cent mentioned a parent among the three other caregivers, 75 per cent mentioned a sibling (including sibling-in-law), 30 per cent mentioned another immediate family member, such as partner, own child or others, and 21 per cent mentioned a non-kin care-giver. Adult children shared care with at least one other care-giver for on average 7.8 months out of 12. Most of the tasks that an adult child carried out were shared with at least one other care-giver. The proportion of shared tasks was on average 0.74 (SD 0.3). About one-fifth (22%) of the adult children reported disagreements with other care-givers.
The fit statistics of the path model were both acceptable (CFI 0.96 and RMSEA 0.03). The estimated unstandardised model parameters derived from the path model are presented in Table 2. The results show that the adult children who gave more hours of informal care were older, did not have partners, performed a large number of tasks, and had parents that could not be alone for more than half-an-hour. Regarding network characteristics, only the number of months over which an adult child shared care with others significantly affected adult children's care-giving hours ; in other words, an adult child provided less hours of care per week if the care was shared for a longer period. In total, the model explained 21.4 per cent of the variance in adult children's care-giving hours.
More cognitive and physical impairments, as well as more hours of informal care, were positively correlated with an adult child's care-giver burden. From an additional bivariate analysis (not detailed here), we know that there was a significant negative correlation between a child care-giver reporting appreciation and support by other care-givers and her or his perceived burden. Despite the significant correlation, perceiving appreciation and support from other care-givers did not influence care-giver burden in the path model, which does not support Hypothesis 1. As postulated in Hypothesis 2, a higher number of care-givers associated with lower care-giver burden although the effect was relatively small given that burden ranged from 0 to 14 (B=x0.24, p<0.05). Contrary to Hypothesis 3, however, the type of relationship an adult child had with other caregivers had no significant effect on her or his care-giver burden. The duration of shared care provision did not directly affect adult child care-giver burden (Hypothesis 4), but sharing a larger number of tasks with others did (Hypothesis 5) : the higher the proportion of the adult child's care tasks that were shared with others, the lower the burden (B=x1.15, p<0.05). The duration of the shared care provision influenced the number of caregiving hours (B=x0.51, p<0.001), and there was a significant correlation between hours of care and care-giver burden (B=0.05, p<0.001). The results suggest an indirect effect (b=x0.03) : the longer that others shared in the responsibility for care, the lower an adult child's care-giving burden. As stated in Hypothesis 6, having disagreements with other care-givers increased an adult child's care-giver burden by 1.37 ( p<0.001). When comparing all network characteristics, having disagreements with other care-givers had the largest effect on adult child's care-giver burden; it had the highest standardised regression weight coefficient (b=0.16)
T A B L E 2. Regression of hours of informal care and child's care-giver burden on care-giver's and care-recipient's characteristics and the characteristics of the care-giving network Hours of informal care (b 1 ) Child's care-giver burden Direct effects (b 1 ) Indirect effects (b ) Care-giver: Hours of informal care per week -0.05*** -Age difference x0.32* -x0.03 Gender (female) 2.58 -0.01 Have a partner (yes) x4.39* -x0.03 Number of own children in household 0.24 -0.01 Paid work in the last 12 months (yes) x1.07 -x0.01 Travel time to parent's residence 0.01 -0.01 Number of types of tasks 4.59*** -0.08 Care recipient: Age 0.01 -0.00 Gender (female) x2.48 -x0.02 Have a partner (yes) x3.28 -x0.02 Care recipient needs: Having dementia x0.85 0.81** x0.01 Having psychiatric problems x2.56 2.01*** x0.01 Physical limitations 0.16 0.06* 0.01 Cannot be left alone 9.75*** 0.71 0.05 Network characteristics: Support and appreciation from network members 2.68 x0.55 0.01 Care-giving network size x0.05 x0.24** 0.00 Having a parent within a network x2.10 0.26 x0.01 Having a sibling within a network x0.38 x0.22 x0.01 Having other family within a network 1.29 x0.15 0.01 Having non-kin within a network 1.43 0.60 0.01 Duration of shared care provision x0.51*** 0.02 x0.03 Proportion of shared task types x2.99 x1.15* x0.01 Disagreement between child care-giver and network members x0.87 1.37*** x0.01 R 2 21.4% 20.0% Notes: 1. The figures are unstandardised regression coefficients and represent direct effects. b: standardised regression coefficient. The sample size was 602. The correlation matrix of all variables is available upon request. Significance levels : * p<0.05, ** p<0.01, *** p<0.001.
(not detailed in Table 2). The model explained about 20 per cent of variance in care-giver burden.
This study has focused on the impact of the informal care-giving network on an adult child's care-giver burden. Whereas much previous research has focused on the characteristics of the parent-child dyad to account for the adult child's care-giver burden, it has been shown that attributes of the informal care-giving network are important. The findings not only corroborate the care-giver stress process model's prediction that the parent's impairments influence the adult child's care-giver burden, as previous investigators have shown (Chappell and Reid 2002;Dwyer, Lee and Jankowski 1994;Yates, Tennstedt and Chang 1999), they also add new knowledge. It has been shown that the informal care-giving network plays an essential role for an adult child care-giver, as his or her care-giver burden partly depends on how supportive it is. The idea of social capital, defined as the array of ties giving access to various forms of support (Gray 2009), proved important in care-giving situations. We demonstrated that an adult child experiences lower levels of care-giver burden when he or she can count on a larger care-giving network, shares tasks with others for a longer period, and shares more types of tasks with others. At the same time, the findings also suggest that the informal care-giving network can increase care-giver burden if there are disagreements among the network members.
The model explored to what degree the impact of the informal caregiving network on burden is influenced by the number of hours of care provided. It seems plausible that when more care-givers are involved and more tasks are shared, an adult child will provide fewer hours of care and, consequently, experience lower care-giver burden. Spitze and Logan (1990) studied sibling care-givers and suggested that the greater the number of siblings, the greater the amount of sharing. The Netherlands data, in contrast, revealed that network size did not affect an adult child's caregiving hours, but nonetheless that both care-giving network size and the number of shared task types directly affected the adult child's burden, consistent with Hypotheses 2 and 5. Wallsten et al. (1999: 145) stated that ' one's perception of the network's helpfulness appears to be more potent than the actual help provided by friends and family'. This suggests that a larger care-giving network and sharing more tasks with others might in themselves be sufficient for an adult child to feel supported and to experience lower care-giver burden, even if she or he provides the same amount of care. There was one indication that the care-giving network affected burden by decreasing the hours of care: the longer an adult child had shared care with others, the fewer the hours of care she or he provided and the lower the experienced care-giver burden. It seems that it takes some time before care-givers come to an agreement on how to divide care among them, or in other words for the predictions of equity theory to take effect (Walster, Walster and Berscheid 1978). This suggests that the informal care-giving network decreases an adult child's care-giver burden, either directly by providing the emotional support of sharing the care with others, or, in the case of extended durations of care, by enabling the adult child eventually to provider fewer hours of care.
Unexpectedly the results did not support the hypothesis that an adult child experienced lower burden when there were family members in the informal care-giving network. It might be that the care-giving network is already highly selected: the adult child had to choose three other caregivers who provided the most care. Given such a selection, one can imagine that any nominated non-kin are people with whom close and supportive relationships exist. Their presence may then be comparable to the presence of siblings or other relatives. The results suggest that it does not matter who is involved in the network of informal helpers, as long as there are multiple helpers. It may be that single adult-child care-givers consider the assistance of non-kin just as valuable as assistance from the family, whereas adult children with siblings may consider their involvement intrinsically important. The data did not have information about the availability of siblings or family composition, so we could not distinguish between adult-child care-givers who do and do not have siblings.
The results support Hypothesis 6 that an adult child experiencing disagreements with other care-givers has higher care-giver burden, which corroborates earlier research showing that a family conflict affects caregiver strain (Scharlach, Li and Dalvi 2006). The bivariate analysis indicated that support and appreciation from others decreased care-giver burden. When disagreement was added to the total model, the effect of perceived network support disappeared. The impact of negative interactions on burden seems larger than the impact of positive interactions. These findings are in line with those of previous studies on the impact of negative interaction on one's psychological wellbeing: negative exchanges occur less often but the consequences exceed those of positive exchanges (e.g. Newsom et al. 2005;Rook 2001). Our results suggest that the negative influence of disagreements in a care-giving network exceeded the positive influence of feeling supported and appreciated by others.
Several comments should be made about the study's methodological limitations. Some measures used in the questionnaire could have been more precise, e.g. parents and parents-in-law were not distinguished. It is possible that caring for a parent-in-law involves a lower contribution than caring for one's own parent, so the care-giver burden might be slightly under-estimated for the latter. Furthermore, we did not obtain much information about interactions within the care-giving network, and are not sure whether communications about care-giving and sharing tasks are through the parent or directly among the members of the care-giving network, or both. We inferred that there is communication across the network and that an adult child is fully aware of all available support. Studying care-giving network members' inter-communication in more detail would make a valuable contribution to the care-giving literature. Another limitation is that our cross-sectional study had no depth information about the processes by which care-giving is shared. Disagreements among care-givers could be a determinant of care-giver burden, but could also be a result of care-giver burden and indeed of sharing tasks. A longitudinal design would clarify how care-giving networks and the process behind sharing care affect an adult child's burden. Finally, many care-giving networks involve a larger group of care-givers with both informal and formal help, which may multiply the potential sources of both support and disagreement (Carpentier and Ducharme 2003). We did not take into account formal care in the current research, but almost all the care recipients received care from professional care-givers.
The study has implications for care-givers, professional helpers and policy makers. First, adult children who provide long-term care to a frail older parent clearly benefit from the availability of an informal care-giving network. This implies that, along with the provision of care, adult child care-givers have to spend time organising the informal care-giving network, co-ordinating care activities, and coping with disagreements among the informal helpers. For primary care-givers, particularly the daughters of single parents, this may imply a mental shift from actually performing care activities to organising an informal structure in which care activities are more equitably distributed among multiple helpers. Given the longterm increase in labour-force participation among women, those who might be future parental care-givers should realise that parent-care is better and more sustainable over time when shared with other informal and formal helpers. Secondly, professional care-givers are accustomed to dealing with one primary care-giver ; usually a spouse or one of their adult children, and preventing care-giver burden is one of their professional activities. Their support to care-givers needs to be extended to the informal care-giving network. This could involve helping to organise the network, co-ordinating its tasks, and intervening when disagreements occur. Finally, policy makers need to realise that the informal care-giving network comprises both kin and non-kin. Increasing numbers of friends and neighbours assist frail older adults who are single or whose adult children do not live close by. Support programmes, financial arrangements or arrangements for work-leave to provide care are generally targeted or restricted to the relatives of people in need of care. An extension of these programmes to non-kin care-givers would acknowledge that informal care for frail older adults has a more variable network perspective.
Natalia Tolkacheva et al.
Acknowledgements This study was funded by the
This annual review focuses on invertebrate model organisms, which continue to yield fundamental new insights into mechanisms of aging. This year, the budding yeast has been used to understand how asymmetrical partitioning of cellular constituents at cell division can produce a rejuvenated offspring from an aging parent. Blocking of sensation of carbon dioxide is shown to extend fly lifespan and to mediate the lifespan-shortening effect of sensory exposure to fermenting yeast. A new study of daf-16, the key forkhead transcription factor that mediates extension of lifespan by mutants in the insulin-signalling pathway in Caenorhabditis elegans, demonstrates that expression of tissue-specific isoforms with different patterns of response to upstream signalling mediates the highly pleiotropic effects of the pathway on lifespan and other traits. A new approach to manipulating mitochondrial activity in Drosophila, by introducing the yeast NADH-ubiquinone oxidoreductase, shows promise for understanding the role of mitochondrial reactive oxygen species in aging. An exciting new study of yeast and mammalian cells implicates deterioration of the nuclear pore, and consequent leakage of cytoplasmic components into the nucleus, as an important cause of aging in postmitotic tissues. Loss of, or damage to, chromosome-associated histones is also implicated in the determination of lifespan in yeast, worms and fruit flies. The relationship between functional aging, susceptibility to aging-related disease and lifespan itself are explored in two studies in C. elegans, the first examining the role of dietary restriction and reduced insulin signalling in cognitive decline and the second profiling aggregation of the proteome during aging. The invertebrates continue to be a power house of discovery for future work in mammals.
One of the most fascinating observations in biology is the production of youthful offspring by aging parents. The mechanisms by which this phenotypic rejuvenation can occur are being investigated in the single-celled, budding yeast Saccharomyces cerevisiae. This organism has an asymmetrical mitotic division, in which the larger mother cell shows progressive deterioration in the capacity to bud while the smaller daughter cells are born with full replicative potential. At least part of this asymmetry is attributable to a mechanism of spatial quality control (SQC), in which accumulated extrachromosomal rDNA circles (Sinclair & Guarente, 1997) and oxidatively damaged proteins (Erjavec et al., 2007) are selectively retained in the mother cell, a process that requires both the sirtuin Sir2p, a protein deacetylase, and the protein-aggregate remodelling factor Hsp104p (Erjavec et al., 2007). A recent study (Liu et al., 2010) has investigated mechanisms involved in SQC, using synthetic genetic array analysis, which screens for interactions between individually viable mutations that result in synthetic lethality. Such an interaction implies that the two gene products may affect the same process, in this case SQC. A screen for interactions with a Sir2p null mutation identified actin and the polarisome, a protein complex that is involved in the organization of the actin cytoskeleton and consequently required for polarized cell growth. These interactions required the deacetylase activity of Sir2p, and cells lacking Sir2p showed elevated acetylation of the chaperonin CCT, reduced efficiency of actin folding by CCT and elevated levels of non-native actin. Mutants in polarisome components and the myosin V motor protein resulted in reduced efficiency of segregation of protein aggregates and lowered replicative lifespan. Cytoskeletal functions and polarity are thus crucial for the asymmetrical segregation of damaged molecules. A second recent study (Eldakak et al., 2010) has identified a new class of proteins that are retained in the yeast mother cell and that appear to be limiting for replicative lifespan, the MDR transporter proteins (so named because they mediate multidrug resistance in some cancers). These proteins are found in the plasma membrane and eject a variety of toxins from cells. At budding, the mother cell retains its own pool of MDR proteins, while the daughter cell inherits newly synthesized MDRs. Mild over-expression of several of these MDRs increased replicative lifespan, implying that their number or function may decline during aging and limit lifespan.
One unresolved paradox is why loss of asymmetric segregation of damaging molecules in budding yeast leads to shortened replicative lifespan of the bud. Normally, the daughter cell is generated in a youthful state, presumably because the mother retains damaging agents during division. Following this logic, if asymmetry is lost in the absence of altered rates of damage repair, the daughter would inherit more damage at birth but would also retain less damage during subsequent divisions as a mother cell. The observation that lifespan is reduced in cells losing asymmetry may imply that damage inherited in the daughter is more highly detrimental than that accumulating in the mother, perhaps because mothers are more equipped to deal with damage or because impaired asymmetry also leads to impaired repair. These issues remain to be resolved. It will also be important to establish whether these mechanisms of SQC through retention of damaged molecules in the parent by cytoskeletal and polarity function are important in the germ lines and stem cells of multicellular organisms and, if so, to identify the classes of molecules that are affected.
Dietary restriction (DR) can extend lifespan in numerous organisms and can also ameliorate aging-related loss of function and pathology (Mair & Dillin, 2008;Fontana et al., 2010). Dietary restriction in the laboratory model organisms can be implemented using a variety of methods, and it is clear that the mechanisms of lifespan extension can vary with the method used. Although altered nutrition per se presumably mediates the extension of lifespan in some cases, it is also becoming apparent that chemosensation, without any change in nutrient intake, can play an important role. For instance, fruit flies subjected to DR show increased lifespan, but if they are exposed to odours derived from live yeast, a standard food source for Drosophila in the laboratory and in nature, then some of the extension of lifespan by DR is lost (Libert et al., 2007). Poon et al. (2010) now show that sensation of carbon dioxide plays a key role in the response to live yeast. Flies sense the gas using a specific subpopulation of neurons that express two gustatory receptor genes. A null mutation in one of these, Gr63a, abolishes the electrophysiological response of the neurons to carbon dioxide, and it also extended the lifespan of female, but not male, flies by 30%. Lifespan was extended mainly by reducing baseline mortality rate rather than reducing the rate of increase with age, similar to the effects of DR in Drosophila. The extension of lifespan was lost if Gr63a was re-expressed in the neurons using a heterologous expression system, and targeted ablation of the carbon-dioxide-sensing neurons also resulted in an increase in lifespan, demonstrating robust causality. Whereas exposure to odorants from live yeast generally shortened the lifespan of female (but not male) control flies, the mutant Gr63a flies did not respond, although flies mutant for a different olfactory mutant, Orb83b, which are generally anosmic but retain the ability to sense carbon dioxide, responded normally. Interestingly, Gr63a null flies themselves responded normally to DR imposed by dilution of a diet containing killed yeast and sucrose (Poon et al., 2010), a standard method of DR in Drosophila in many laboratories, which robustly increases lifespan in multiple strains of flies (Bass et al., 2007;Grandison et al., 2009;Wong et al., 2009;Piper et al., 2010). These results imply that chemosensation of the carbon dioxide from fermentation by live yeast plays an important role in the response of lifespan to this food source, but that other systems, possibly also involving chemosensation, mediate the response of lifespan to killed yeast and ⁄ or sucrose. Chemosensation of food can also shorten lifespan in worms subjected to DR (Smith et al., 2008), and it would be interesting to know if olfactory or gustatory mutant mice have longer lifespans or if their lifespans respond differently to DR.
The insulin ⁄ Igf ⁄ TOR signalling network plays an evolutionarily conserved role in the determination of lifespan in budding yeast, the nematode worm Caenorhabditis elegans, the fruit fly Drosophila melanogaster, the mouse and, possibly, humans (Fontana et al., 2010;Kapahi et al., 2010;Kenyon, 2010). In the worm, extension of lifespan by reduced activity of the upstream pathway requires the forkhead transcription factor daf-16 (Kenyon et al., 1993). Much interest has therefore centred on the roles of orthologous forkhead box O transcription factors in other organisms, including humans (Greer & Brunet, 2008;Partridge & Bruning, 2008;Greer et al., 2009), and in the mechanisms by which altered activity of daf-16 increases lifespan in the worm. It now transpires that expression of two different, tissue-specific splice variants of daf-16 are required for the full extension of lifespan (Kwon et al., 2010). The worm has three daf-16 isoforms, with two (a and b) discovered some time ago and a third (d ⁄ f) reported more recently. Some evidence suggested that one of these isoforms, daf-16a, was particularly important for the increased longevity of insulin pathway mutants, which is lost in a daf-16 null background. However, experiments where daf-16a alone was re-introduced to worms double mutant for the insulin receptor daf-2 and for daf-16 resulted in incomplete rescue of the increased lifespan, leading Kwon et al. to take a closer look at the situation. New isoformspecific constructs for RNA interference were used to show that the new daf-16 isoform, daf-16d ⁄ f, as well as daf-16a, was involved in the extension of lifespan by a mutant insulin receptor. This conclusion was further supported by rescue experiments for the extended lifespan of an insulin receptor mutant in a daf-16 null background, where both daf-16a and daf-16d ⁄ f were required for full restoration of the lifespan extension, with daf-16d ⁄ f playing the preponderant role. These two daf-16 isoforms have different patterns of tissue specificity, with daf-16d ⁄ f particularly enriched in the intestine, a tissue already known to be important for the action of daf-16 in the extension of lifespan by reduced insulin signalling (Libina et al., 2003). Rescue experiments with the promoter regions of the daf-16 isoforms swapped indicated that the promoter region, and hence the different patterns of tissue-specific expression, mediated the differing effects on lifespan. The two isoforms also responded differently to altered activity of the two upstream AKT kinases, with daf-16d ⁄ f more strongly regulated by AKT-1. The differential inputs and tissue-specific expression of these daf-16 isoforms may explain the highly pleiotropic effects of this pathway, by allowing specific and spatially differentiated responses to different inputs to insulin signalling, such as nutrition and stress. Worms and flies have only a single gene encoding a forkhead box 0 transcription factor, while mammals have several, which may show similar patterns of differentiation in function to the daf-16 isoforms.
Aging is accompanied by the accumulation of damage to molecules, cells, tissues and the whole system. At least some of this damage is thought to be causal for loss of function during aging, although an intriguing alternative perspective has been suggested, based upon deleterious effects of TOR activity later in life (Blagosklonny, 2010). Oxygen free radicals (Harman, 1956) have occupied central stage in discussions of damage-induced aging, although the oxidative damage theory has been undermined by recent evidence, as described in last year's Hot Topics (Partridge, 2009). Two studies with Drosophila have taken a novel approach, with interesting, but somewhat conflicting, results. Complex 1 of the mitochondrial electron transfer chain, NADHubiquinone oxidoreductase, is strongly implicated in the production of superoxide in mitochondria isolated from Drosophila and rodents. Direct manipulation of the activities of this complex is hampered by the fact that it consists of over 40 subunits that are encoded by both the nuclear and the mitochondrial genomes. However, the yeast NADH-ubiquinone oxidoreductase Ndi1 is composed of a single, nuclear-encoded polypeptide. Two studies (Bahadorani et al., 2010;Sanz et al., 2010) have examined the effects of introducing this yeast enzyme into Drosophila mitochondria in vivo. When ubiquitously expressed, enzyme activity, ATP levels and NAD + ⁄ NADH ratio were all increased. However, one study (Sanz et al., 2010) found reduced superoxide production from isolated mitochondria while the other (Bahadorani et al., 2010) found no change, possibly because different respiratory substrates and methods of measurement were used in the two studies. Ubiquitous expression extended lifespan in one study (Sanz et al., 2010), but only neuronal expression did so in the second (Bahadorani et al., 2010). These somewhat inconsistent results could have been a consequence of differences in expression level, and there are now fly stocks that allow standardization of insertion site and expression level of transgenes (Markstein et al., 2008). This approach to manipulating mitochondrial activity is a promising one for understanding effects of respiratory activity and oxygen free radicals on aging.
Although molecular damage is a marked feature of the aging process, damage at higher levels of organization may be equally important in causing aging. A fascinating new study (D' Angelo et al., 2009) has revealed a new type of damage that occurs at the cellular level and that could play a key role in many postmitotic tissues. Nuclear pores are formed of complexes of nucleoporin proteins, which in dividing cells disassemble at mitosis and reassemble into the newly forming nuclei. D'Angelo et al. showed in C. elegans that expression of the genes encoding a subset of scaffold nucleoporins is confined to dividing cells, with scaffold proteins produced in the embryo showing life-long sta-bility in the postmitotic cells of the adult worm. RNA interference did not reduce the level of these scaffold nucleoporins in adult somatic cells, nor did it reduce lifespan, even of long-lived insulin receptor mutant worms. A similar down-regulation of expression of scaffold nucleoporins at exit from the cell cycle, and long-term stability of the proteins, was found in mouse cells. Furthermore, these scaffold proteins showed very low turnover compared with other components of the nuclear pore and of the nuclear envelope, such as lamins. D'Angleo et al. therefore assessed nuclear pore function during aging, by measuring permeability to fluorescently labelled dextrans of a molecular weight that would normally be excluded. Nuclei isolated from old, but not young, worms and rat brains allowed passage of these molecules, which was prevented by chemical blocking of the nuclear pores. Nuclei from rat cells that failed to exclude dextrans also showed leakage into the nucleus of normally strictly cytoplasmic proteins such as tubulin, which aggregated into large filamentous structures that caused severe chromatin aberrations, and one of the scaffold nucleoporins, Nup93, was lost from these nuclei. It would be interesting to know whether interventions that extend lifespan, such as DR and reduced activity of the insulin ⁄ Igf ⁄ TOR network, tend to protect scaffold nucleoporins and nuclear pore function during aging.
It is also noteworthy that mutations in LMNA cause Hutchinson-Gilford progeria syndrome (HGPS), a severe progeroid disorder characterized by pathologies resembling premature aging (Burtner & Kennedy, 2010). While it is unclear that the molecular causes of HGPS overlap with those of normal aging, it has been reported that A-type nuclear lamins, encoded by LMNA and juxtaposed to the nuclear envelope, may regulate nuclear pore function. Moreover, an altered form of lamin A resembling the HGPS mutant is detectable in normal cells and tissues and may increase with age, while defects in nuclear organization increase with age. More studies need to be conducted to determine whether the LMNA mutants linked to progeria affect the function of scaffold nucleoporins.
Another potentially widespread type of damage in the cell nucleus during aging has been found in yeast, through loss of chromatin-associated histones (Feser et al., 2010;McCormick & Kennedy, 2010). Levels of histone proteins including histone H3 decline with age, and chromatin immunoprecipitation showed less H3 associated with specific gene promoters, including promoters that are normally both silent and active in aging yeast cells. Tellingly, over-expression of histones led to increased replicative lifespan. This result could imply that looser chromatin packing, and hence elevated expression of many genes, contributes to aging or that histones become damaged during aging and a higher level of histone proteins allows for more frequent replacement. Loss of chromatin structure and de-repression of gene expression could contribute to aging in multicellular organisms, where transcriptional dysregulation during aging has been reported. However, recent findings in the worm and the fly appear at first sight mutually contradictory. A recent report on C. elegans (Greer et al., 2010) showed that reduced H3K4 trimethylation, which is usually characteristic of inactive chromatin, increased lifespan of the worm. However, another recent study, in Drosophila, showed extended lifespan by reduced chromatin silencing, through lowered trimethylation of lysine 27 in histone H3 by the Polycomb Repressive Complex 2 (Siebold et al., 2010). These results could imply that the effects of chromatin silencing on lifespan are opposite in worm and fly. However, the fly study used heterozygous mutations, which can be a problem in Drosophila because of the possibility of heterosis. Nonetheless, it will be important to understand the exact role of chromatin silencing in aging in these two invertebrates.
Aging is the major risk factor for the predominant killer diseases in humans. Animal models of increased lifespan often show evidence of improvement in function and reduction in impact of aging-related diseases, implying that slowing aging also ameliorates the diseases of aging. However, the correlation between increased lifespan and specific improvements in function is not perfect (Partridge, 2010), and it is clear that some models of increased lifespan are not associated with universally improved function during aging (Bhandari et al., 2007) and can even have negative side effects, such as reduced resistance to infection (Clinthorne et al., 2010). It is therefore important to understand the mechanisms linking improved function during aging, resistance to aging-related disease and lifespan. Two recent studies in C. elegans are illuminating. In the first, Kauffman et al. (2010) developed useful new paradigms to assess associative learning and memory in the worm. During aging, learning ability and long-term memory declined. Reduced insulin signalling and DR both alleviated cognitive decline, but in subtly different ways. A mutation in the insulin receptor daf-2 improved memory in young adults and maintained the ability to learn during aging, but did not extend long-term memory with age. In contrast, eat-2 mutants, which reduce pharyngeal pumping and are a model of DR in C. elegans, impaired longterm memory in young adults but maintained the memory for longer into older age. The C. elegans homologue of the mammalian transcription factor CREB was required for long-term memory. CREB expression declined with age and was higher in young insulin receptor mutant worms and lower in young eat-2 worms and was maintained for longer with age in eat-2 mutants. The differing effects of these two mutants on longterm memory thus corresponded with their differing effects on expression of CREB. These findings point to the value of understanding the whole signalling network, to identify the modulations that are likely to provide the maximum benefit to specific aspects of aging-related decline.
Also searching for valid generalizations about aging as a risk factor for disease, David et al. (2010) investigated protein aggregation as part of the normal aging process in C. elegans. Aggregated proteins are present in many neurodegenerative and amyloid diseases, but their possible role in the normal aging process has barely been investigated. David et al. found that a significant fraction of the proteins in young worms were insoluble in a strong detergent buffer, and that this fraction increased over threefold with age. Some proteins showed no change in insolubility with age and over-represented among these were cytoskeletal proteins. A subset of 461 proteins showed evidence of a repeatable increase in insolubility with age, which was not associated with an increase in total levels of individual proteins. Rather, those aggregation-prone proteins that were present became more likely to be aggregated as the worms aged. These proteins were enriched for b-sheets, already strongly implicated in promoting protein aggregation, and for involvement in developmental processes, particularly proteostasis itself, including components of the proteasome and chaperones. David et al. went on to test the effect on age-related protein aggregation of a mutant in the worm insulin receptor daf-2, already shown to double lifespan and to delay proteotoxicity associated with genetic models of polyglutamine and amyloid beta toxicity in C. elegans. Although there was no detectable difference in protein aggregation in young animals, as they aged the daf-2 mutants showed no increase in protein aggregation with age, and again measurement of the concentrations of specific proteins showed no reduction in total protein present. Thus, in worms, reduced activity of the insulin-signalling pathway may exert a very general beneficial effect on the cellular environment during aging, by reducing the propensity of aggregation-prone proteins to form aggregates. Understanding exactly how this occurs is an interesting future challenge.
These little invertebrates have, for another year, provided many fascinating discoveries for future follow-up in mammals, and they show no sign of losing their ability to provide novel biological insights into mechanisms of aging.
I thank
Though the clinical significance of testosterone deficiency is becoming increasingly apparent, its prevalence in the general population remains unrecognised. A large web-based survey was undertaken over 3 years to study the scale of this missed diagnosis.
Methods. An online questionnaire giving the symptoms characterising testosterone deficiency syndrome (Aging Male Symptoms -AMS -scale) was set up on three web sites, together with questions about possible contributory factors. Results. Of over 10,000 men, mainly from the UK and USA, who responded, 80% had moderate or severe scores likely to benefit from testosterone replacement therapy (TRT). The average age was 52, but with many in their 40s when the diagnosis of 'late onset hypogonadism' is not generally considered. Other possible contributory factors to the high testosterone deficiency scores reported were obesity (29%), alcohol (17.3%), testicular problems such as mumps orchitis (11.4%), prostate problems (5.6%), urinary infection (5.2%) and diabetes 5.7%. Conclusions. In this self-selected large international sample of men, there was a very high prevalence of scores which if clinically relevant would warrant a therapeutic trial of testosterone treatment. This study suggests that there are large numbers of men in the community whose testosterone deficiency is neither being diagnosed nor treated.
There is increasing recognition among doctors that testosterone deficiency is a common and important condition. As well as the key symptoms of loss of energy, drive, libido and impaired erectile function, there is often depression, memory loss and increased irritability which can impair work, home and social life.
This has also been causally linked to a wide range of medical conditions [1] including metabolic syndrome [2], diabetes [3], cardiovascular disease [4], osteoporosis [5], Alzheimer's disease [6], depression [7], frailty [8] and even premature death [9].
An even more important point is that not only are there strong theoretical and epidemiological links to these disorders, but testosterone treatment has also been shown to be effective in the prevention and treatment of many of them, especially diseases linked with the recent epidemic of obesity such as metabolic syndrome, Type 2 diabetes [10] and coronary heart disease [11]. Also many of the conditions seen in aging populations world-wide which lead to frailty and disability can be prevented by such treatment, making it an effective form of preventive medicine. In view of these benefits, it was decided to undertake a webbased survey of the prevalence of symptomatic testosterone deficiency.
The standard English version of the 17-item Aging Male Symptoms (AMS) scale [12], which is well validated and widely used in research studies both to detect symptoms which might suggest testosterone deficiency syndrome (TDS) and monitor its treatment [13,14], was made available on the web-site of The Society for the Study of Androgen Deficiency (SSAD -Andropause Society) and two other Centre for Men's Health web sites.
Additional questions were asked which were thought would give information about factors which might be related to androgen deficiency. As well as age, these included questions about whether the respondent had adult mumps, orchitis or other testicular problems, prostate operation or inflammation, persistent urinary infection, vasectomy, recent weight gain, diabetes or a high alcohol intake. For a representative sub-sample, demographic data on city and country of domicile and occupation were available.
On completion of the full questionnaire, the subject could immediately learn his total score and rating according to the standards for this scale: 17-26 none, 27-36 little, 37-49 moderate, over 50 severe. Given the high response rates to testosterone treatment of scores at these levels, moderate and severe scores were considered to represent testosterone deficiency, and the advice given to consult a physician.
Of the 13,861 respondents taking the online AMS questionnaire between August 2007 and January 2010 (28 months), 10,896 gave permission for their data to be analysed, and this was used in the study. The information from the full questionnaire from all three web-sites was down-loaded to an Excel spreadsheet for analysis, using PASW 18 statistics programme.
With the exception of the respondents to the SSAD site who gave a higher 'Permission to use for analysis' rate than either of the two commercial sites, i.e. was more trusted, the data were very similar, from all three web-sites, as shown by the mean age and total AMS scores, as well as in the more detailed analyses, two-thirds coming from the SSAD site (Table I). The combined data from the three sites are therefore used in this analysis.
The age range of the respondents was 16-89 (mean 52 years), which is close to the mean age of patients attending the Centre for Men's health prior to treatment, and recognised as the time when these symptoms most commonly present as the 'Andropause' or 'Male Menopause' (Figure 1). Though symptoms of testosterone deficiency rarely present below the age of 30, and can be recognised even in the very elderly, the AMS scale, despite its name is not significantly age-related, unlike other commonly used scales [15].
While there may have been a few respondents in their 20s or earlier who were just experimenting with the AMS questionnaire, it is remarkable that because of spread of the normal distribution curve, 13% were below the age of 40, and another 27% in their 40s, i.e. 40% at ages when 'Late Onset Hypogonadism' is generally not recognised as occurring. However it has been found that the AMS scale can be used to measure and compare the health-related quality of life in those of less than 40 years of age on the basis of similar European normative values [16]. The surprisingly high prevalence of raised scores in the younger age groups may be due to the increasing impact of work stress in these age groups, which has been shown to decrease testosterone synthesis [17,18] and increase resistance to its action [19].
Seventy percent of respondents were from the UK, 10% from America, and the remainder from mainly English speaking countries all over the World. There was a wide scatter of occupations, and about 25% were retired as would be expected from the age distribution.
As rated by their total AMS scores overall 30.3% had 'moderate' and another 49.7% 'severe' symptoms, i.e. 80% having symptoms which gave them a high probability of having TDS and which would warrant a therapeutic trial (Figure 2). The proportion of 'severe' symptoms was slightly lower in the 30s and less and 70s and over age groups, peaking in the 50s, but this was balanced by raised 'moderate' scores in the former groups.
Of the individual questions, as might be expected, the three most highly rated questions were in the sexual subscale of the AMS questionnaire, i.e. decreases in ability and frequency of sexual activity, morning erections and libido.
Table I. Analysis of data from three web-sites. All data SSAD CMH MHC Number with permission 108,96 6885 1001 3010 Source of data 100% 63% 9% 28% % Refusing permission 22% 17% 32% 26% Mean age 52.0 51.7 51.8 52.5 Total AMS score 48.6 48.9 50.3 52.4 Figure 1. Age distribution of web AMS respondents.
The prevalence of related factors is described in the discussion section.
This was not a random survey of the general population, but was biased towards those searching the web for an explanation of symptoms they thought might be related to testosterone deficiency. Also there could be over-representation of younger and more computer literate age groups.
However, of this self-selected group of men who answered an online version of the widely studied and well validated AMS questionnaire, the results clearly show that a very large number of men identify with these classic symptoms of TDS. If they had been seen in a clinic setting where other causes of these symptoms could be excluded, as many as 80% might warrant a therapeutic trial of testosterone treatment, with a high probability of a positive response [20].
In order of frequency, positive responses to the additional questions following those of the AMS scale (Figure 3) were as follows:
Recent weight gain (29% overall). This was noted most frequently in the 40-and 50-years-old age groups. Whether this could be considered as a cause or a symptom is unclear, but this common complaint of patients presenting for testosterone treatment can be explained by the effect of testosterone deficiency on the multipotent stem cell [21]. This promotes differentiation from progenitor cells for muscle, bone and epithelium, towards the pre-adipocyte and adipocyte lineage. The resulting increase in adipocytokines contributes to the greater prevalence of metabolic syndrome, Type 2 diabetes, hypertension and circulatory diseases seen in obese individuals both in population surveys and in medical clinics. This has caused the adipocyte to be termed 'The axis of evil' [21].
The reported frequency of this operation rose rapidly from 5% in the 30s age group, to a peak of 34% in the 60s group, falling away to the over 70s age group when vasectomy was less common.
This is thought to be about twice the frequency of this operation in the general population, and is a significant finding of this study. It is closely similar to the 24% found in the UK Androgen Study (UKAS) [22]. This could be because vasectomy is a marker for the 'High-T' men who are more aware of symptoms when their testosterone drops later in life. Alternatively, vasectomy could cause long-term damage to the testosterone-producing capacity of the testes many years after the operation for reasons yet to be explained [22].
Alcohol (17.3% overall). This response to the question about 'High alcohol intake' rose to a maximum of 20% in the 40s, and is likely to have been an underestimate because the population is modest about reporting what it considers a reasonable intake in relation to the proven reduction in testosterone caused by alcohol, particularly with binge drinking. Also, unlike the liver, the testis has limited powers of regeneration, and high alcohol intake, even many years previously, can cause lasting damage to both spermatogenesis and testosterone production.
'Testicular problems and orchitis' (11.4% overall). While mumps orchitis has been described as the classic example of an infection causing an endocrine disorder, other viral infections such as glandular fever can also cause testicular damage and orchitis, as can a variety of sexually transmitted diseases. The reported prevalence of these testis-related conditions was as might be expected highest in the youngest group, but remained relatively constant at about 12% at other ages. 'Prostate operations and infections' (5.6% overall) and urinary infections (5.2% overall). As well as being stressful events likely to interfere with sexual activity, and directly or indirectly suppress testosterone production, benign prostatic enlargement has been linked with higher oestrogen levels, which especially in the obese individual can reduce androgen production.
Prostate related problems and urinary infections combined rose with age as would be expected to 17% in the 60s and 35% in the over 70s. Diabetes (5.7% overall). This rose from 2% in the 1930s to 10% in the 1970s. Together with obesity, it is recognised as one of the major predisposing causes of androgen deficiency [23]. Figures available for Ireland show a population prevalence for all cases of diabetes in 2005 of 5.4% in Northern Ireland and 4.7% in the Republic of Ireland. Taking into account population change and the linear rise in obesity, these figures are projected to rise to 6.3% in Northern Ireland and 5.6% in the Republic of Ireland by 2015 [24]. Significantly, low testosterone levels have been found in up to 50% of diabetics, and more importantly, the complications of this condition have been reduced by testosterone treatment [25].
This level of symptoms at all ages is similar to that seen in two studies carried out in urological clinics in Germany where those men with moderate or severe AMS scores were treated with two different testosterone preparations, testosterone injections and gel, with equally good response rates [20,26] (Figure 4).
The uniform and sustained relief of these symptoms given by testosterone treatment has been confirmed by the experience over more than 15 years in 1700 men in the UKAS treated predominantly on the basis of their symptoms rather than androgen levels [19].
Though the AMS was not originally developed as a tool for diagnosing androgen deficiency, the majority of the questions in it overlap with the typical symptoms of TDS, which have been recognised and consistent since they were first described over 60 years ago by Dr. August Werner [27]. This questionnaire makes it possible to suspect TDS in the clinical setting where other causes of these symptoms can be excluded.
The paradox is that these characteristic symptoms of testosterone deficiency are very poorly correlated with total testosterone (TT) or other androgen levels in the blood. This can be explained by the concept of 'androgen resistance' [19]. As with insulin in maturity onset diabetes mellitus, there can be both insufficient production and variable degrees of resistance to the action of androgens operating at several levels in the body simultaneously, with these factors becoming progressively worse with aging, adverse lifestyle, other disease processes, and a wide range of medications. The mechanisms by which androgen deficiency acts can be considered at five different levels:
1. Impaired androgen synthesis or regulation. 2. Increased androgen binding. 3. Reduced tissue responsiveness. 4. Decreased androgen receptor activity.
This provides an explanation of why scores on a range of questionnaires listing symptoms of testosterone deficiency consistently fail to predict low androgen levels.
For example, a recent report on the AMS scale, states that 'the total AMS score was not significantly associated with TT' [28]. Similarly, using three questionnaires, including the AMS and Androgen Deficiency in the Adult Male -St Louis (ADAM) scales, no relationship was found between symptomatology and any of a battery of eight endocrine assays, including TT and free testosterone (FT), other than possibly age-related declines in dehydroepiandrosterone (DHEA) and insulin-like growth factor 1 (IGF-1) [29]. Further, investigation of a group of 81 Belgian men aged 53-66 (mean 59) concluded 'there was no correlation between AMS (total and subscales) and testosterone levels [30], while the same group in a study of 161 more elderly men aged 74-89 (mean 78) also showed no correlation between symptom scores and TT, FT, or bioavailable testosterone (BT)' [31].
One of the studies most clearly highlighting the paradox dividing androgen deficiency symptom scales and laboratory measures is that of Miwa et al. in 2006, who found no correlation between the total and psychological, somatic or sexual domain scores of the AMS and serum levels of TT, FT, estradiol (E2), luteinizing hormone (LH), follicle stimulating hormone (FSH), dehydroepiandrosterone-sulphate (DHEAS), or growth hormone (GH) [32].
Because of this high sensitivity but low specificity of questionnaires to detect patients with low levels of androgens, the complexity of factors involved in androgen resistance, and the lack of validity of androgen assays [33], it seems logical to adopt the suggestion endorsed by Black et al. [34], which is where typical symptoms or conditions known to be related to androgen deficiency occur, that a 3-month therapeutic trial of testosterone treatment be given, providing it appears clinically justified.
This coincides with the emerging view that 'An emphasis and reliance on serum T alone hinders the clinician's ability to manage testosterone deficiency syndromes (TDS)' [29]. Low total testosterone is just the tip of the iceberg of androgen deficiency, and is a poor marker of the underlying larger mass of symptoms and metabolic disturbances.
The high level of symptoms reported by these men in the web survey highlights the findings in a recently published study by one of the report's authors 'Time for International Action on Treating Testosterone Deficiency Symptoms' [35]. This gives world-wide data showing that testosterone deficiency is the commonest hormonal disorder in men, but the least commonly treated. In the UK and all European countries, as well as Australia, Japan and Russia, of the 20% of men over the age of 50 who on a symptomatic basis can be rated as deficient in this hormone, less than 1% are being treated. The USA treats about 8%. This makes testosterone deficiency the most common endocrine disorder in men, and yet the least often diagnosed and treated.
Diagnosis and treatment of testosterone deficiency is economic and safe, and represents an important part of the preventive medicine of the future. This is not a condition which we can afford to leave untreated for either the health or welfare of an aging society.
In the light of the poor quality of life, energy, vitality and sexual function of the testosterone deficient man, as well as the associated conditions such as metabolic syndrome, diabetes, cardiovascular disease, and osteoporosis, how long can the medical community or health services afford to let men continue to go without this treatment?
The authors report no conflicts of interest. The authors alone are responsible for the content and writing of the paper.
Objectives: To assess the concurrent and the construct validity of the Euro-D in older Thai persons. Method: Eight local psychiatrists used the major depressive episode section of the Mini International Neuropsychiatric Interview to interview 150 consecutive psychiatric clinic attendees. A trained interviewer administered the Euro-D. We used receiver operating characteristic (ROC) analysis to assess the overall discriminability of the Euro-D scale and principal components factor analysis to assess its construct validity.
Results: The area under the ROC curve for the Euro-D with respect to major depressive episode was 0.78 [95% confidence interval (CI) 0.70-0.90] indicating moderately good discriminability. At a cut-point of 5/6 the sensitivity for major depressive episodes is 84.3%, specificity 58.6%, and kappa 0.37 (95% CI 0.22-0.52) indicating fair concordance. However, at the 3/4 cut-point recommended from European studies there is high sensitivity (94%) but poor specificity (34%). The principal components analysis suggested four factors. The first two factors conformed to affective suffering (depression, suicidality and tearfulness) and motivation (interest, concentration and enjoyment). Sleep and appetite constituted a separate factor, whereas pessimism loaded on its own factor. Conclusion: Among Thai psychiatric clinic attendees Euro-D is moderately valid for major depression. A much higher cut-point may be required than that which is usually advocated. The Thai version also shares two common factors as reported from most of previous studies.
Depression in older persons is reported to be associated with substantially reduced quality of life and increased mortality (Penninx et al., 1999;Penninx, Leveille, Ferrucci, van Eijk, & Guralnik, 1999) and increased use of all health and social service resources (Watts et al., 2002). In spite of its public health significance, levels of recognition and treatment of depression by physician are reported to be rather low (Crawford, Prince, Menezes, & Mann, 1998;Dearman, Waheed, Nathoo, & Baldwin, 2006;Koenig, 2007). There has been little research from Thailand on depression in older people. Given the importance of detection of depression in older people, there is a need to validate a Thai version of a depression screening instrument.
Euro-D was an instrument developed by 14 European countries for screening depression in the elderly (Prince et al., 1999). It has also been used in low income and middle income countries in other regions (Castro-Costa et al., 2007;Prince et al., 2004). Its concurrent and criterion validity across 14 different settings in Europe was satisfactory (Prince et al., 1999). Most of the previous Euro-D studies suggest two common factors; affective suffering (including depression, tearfulness and wishing to death) and motivation (including loss of interest, poor concentration and lack of enjoyment). Evidence for its cross-cultural validity in developing countries including China, India, Latin America and Africa suggest a similar factor structure across these settings (Prince et al., 2004). There is as yet no evidence for the validity of Euro-D in Thailand.
In the current study we attempted to test Euro-D's internal reliability, its concurrent validity against the diagnosis of major depressive episodes provided by the mini international neuropsychiatric interview (MINI) (Sheehan et al., 1998) and its construct using principal component analysis (PCA) on a sample of psychiatric clinic elderly attendees.
We sampled consecutive patients attending a busy outpatient practice in a psychiatric hospital in the suburb of Eastern Bangkok. Patients aged 60 or above were approached in the waiting room when they attended the clinic to see their doctor. There were eight psychiatrists participating in the study. A research worker, who was a postgraduate student in population and social research, approached the patients and asked for informed consent. Study was conducted according to the guidelines issued by the institutional review board of the hospital involved.
Procedures After reading a study information sheet, those who agreed to take part were interviewed either before or after seeing their doctor in a private room on the premises. Those with evidence of severe cognitive impairment, dangerous behaviours, very frail, or with limitations of comprehension were excluded. We continued recruitment until our target of 150 completed interviews was met. The majority of the patients were presenting with common conditions such as anxiety, depression, dementia, and substance-related disorders.
Interviews with Euro-D were carried out before or after the participants saw their doctors. Each of the eight psychiatrists administered MINI-Thai version, the major depressive episode section (Kittiratpaiboon & Wongkampin, 2004), while the trained research worker administered Euro-D. We blinded the MINI interviewers to the Euro-D responses obtained by the research worker and vice versa.
The Euro-D is a structured common depressive symptoms scale derived from the geriatric mental state-AGECAT (GMS-AGECAT) interview (Copeland, Dewey, & Griffiths-Jones, 1986), SHORT-CARE (Gurland, Golden, Teresi, & Challop, 1984) and other measures including the Centre for Epidemiological Studies Depression Scale (CES-D) (Radloff, 1977), Zung Self-Rating Depression Scale (ZSDS) (Zung, 1965), and the Comprehensive Psychological Rating Scale (CPRS) (Asberg & Schalling, 1979). It was designed to be administered by trained lay interviewers and only consists of 12 items dealing with: depression, pessimism, wishing death, guilt, sleep, interest, irritability, appetite, fatigue, concentration, enjoyment and tearfulness. A cut-point of 3/4 was identified using receiver operating characteristic (ROC) analysis in studies carried out in 14 European countries and produced a range of sensitivity (63%-83%) and specificity (49%-95%). Principal components analysis generated two factors (affective suffering and motivation) that were common to nearly every participating European country (Prince et al., 1999) and for Indian, Latin-American and Caribbean centres (Prince et al., 2004). Internal consistency was universally satisfactory, ranging from 0.83-0.93 (Prince et al., 2004).
In Thailand, the Euro-D items appeared to cover symptoms recognised locally as common in psychological disorders in older adults. The original version was carefully translated and back-translated into English. First, a team of bilingual mental health professionals and social scientists developed the first translation, paying particular attention to conceptual and semantic equivalence. Two English speaking old age psychiatrist with training and extensive experience with the GMS assisted to answer any questions about the original version. The first translation was piloted on an urban community sample of 12 elderlies to test its clarity and comprehensibility. Some questions (e.g., sleep, irritability, appetite, weight loss) simply required translation. Others were adapted to ensure local equivalence or to aid comprehension (i.e. depression, concentration and pessimism). 'Can you concentrate on entertainment or reading' was replaced with 'Can you concentrate on daily activities that you like, such as listening to the monks' teachings, the radio or watching television'. 'Have you been feeling depressed?' was replaced with 'have you been feeling sad, gloomy or in despair'. It was most difficult to translate the question on pessimism. The original question asks 'tell me your hopes for the future' and rates pessimism if the person cannot describe at least one hope. In the Thai context we anticipated that older people, most of whom were Buddhist, would not mention any hopes because they were expected at their age to view life with contentment and to take each day as it comes. Asking about 'future hopes' might therefore elicit feelings of unfamiliarity and discomfort. This could lead to misinterpreting true pessimism. We changed the wording to 'do you have any hope that in the near future good things might happen to you?'. We hoped that the modified closed question would enable to valid response.
The MINI was used as gold standard criterion measure. MINI is a short structured diagnostic interview, developed for DSM-IV and ICD-10 psychiatric disorders. The section on major depressive episode was used. The Thai version of MINI was validated against diagnoses made by local psychiatrists in a clinical setting (Kittiratpaiboon & Wongkampin, 2004) and used as a gold standard criterion against a newly developed screening test in a previous study (Arunpongpisal et al., 2006). It was also used as a main diagnostic instrument in a recent national mental health survey (Department of Mental Health, 2003).
The output data from the MINI and Euro-D were read into STATA and the following parameters were assessed.
(1) The prevalence of major depressive disorder according to each method of assessment.
(2) The absolute numbers of participants in each cell of the cross tabulation by test and criterion status, that is true positive, false negative, false positive and true negative. (3) The kappa with 95% confidence intervals (CI) as the primary measure of agreement. (4) Coefficient of agreement (the proportion of participants classified as positive by either assessment that are classified as positive by both. (5) To assess concurrent validity, the scores of the Euro-D were compared with the diagnosis of major depressive disorder provided by the MINI. A positive case was defined as an individual who was diagnosed as having a major depressive episode. The area under the rROC, with the correspondent 95% CI, was estimated for major depressive disorder criterion. The optimal cut-off point was identified and sensitivity, specificity, positive, and negative predictive values were calculated. (6) The inter-item correlations and Cronbach's alpha coefficients were calculated as measures of the internal consistency of the Euro-D. (7) A PCA of the Euro-D was performed, in an exploratory approach.
Over a 2-month period we approached 167 patients of whom 150 completed the interview (response rate 89.8%). The demographic characteristics of the sample are shown in Table 1. The principal diagnoses were mood disorders (29.3%), anxiety disorders (23.3%), schizophrenia and delusional disorders (16.7%), organic mental disorders (16%), headache (4%), substance abuse and dependence (2.7%) and other disorders (7.3%).
Table 2 shows the prevalence of each Euro-D symptom across the whole sample and the prevalence of each Euro-D symptom comparing patients who were diagnosed and not diagnosed with depression. As shown in Table 2, the mean number of items rated as present during the past month per subject was 4.9 (95% CI 4.4-5.3). The frequencies of items ranged from 16.7% (concentration) to 58.7% (irritability). Only one item (pessimism) did not significantly discriminate patients depressed from those who were not. Over 75% of depressed patients reported that they had been depressed, had sleep trouble, had been irritable, and felt fatigue.
The concordance between Euro-D and MINI for major depressive episode was fair. The level of agreement for the sample (kappa) was 0.37 (95% CI 0.22-0.52). Euro-D tended to overdiagnosis with respect to MINI with a prevalence of 39.3% compared with 34% for major depressive disorder according to MINI. The area under ROC curve for the overall discriminability of the Euro-D scale against the criterion of MINI major depressive disorder was 0.78 (95% CI 0.70-0.85) (Figure 1). Inspection of psychometric indices at different cut-points suggests a cut-point of 5/6 with sensitivity of 84.3% and specificity of 58.6%. The positive and negative predictive values were 55.9 and 80.2, respectively (Table 3). A higher cut-point of 6/7 will yield lower sensitivity (64.7%) and higher specificity (73.7%) (Table 3).
The Cronbach's alpha for the total scale was sufficiently high at 0.72, although the internal consistency could be improved upon by the omission of EURO4 (pessimism), which marginally increased the standardised alpha to 0.74.
For factor analysis, the adequacy of the data for factor analysis was assessed by inspecting the correlation matrix, which showed that 10 coefficients in 66 were 40.3. Moreover, the Kaiser-Mayer-Oklin value was 0.77 and the Barlett's test of sphericity achieved a high level of statistical significance (p 5 0.001), supporting the factorability of the correlation matrix.
Principal component analysis revealed the presence of at least three components with eigenvalues exceeding 1, following inspection of the scree plot (Table 4). Factor one conformed to affective suffering (depression, suicidality and tearfulness) with smaller contribution (0.4-0.5) from guilt, irritability and fatigue, and factor two indicated motivation (interest, concentration and enjoyment) but negligible loadings from other items. The affective suffering factor account for 26% of the variance in the items, and the motivation factors for 12%. The third factor, with eigenvalue marginally over 1, was largely loaded on by sleep and appetite.
Depression is an important condition affecting older people's health and quality of life which needs a screening instrument to identify individuals at risk. Euro-D is a useful instrument which serves this purpose. To our knowledge, this is the first study to test the reliability and validity of Thai version of Euro-D. Attempts were made to ensure that content and semantic equivalence were achieved by strict translation and back-translation procedures with minor modifications to make the items more comprehensible and relevant to the Thai context. It was found to be feasible and easily administrable by trained interviewers in this clinical setting.
Our results show that the frequency of individual items rated as present during the previous month in this study was somewhat higher than those reported in a previous Euro-D studies in European countries (Prince et al., 1999). The prevalence of the pessimism item was particularly high and it did not discriminate well between the depressed and the nondepressed groups. Although we altered the original version of the item it was still the case that few older people volunteered any hopes for the future. The Buddhist doctrines of contentment and acceptance of the current life situation are likely to explain this view. The hope of better living in the future, for instance in terms of gaining material good, contradicts the traditional path toward merit in the afterlife (Pfanner & Ingersoll, 1962). Guilt was also highly prevalent. This may be due to people misunderstanding the guilt item or the tendency to volunteer feeling a burden on children and this being rated as a symptom. It is widely perceived among Thai older people that they should rely on their children when they get old and dependent (Knodel, Chayovan, & Siriboon, 1992), but many expressed guilt about this. Fatigue was also high, possibly because there were extra untreated physical illnesses or because many older Thai people are still working hard to make ends meet even after their retirement. Further studies using item response theory analysis to explore for differential item functioning (DIF) may be helpful to investigate whether there is item bias in the Thai version and to explain the high levels of these symptoms. The area under the ROC curve for the overall discriminability of the Euro-D scale against the criterion of any MINI depressive disorders was 0.78 (0.70-0.85). This rather impressive degree of overall prediction, coupled with the high sensitivity (94%) and poor specificity (34%) of the Euro-D cut-point of 3/4 suggested that this cut-point is too low for the Thai clinical sample. A cut-off score of 5/6 had lower sensitivity (84.3%) but higher specificity (58.6%) to screen major depressive episode compared with the usual cut-off reported previously. This higher cut-off might reflect more depressive symptoms among psychiatric patients. The majority of cases of depressive disorders in clinical settings obviously tend to be more severe compared with cases in general population. In community settings, cases may tend to have lower screening score. Further validation studies in the community setting will be needed. Taking account of the Thai context and improving the clarity of the wordings may help to improve the overall sensitivity and specificity.
The Cronbach alpha coefficient shows a high internal consistency for the whole Euro-D questionnaire, although the item on pessimism would appear to improve the overall internal consistency if deleted. As discussed earlier, the concept of pessimism may be different among Thai elderly. Finding a better way to address the issue of pessimism may help to improve the overall psychometric properties. One of the questions raised in a validation study is about what is the real construct behind the instrument to be validated. The Euro-D was mainly derived from GMS/AGECAT and designed to be used by trained lay interviewers, whereas MINI was based on ICD-10 and DSM-IV criteria and used by mental health professionals as a diagnostic tool for major depressive episode. These discrepancies may provide an explanation for the fair kappa for the agreement between Euro-D and MINI cases.
It is striking that the internal structure of the Euro-D as suggested by principal components analysis is consistent with most of the previous studies (Castro-Costa et al., 2007;Prince et al., 2004). Our version of the Euro-D also shares two common constructs, the first one being affective suffering, including depression, tearfulness and wishing to die, and second one being motivation including loss of interest, poor concentration and lack of enjoyment. They were mapped well onto the two common main factors reported previously, although there were slight differences in the individual items included in each factor (Castro-Costa et al., 2007;Prince et al., 1999;Prince et al., 2004). The two extra factors were identified in our study, one including sleep and appetite and the other one including pessimism. The three items had eigenvalues of about 1. This may not be so surprising as three-or four-factor structure was also reported in some of the European settings. For example, sleep and appetite were found to be loading on another factor in Sweden, Netherlands and Iceland and pessimism loading on a separate factor in Germany (Prince et al., 1999). Although the four-factor structure is somewhat different, it remains difficult to establish that it does vary from the European studies. Confirmatory factor analysis may be a useful technique to test whether the four-factor model has better overall goodness-of-fit than the two-factor solution.
There were some limitations in this study. First, the study was conducted in a psychiatric hospital on a sample of patients with a variety of mental conditions. The study sample was obviously different from those based in communities or general practice. Second, the sample size available for this validation study was relatively small. More studies with a larger sample size in a variety of clinical settings are warranted. In addition, item bias cannot be excluded as an explanation for the relatively high prevalences of many Euro-D items in our study. Further studies in crosscultural settings using item response theory analysis should provide more clues.
In conclusion, this is the first study to test the reliability and validity of Thai version of Euro-D. This validation study reports the use of Euro-D in a clinical setting. The instrument has demonstrated rather different properties to those shown in previous studies. The different cut-point found in the current study suggests the need for caution in interpretation of Euro-D scores in Thailand. Further studies are still needed to validate Euro-D in the Thai community.
Aging & Mental Health
This research was supported by the
Depression 76 50.7 79.7 31.9 0.000 Pessimism 85 56.7 64.4 51.7 0.123 Suicidality 40 26.7 50.9 11.0 0.000 Guilt 44 29.3 52.5 14.3 0.000 Sleep 84 56.0 76.3 42.9 0.000 Interest 47 31.3 54.2 16.5 0.000 Irritability 88 58.7 89.8 38.5 0.000 Appetite 64 42.7 71.2 24.2 0.000 Fatigue 86 57.3 93.2 34.1 0.000 Concentration 25 16.7 33.9 5.5 0.000 Enjoyment 40 26.7 44.1 15.4 0.000 Tearfulness 52 34.7 59.3 18.7 0.000 0.00 0.25 0.50 0.75 1.00 Sensitivity 0.00 0.25 0.50 0.75 1.00 1-Specificity Area under ROC curve = 0.7753
A partial solar eclipse occurred in South Korea on 22 July 2009. It started at 09:30 a.m. and lasted until 12:14 LST with coverage of between 76.8% and 93.1% of the sun. The observed atmospheric effects of the eclipse are presented. It was found that from the onset of the eclipse, solar radiation was reduced by as much as 88.1∼ 89.9% at the present research centre. Also, during the eclipse, air temperature decreased slightly or remained almost unchanged. After the eclipse, however, it rose by 2.5 to 4.5°C at observed stations. Meanwhile, relative humidity increased and wind speeds were lowered by the eclipse. Ground-level ozone was observed to decrease during the event.
Atmospheric effects of solar eclipses have been discussed in earlier studies (Anderson and Keefer 1975;Mims and Mims 1993;Segal et al. 1996;Founder et al. 2007;Hakan 2009) and relatively accurate predictions of solar eclipses in the near future have been provided by astronomers (e.g. Espenak and Anderson 2006).
As predicted, a partial solar eclipse occurred in Korea on 22 July 2009. The phases of the solar eclipse are shown in Fig. 1. The purpose of the present study is to describe atmospheric disturbances produced by the partial solar eclipse which occurred in South Korea.
The solar eclipse began at the southernmost island, Jeju, at 09:30 Local Standard Time (LST), and it ended at 12:14 LST at Dok-dho Island in the Korea East Sea. At the research centre (KCAER) in Cheongju-Cheongwon, in central South Korea, the eclipse started at 09:34 and it lasted until 12:08 LST, a total of 2 h and 42 min. The percentage coverage of the sun by the moon was recorded as ranging from 76.8% to 93.1%. In Cheongju-Cheongwon, the maximum coverage of the sun was 81.3%. The largest part of the sun was hidden as seen from the southernmost island at 10:48 LST, shown in Fig. 1. It was the first significant eclipse in Korea for 61 years in which such a large part of the sun was covered. A total eclipse was seen in Korea on 19 August 1887.
A meteorological map of 22 July 2009 is included in Fig. 2. Korea was under the influence of an anticyclone and fair weather with variable low clouds prevailed during the period of the partial eclipse. There were more clouds in southern Korea than in central Korea (cloud amount in Table 1). This was related to a stationary and rainy front situated along the line from Shanghai, China to the south of Japan. Likewise, there were considerable amounts of clouds over the Korea South Sea (Fig. 3).
Meanwhile, satellite monitoring was carried out at the research centre, and satellite images are shown in Fig. 3. It can be seen that the visible Channel 1 image from the geostationary satellite MTSAT (Fig. 3b) clearly shows an extended dark area in Shanghai, the East China Sea and the Korea South Sea areas during the solar eclipse; sunlight in the dark phase of the solar eclipse, the visible Channel 1 camera could not detect low-level clouds in these areas. Nevertheless, the IR Channels in Fig. 3a indicate a cloud band over these areas. The long cloud band in Fig. 3a is associated with the stationary weather front shown in Fig. 2. The polar-orbiting NOAA satellite image (not shown) also indicates fair weather over the Korean Peninsula after the solar eclipse.
Figure 4 shows variations of solar radiation (W m -2 ) and air temperature (°C). The meteorological elements were measured with standard instruments at the KCAER. The centre is located in a rural area 80 m above sea level. We found that the amplitude of radiation variations was large in comparison with air temperature. On the morning of 22 July, in Cheongwon, there were scattered stratus clouds and fair weather prevailed during the eclipse.
The radiation curve in Fig. 4 suggests that radiation was very sensitive to the eclipse and low-level stratus clouds, whereas there were gaps and delays in air temperature variations from radiation. A small jump of radiation energy between 10:05 and 10:20 LST resulted from diffusive radiation under stratus clouds. Otherwise a steep reduction in radiation was recorded during the eclipse. Radiation was measured with Pyranometer (Li-cor PY45671). At 09:40 LST, solar radiation was 760.0 Wm -2 , and a minimum value of 90.1 Wm -2 was observed at 10:35 LST. With the recovery of the full sun it went up to 887.0 Wm -2 at 12:10 LST. The minimum value and satellite images clearly show the striking negative impact on atmospheric radiation to which the lower atmosphere is exposed.
Air temperature at the research centre at 09:25 and 09:30 LST was steady at 22.2°C, and it went up to 22.9 and 23.0°C at 09:40 and 09:45 LST, respectively. At 09:55, however, air temperature decreased to 22.2°C, and at 10:00 it was 22.3°C (Founder et al. 2007;Segal et al. 1996). At 10:35 and 10:40 LST, air temperature was 22.3°C, coinciding with the smallest phase of the sun. Air temperature was slowly increasing from 11:30 LST, and a quasi-steady air temperature was maintained by the eclipse in the range of 22.2∼23.0°C for about 2 h. Around the time that the eclipse ended at 12:00 LST, air temperature at the research centre rose promptly to 24.7°C (and at 13:00 LST it was 25.7°C). This translates to the partial eclipse Fig. 2 A surface meteorological map (of the Korean Meteorological Administration) at 09:00 LST, 22 July 2009
preventing the air temperature from increasing by at least 2.5°C. The estimation was done with the actual data between pre-and post-eclipse air temperature. Moreover, from 01:00 to 06:00 LST, 21 July, KCAER received rainfall as much as 48.3 mm, and this could have resulted in the smaller decrease in air temperature as the radiation used for evaporation. Meanwhile, it should be noted that the maximum air temperature at the research centre was 28.6°C at 17:00 LST.
Furthermore, at two nearby air bases, air temperature from 10:00 to 11:00 LST decreased by 0.2 and 0.7°C, respectively. From 11:00 to 13:00 LST, two air bases recorded an increase of 4.3 and 3.7°C, respectively. In turn the air temperature decreases at the air bases during the eclipse were 4.5 and 4.4°C, respectively. These are comparable with the 2.5°C decrease observed at the research centre in Cheongwon.
Until 09:00 LST, it was misty and visibility was 4∼5 km in the Cheongwon area. Observed relative humidity at 09:00 LST was 88% and it went down gradually to 79% at 09:45. When the eclipse was in progress, the relative humidity slowly increased to over 79% until 11:15 LST. From this time on, relative humidity decreased steadily with the increase in radiation and air temperature. Relative humidity depends strongly on the air temperature. This relation is seen in Fig. 5, but is reversely correlated (Hakan 2009). When air temperature was low prior to the eclipse, relative humidity was recorded at a high level. This reverse correlation is also observed after 11:20 LST. During the mid-stage of the eclipse, the changes in relative humidity were in the opposite direction to the changes in air temperature.
The atmospheric pressure was steady at 1,003 hPa with no change during the eclipse, but we found that surface wind speed decreased from 1.2 to 0∼0.2 ms -1 (Fig. 6). In the earlier part of eclipse until 10:00 LST, measured winds were 0.8∼1.5 ms -1 . After 11:40 LST, observed horizontal winds increased to 0.4∼1.3 ms -1 . In general, solar heating in the morning hours generates advective and convective airflows, and in turn these were negligible during the eclipse without radiation energy inputs. Also, the air during the morning hours was relatively stable with stratus clouds for the decrease in winds.
Atmospheric ozone is produced by solar radiation and, in the warm air, ozone precursors, like NO x , hydrocarbons etc., are agents for the generation of ground-level ozone (Chung 1977;Finlayson-Pitts and Pitts 2000). In order to find negative effects of solar eclipse on ozone production, measured values with ozone monitors (TECO 49C) were examined with regression estimates. Figure 7a shows an irregular increase in ozone values, but ozone values at KCAER during the eclipse were lower than on a regression curve, especially around 10:40 to 12:20 LST. It should be
Element Radiation (MJ m -2 ) Temperature (°C) RH (%) Cloud amt (*/10) Station 11-10h 13-11h 11-10h 13-11h 11-10h Seoul -0.93 -2.31 -0.6 -2.2 0 3∼4 Cheongju -0.10 -2.18 -0.1 -3.9 0 3∼5 Busan -1.25 -2.35 -1.9 -1.9 10 5∼8 Mokpo -0.92 -2.72 -0.8 -2.1 6 5∼6 Seogwipo ---0.8 -3.6 1 7∼8 Mean -0.80 -2.39 -0.84 -2.74 3.4 5.4 Table 1 Atmospheric impacts of the solar eclipse observed at five KMA stations in Korea Fig. 3 Satellite images of 22 July 2009: a MTSAT composite image of Channels 1, 2 and 4 at 10:30 LST. b MTSAT image of visible Channel 1 at 10:30 LST noted that in general ozone values do not necessary to vary linearly. According to 5-min average ozone value at 10:40 LST; however, it was at 49 ppb and decreased to 41∼ 42 ppb until 11:15 LST.
During the morning hours, NO and NO 2 values were in the range of 1∼2 ppb and CO values were in the range of 380∼420 ppb. Concentrations were all measured with TECO analysers (Kim and Chung 2008). With ozone precursors not high enough, during the afternoon hours 96 ppb of ozone were still produced under conditions of high radiation, above 800 Wm -2 with the south-westerly air warmer than 28.0°C at KCAER. Scientists are generally concerned about ozone levels above 80 ppb. KCAER also measured ozone at the western coastal site in the Tae-ahn Peninsula (TAP). This site is located 120 km west of KCAER in central Korea. Observed ozone in TAP at 09:00 LST was 60 ppb, and ozone value measured there at 11:00 LST decreased to 51 ppb with the eclipse. In addition, ozone data obtained in Cheongju City by the provincial government were also studied. According to Fig. 7b, hourly ozone values in the city also decreased during the period of partial eclipse.
Atmospheric variables observed at the research centre were studied in detail along with the data provided from two airforce bases in Cheongwon.
In order to investigate national disturbances caused by the partial eclipse, we obtained additional data from the Korean Meteorological Administration (KMA). Table 1 lists observed meteorological elements from five KMA stations. The observed cloud amounts (amt) varied from 3/ 10 to 8/10, and the average cloud amount was 5.4/10 for all stations. From 10:00 to 11:00 LST, mean values of solar radiation at the KMA stations decreased to 0.80 MJ m -2 and mean air temperature decreased to 0.84°C. Meanwhile, in comparison with the observed values at 13:00 LST, mean solar radiation at 11:00 LST was as much as 2.39 MJ m -2 lower. This change can be compared with the maximum value of 3.42 MJ m -2 observed at 14:00 LST.
From 10:00 to 11:00 LST observed mean air temperature at the five KMA stations decreased by 0.84°C (Segal et al. 1996) i.e. around the time of maximum eclipse at 10:50 LST. Therefore, we have used the observed mid-eclipse data at 11:00 LST. By 13:00 LST mean air temperature had increased rapidly by as much as 2.74°C (Table 1). In turn, from 10:00 to 13:00 LST, the mean air temperature during the eclipse decreased by 3.58°C at the five KMA stations.
On the other hand, observed mean relative humidity from 10:00 to 11:00 LST increased by 3.4% at five stations in relation to the decrease in air temperature. The observed trends of variations in weather variables at the five KMA stations agree well with the trends of weather variables observed at the research centre.
From air quality analyses, we have found that the concentrations of ground-level ozone decreased during the eclipse. During the eclipse period of about 2.5 h, clearly the destruction of ozone occurred with the lack of solar radiation in Cheongju-Cheongwon.
The partial eclipse that took place on 22 July 2009 generated significant atmospheric disturbances in South Korea. The solar coverage by the moon that morning was chiefly in the range 76.8∼93.1%.
The observed minimum radiation at Cheongwon in central Korea was 90.1 Wm -2 compared with 760.0 just before and 887.0 Wm -2 just after the eclipse. In turn, the radiation energy was down by a mere 10.1∼11.9% from the expected radiation without the eclipse. At the KMA Cheongju station, however, this stood at 15.8%, whereas the northern station in Seoul recorded 31.7% of the expected value. The average minimum radiation at all four KMA stations at the peak of the eclipse was 22.9% from the highest value of hourly radiation received during the day.
During the partial eclipse, the air temperature was lowered by as much as 2.5°C at the research centre. At nearby airforce stations, air temperature decreased by 4.4∼ 4.5°C. On the other hand, at five KMA stations the magnitude of mean air temperature decrease was 2.74°C.
It was also observed that relative humidity increased by up to 3.4% at the research centre during the period of the partial eclipse. Owing to the lack of advection and convection during the eclipse, winds, in general, decreased (Anderson and Keefer 1975) below 0.2 from 0.8∼1.5 ms -1 during the eclipse.
There is, however, a certain discrepancy in observed values of meteorological variables from several stations. This was largely the result of clouds and different types of instruments and observational practices at the various stations.
Acknowledgements Research fund is provided by the
To examine whether individual changes in alcohol consumption among female alcoholics under treatment are predicted by level of and changes in depression and dysfunctional attitudes. Method: A total of 120 women who were treated for alcohol addiction at the Karolinska Hospital in Stockholm (Sweden) were assessed twice over a 2-year period using the Depression scale from the Symptom Checklist-90, the Alcohol Use Inventory and the Dysfunctional Attitude Scale (DAS). Latent growth curve analysis was used. Results: Decrease in alcohol consumption, depression and dysfunctional attitude variables were found at group level. The results also showed significant individual variation in change. Changes in alcohol consumption were predicted by baseline alcohol drinking, as well as by level and changes in depression. Stronger reduction in depression was related to higher level of depression at baseline, and with reduction in dysfunctional attitudes. Different DAS sub-scales resulted in different magnitude of the model relations. Good treatment compliance was related to lower baseline level in depression, but also with higher baseline level in dysfunctional attitudes, and predicted stronger reduction in alcohol consumption. Conclusion: This paper shows the importance of incorporating both individual level and change in depression as predictors of change in alcohol consumption among subjects treated for alcohol addiction. Also, dysfunctional attitudes are both indirectly and directly related to treatment outcome. By incorporating alcohol consumption, depression and dysfunctional attitudes as targets of intervention, treatment compliance and outcome may be enhanced.
Alcohol problems manifest in several ways: in high level of alcohol consumption, harmful drinking patterns and relapse to alcohol drinking after treatment. Alcohol problems and depression co-vary in women, as assessed by cross-sectional and longitudinal epidemiological and clinical studies (Burns et al., 2005;Dixit and Crum, 2000;Haynes et al., 2005;Ramsey et al., 2005). Even if related, improvement in one disorder does not always reduce the other (Burns et al., 2005). Dysfunctional attitudes may interact with depression and thereby also with alcohol problems. The relation between depression and dysfunctional attitudes has mostly been analysed in the context of cognitive therapy studies, in which patients with alcohol addiction have usually been excluded (Oei and Free, 1995). The present study, which is related to an earlier study of the same sample (Haver and Gjestad, 2005), explores the relation between dysfunctional attitudes, depression and alcohol consumption, both at the group and individual level.
Changes in alcohol use disorders (AUDs) over time correlate with baseline levels of alcohol use (Dawson et al., 2007) and with major depression (Greenfield et al., 1998). Depression at intake, as measured by diagnosis or continuous measures of depression severity, predicts the alcohol outcome after treatment among women who suffer from AUDs (Charney et al., 1998;Zilberman et al., 2003). Measurement differences may partially explain different results of co-morbidity analyses, with strongest relations often found when using diagnostic categories of depression (Bradizza et al., 2006;Burns et al., 2005;Haynes et al., 2005). Depression is a strong predictor of relapse to alcohol drinking, with prolonged depression being a stronger predictor than the baseline level of depression (Bradizza et al., 2006;Haver and Gjestad, 2005). Thus, changes in depression are important for later alcohol consumption (Kodl et al., 2008). The latter study showed reduction in depression to occur both spontaneously and as a result of intervention. The degree of alcohol dependence and alcohol consumption at baseline did not correlate with alcohol measures at follow-up. A decrease in co-morbid depression may thus explain why higher levels of alcohol consumption and dependence at baseline do not always correlate with poor outcomes of addiction treatment. One study found an equal decrease in alcohol consumption over a 3-month period in men and women with or without current depression and/or anxiety (Burns et al., 2005). Changes in alcohol consumption and depression have also been studied using autoregressive models (Aneshensel and Huba, 1983;Peirce et al., 2000). The effect of depression on later residual change in alcohol use was of relatively short duration. Another study found that higher baseline depression levels were associated with a decrease in alcohol consumption at Year 1, and with an increase in alcohol consumption at Year 3, indicating a nonlinear relation over time (Schutte et al., 1995(Schutte et al., , 1997)). The close relation between alcohol consumption and depression over time was confirmed in one meta-analysis (Hartka et al., 1991), but not confirmed by others (Nolen-Hoeksema et al., 2006). More advanced statistical methods, such as multilevel modelling, are available for the analysis of change (Curran and Muthén, 1999;Singer and Willett, 2003). For example, one study reported an increase in alcohol consumption over time among clinically depressed patients (women and men) and community controls (Holahan et al., 2004).
These studies show that both the baseline level and especially changes in depression are important for prediction of changes in alcohol problems in general, and specifically in alcohol consumption. These relations will be influenced by the time between measurements, as well as types of measurements used.
Are dysfunctional attitudes related to the level and change in depression and alcohol consumption? Dysfunctional attitudes, defined as rigidly and relatively stable over generalizations within the domains of perfectionism, performance, need for approval and love, omnipotence and autonomy, are related to depressed mood (Beevers and Miller, 2004;Kwon and Oei, 2003). Such attitudes contribute to a cognitive bias where information is processed in an unrealistically negative manner, which increases the risk for development of a negative mood (Beevers and Miller, 2004). Dysfunctional attitudes are seen as a stable vulnerability factor, and predict changes in depression, including onset and repeated relapses (Elkin et al., 2006;Furlong and Oei, 2002;Hamilton and Dobson, 2002;Weich et al., 2003). Conversely, dysfunctional attitudes may be mood dependent, as depression increases the tendency for negative cognitions (Beevers and Miller, 2004). Improvement in depression could thus be related to an attenuation of dysfunctional attitudes. Attitude changes have been found after various interventions in depressed patients and even in waiting list patients to a lesser degree (Oei and Free, 1995). A change in the Dysfunctional Attitude Scale (DAS) was associated with change in the depression score; thus, dysfunctional attitudes may be seen as a relatively stable parameter, but also as a changeable trait (Oei and Free, 1995). Cognitions have been found to directly predict stability and change in alcohol disorders if relevant cognitive items from existing measurements were selected (Ramsey et al., 2002). Another study found the dysfunctional attitudes scale to be related to later problem drinking in a college sample, even when controlling for level of alcohol consumption, gender, age and depression (Heinz et al., 2009). After controlling for cognitive factors, depressive symptoms were not a significant predictor of problem drinking. Thus, depression-specific dysfunctional attitudes relate to alcohol problems, also problematic alcohol consumption, directly or via the level of and changes in depression, and may be an important factor to include also in samples of problem drinkers.
In this study, alcohol consumption was used as a continuous indicator of alcohol problem severity. The use of a continuous alcohol use variable increases the statistical power and possibilities when analysing multivariate statistical models. This study does not take the clinical cut-off between the presence of a clinical diagnostic condition and the nonclinical condition into consideration. However, nondependent sub-threshold levels of drinking among depressed patients are suggested to be clinically important (Ramsey et al., 2005). On the basis of the literature, we expected changes in alcohol consumption to be predicted by its baseline level and by the level of and changes in depression. In addition, it was hypothesized that changes in depression should be predicted by changes in dysfunctional attitudes.
The baseline dysfunctional attitudes score is supposed to be related to the baseline level of depression, and thus being indirectly related to changes in depression. Because of the depression-related content of dysfunctional attitudes, we expected dysfunctional attitudes as a total scale to predict changes in depression exclusively and thereby moderate changes in alcohol consumption indirectly. This would show dysfunctional attitudes to be of clinical importance for alcohol consumption. Sub-dimensions of the total DAS were tested in separate models to explore potential differences in strengths of the predictors. These constructs measured 'need for other's approval', 'need for love', 'need for perfection' and autonomy. Dysfunctional attitudes directly relevant to alcohol problems could relate to alcohol consumption directly (Ramsey et al., 2002). Since the measurement of dysfunctional attitude never was developed to measure relevant attitudes for alcohol problems, dysfunctional attitudes at indicator level were explored as direct predictors of alcohol consumption. Both baseline-change and change-follow-up level models were specified. These models controlled for age (Hamilton and Dobson, 2002;Manninen et al., 2006). These are prediction models and no conclusions about causes and effects can be made, as this would require more measurement points to test competing models (Bollen, 1989;Bollen and Curran, 2006).
The Early Treatment of Women with Alcohol Addiction (EWA) treatment project started in 1981 at the Karolinska Hospital in Stockholm, Sweden. This project included 420 women who entered treatment consecutively. The first subsample ( pilot study) was included between 1981 and 1982 (n = 100), the second group (n = 200) was included during 1983-1984 and the third group of 120 women during 1991-1994. The study of the second group compared EWA treatment with treatment as usual. The present study involves the last 120 women. The focus in this part of the EWA project was the study of co-morbidity factors at intake and at 2-year follow-up. This last sub-sample was quite similar to the two first sub-samples with respect to drinking and sociodemographic variables, although some changes were found, reflecting what had taken place in the Swedish society in general during the time-frame of the study. A somewhat higher general level of unemployment, more beer consumption and less spirits were found among the last sub-sample. However, all women preferred to drink wine (Haver et al., 2009). Median age at intake was 44 years (range 23-63). Educational and occupational level, marital status and number of children were representative of the women in the general population. Almost all women (96%) suffered from a DSM-III-R diagnosis of alcohol dependence (Haver et al., 2001), and had an average alcohol consumption ~130 g of alcohol on days with heavy drinking. Only women not previously treated for alcohol abuse were included in the EWA study; however, about half the sample had received psychiatric treatment. Increased levels were found with regard to depression and other mental health dimensions compared with a general population control sample (Haver, 2003). Attrition: of the 120 women enrolled, 98 were studied at the 2-year follow-up (82%). Three women died during the study period (from cancer and alcohol-related causes) and 19 women either refused to participate or could not be located. These 19 women differed statistically from women contributing to the last measurement in baseline measures on some relevant variables, with lower scores on alcohol consumption measured by three different alcohol consumption variables, lower prevalence of psychiatric diagnoses in general and depression diagnoses specifically, but higher scores on failing in goal achievements. Thus, the missingness was probably not at random (Bollen and Curran, 2006), and women with baseline data only were therefore excluded from the analyses (list-wise deletion) instead of imputing follow-up scores.
The psychotherapy focused mainly on alcohol-related problems with a reduction of harmful drinking or total abstinence as the intervention goal; however, pharmacological treatment was also given for alcohol addiction and psychiatric disorders. Further descriptions of the sample, treatment and measurements were reported earlier (Haver, 2003;Haver et al., 2001). The study was approved by the Stockholm Regional Ethical Review Board and by the Swedish Data Inspection Board.
The instruments used were Swedish versions of the Alcohol Use Inventory (AUI; Berglund et al., 1988;Wanberg et al., 1977), a 25-item version of the DAS-25 (Weich et al., 2003) and the Depression scale from the Symptom Checklist-90 (SCL-90; Derogatis and Cleary, 1977;Zack et al., 1998). The instruments were administered twice, at intake after the abstinence phase and at the 2-year follow-up.
The AUI measures used were maximum alcohol consumption on any drinking day in grams of alcohol, average consumption on drinking days and average daily consumption over 1 week. Maximum consumption on any drinking day was used as the main alcohol variable. The two other indicators were also tested in similar models. The variables were normally (or nearly normally) distributed.
The DAS-25 is a valid and frequently used instrument of depression-related cognitions (Nelson, 1992;Oei and Free, 1995;Ramsey et al., 2002). The items are within the domain of approval, love, achievement, perfectionism, entitlement, omnipotence and autonomy. The internal consistency (Cronbach's alpha) for DAS was 0.91 at baseline and 0.90 at follow-up. In explorative bivariate analyses, some follow-up DAS indicators predicted follow-up level in alcohol consumption. These items indicated a need of achievement to gain respect, difficulty taking risks because of fearing the consequences, experiencing inferiority if not succeeding as well as others, afraid of losing others' respect, the need of expecting success before making any effort to solve problems, problems of trusting others because of being afraid to be hurt, believing that happiness is more dependent on other persons than oneself and a belief that problems will disappear without any effort. These items were grouped as a scale [here called alcohol-related DAS (DAS-ALC)] to explore the relations within the model with the follow-up level specified as the intercept factors for DAS-ALC and alcohol consumption.
The SCL-90 Depression scale is a symptom measure of depression. Depression as a theoretical construct is related to observed indicators, both within the diagnostic and dimensional diagnostic perspective (Borsboom, 2008). In this paper, we used the term 'Depression' as the label of the continuous latent underlying variable being reflected by the SCL-90 symptom indicators. Even if these symptom indicators would be related to the diagnosis of depression, this measure is not a measure of clinical diagnosis. As for the alcohol measure, a continuous measure was best suited for the research problems in this study. The internal consistency for this scale was 0.90 and 0.93 at baseline and follow-up, respectively. On the basis of the internal consistency results, the latent depression variable was corrected for measurement error in the variables (Stoolmiller, 1995).
Although it is not possible to determine how much of the change was caused by the intervention, some treatment variables were added in the model. Treatment compliance was indicated by still being in treatment or having terminated in accordance with the treatment staff. About 59% women confirmed this kind of compliance. Other variables were number of visits at the EWA centre during intervention years 1 and 2, number of months participating in treatment and number of days in inpatient treatment.
Our study was based on a one-group pre-post design, which may be described as a longitudinal design (Rogosa, 1995). Problems related to the use of difference score models have been debated, but nevertheless described as relevant for longitudinal analyses (Rogosa et al., 1982). Although two observations do not reveal the nature or shape of change, they provide information about the amount of change (Duncan et al., 2006). Inclusion of a reference group would have enhanced the possibility of using latent growth curve (LGC) modelling to compare the different sources of changes (Muthén and Curran, 1997). However, as this part of the EWA treatment project was the study of psychiatric co-morbidity and intervention effects were studied prior to this phase, decreases and increases in variables may reflect both intervention effects and other causes.
The analyses used were descriptive statistics, reliability (Cronbach's alpha) and t-test for dependent samples testing of group means over time. Structural Equation Modelling (SEM) was used to test confirmatory factor analysis (CFA) models and LGC models of the level and change in the variables.
A CFA of the depression sub-scale did not support the fact that this scale represented one dimension. Eleven residual co-variances had to be estimated to achieve an acceptable fit between model and data, which indicates several dimensions. This probably reflects some heterogeneity in the scale content and an underrepresentation of the depression construct. Owing to low sample size, the DAS could not be tested with CFA.
LGC models take into account not only group mean differences, but also individual differences in change.
LGC is well suited for longitudinal data (Bollen and Curran, 2006), and is preferred over traditional longitudinal analyses (e.g. repeated ANOVA), as it is more flexible regarding complexity in model testing and is the only method that may be used to control for measurement error. On the basis of two measurement points, this model is a difference score model with latent variables (Duncan et al., 2006;Raykov, 1993). Ideally, measurement should be analysed with latent variables reflected by several observed indicators. However, this procedure places demands on the sample size. In this study, the sample size did not allow for such estimation. An approximation of the measurement model could be to pre-specify measurement errors in the models based on Cronbach's alpha (Stoolmiller, 1995). Beyond these restrictions, our analyses follow general guidelines for LGC models within the SEM literature (Bollen and Curran, 2006;Duncan and Duncan, 2004;Willett and Keiley, 2000).
The models were step-wise analysed, separate growth models first, followed by complete models. Based on the hypotheses, the tested structure was tentatively specified. Then non-significant parameters were removed and the model re-estimated. This analytic strategy, which is both confirmatory and exploratory, is described as a modelgenerating procedure (Jöreskog, 1993). The relation between the level and change in DASs, and change in depression, was analysed with the baseline level of depression entered as a predictor and not as a residual covariance. This strategy was used to control for a possible confounding effect between depression and DASs.
The estimation method was Full Information Maximum Likelihood, which uses information from all cases, even if some data are missing (Arbuckle, 2007). These results were compared with Maximum Likelihood on a total imputed sample with Expectation Maximization imputation for missing data. The two estimation methods produced similar results.
We used χ 2 with the significance test, Comparative Fit Index (CFI), Normed Fit Index (NFI), Non-Normed Fit Index (NNFI) and Root Mean Square Error of Approximation (RMSEA) with confidence intervals to evaluate model fit. Ideally, CFI, NFI and NNFI should be beyond 0.90, and RMSEA should be below 0.08 or preferably 0.05 (close fit) (Bollen and Curran, 2006;Kline, 2005). SPSS 15 and AMOS 16 were used for analyses, and Microsoft Excel to compute predicted estimates and generate plots. Plots were estimated at the mean level to describe group baseline and change means, but also at plus/minus one standard deviation to show individual variation around this group level.
Table 1 lists the descriptive statistics obtained from the DAS, SCL-90 Depression scale and AUI Maximum alcohol consumption in grams. All group level changes from baseline (T1) to follow-up (T2) were statistically significant. AUI Maximum alcohol consumption and DAS did not correlate at intake, nor did AUI Maximum and SCL-90 Depression. At follow-up, there was a statistically significant correlation between AUI Maximum alcohol consumption and SCL-90 Depression (r = 0.36, P < 0.05). The correlation between DAS and Depression was stronger at the baseline (r = 0.53, P < 0.05) than at the follow-up (r = 0.36, P < 0.05). Analyses of time-lagged correlations were stronger with DAS (r = 0.54, P < 0.05) and weaker with AUI Maximum alcohol consumption (r = 0.34, P < 0.05).
The results of separate growth analyses were used to estimate the predicted time scores, as presented in Fig. 1. A mean group decrease was found for all three variables even though considerable variation in both the baseline level and the change was found.
The relation between baseline level and change was negative for all variables, indicating that most reduction occurs among those subjects displaying the highest baseline level, while lower baseline levels were associated with a less pronounced decrease and thus increased stability. The depression measure showed more stability than alcohol consumption. The other two alcohol variables showed similar results, with some minor differences regarding the decrease
Table 1. Descriptive statistics (mean, SD) at baseline and at the 2-year follow-up for AUI Maximum consumption on any drinking day, DAS and SCL-90 Depression scale (n = 98) Variable Baseline (T1) Follow-up (T2) Change t-test Mean SD Mean SD AUI Max alcohol cons 150.92 70.86 91.39 70.12 59.53 6.95* DAS 85.94 25.65 79.81 24.47 6.13 2.54* SCL-90 Depression 1.28 0.80 0.90 0.86 0.38 4.18* *P < 0.05. and co-variance between baseline levels and changes. Low DAS levels at baseline were, however, associated with an increase in scores over time.
Figure 2 shows the final model, which reached a close fit between the model and data (χ 2 = 11.37, df = 15, P = 0.73, CFI = 1.00, NFI = 0.93, NNFI = 1.06, RMSEA = 0.00, RMSEA c.i. = 0.00-0.07, RMSEA 1<0.05 = 0.88). Prediction of change in Maximum alcohol consumption was more strongly related to change in SCL-90 Depression than baseline depression level. SCL-90 Depression scores at intake were statistically significantly predicted by DAS scores at intake. The model also showed that the DAS change predicted SCL-90 Depression changes. A plot based on these structural equations is presented in Fig. 3. As the AUI Maximum alcohol consumption and the DAS at the baseline level were predicted by age at treatment intake, this plot represents a scenario for women at the mean age. The model-generated results showed younger or older women to have similar changes over time; they would, however, exhibit higher or lower levels. Age did not predict the baseline level of depression and changes over time for any of these three variables.
The model-generated estimates in Fig. 3 shows that women with an elevated baseline level of SCL-90 depression decreased their maximum alcohol consumption to a lesser extent than women with a low baseline level of depression.
The figure shows two other conditions. Women with lower baseline level of depression and a decrease during the study exhibited the strongest decrease in maximum alcohol consumption. A much less pronounced change in maximum alcohol consumption was related to a higher baseline level of depression and an ongoing high level over time. The figure also shows that the degree of change was dependent on the baseline level of maximum alcohol consumption. This illustrates the importance of building the relation between depression level and change into the model. Dysfunctional attitudes were related to the relation between depression and alcohol consumption, as decrease in the DAS was associated with a greater decrease in depression. Both level and change in dysfunctional attitudes moderated changes in alcohol consumption through the level of and changes in depression. Models using changes in average consumption on drinking days and average daily consumption over 1 week yielded almost identical structural parameter values and goodness-of-fit results.
DAS sub-scales showed differences in parameter values for the prediction of depression, both the relation between baseline levels and the relation between changes in these two factors. Such findings indicate that the scale content is important for prediction strengths. Table 2 shows the 'need of approval' factor to be associated with strongest values, while the achievement scale did not result in a well-fitted model. In this last model, the regression effect of age was not statistically significant.
A model with the empirically derived alcohol-related DAS scale (DAS-ALC) showed that follow-up in DAS-ALC predicted follow-up alcohol consumption (standardized β = 0.21), even after controlling for the relation between change in depression and follow-up alcohol consumption (β = 0.37). As before, change in depression was related to change in alcohol consumption (β = 0.23). In addition, the depression baseline level was related to both changes (-0.44) and the follow-up level in DAS-ALC (0.37). Changes in DAS-ALC was related to changes in depression (β = 0.34). All parameter values were statistically significant and the model fitted the data well (χ 2 = 5.84, df = 7, P = 0.56, CFI = 1.00, NFI = 0.95, NNFI = 1.02, RMSEA = 0.00, RMSEA c.i. = 0.00-0.07, RMSEA 1<0.05 = 0.70). The treatment compliance variable was added to the original model presented in Fig. 2. Compliance was predicted both by the DAS (0.27, P < 0.05) and depression (-0.34, P < 0.05) baseline factors, with higher levels of the DAS and lower levels of depression being associated with the presence of compliance. The direct relation between the baseline symptom level of depression and changes in alcohol consumption was not supported in this model. Also, the compliance variable predicted changes in alcohol consumption (standardized β = -0.20, P < 0.05), that is, greater reduction in alcohol consumption among compliant women. The model fitted the data very well (χ 2 = 14.89, df = 19, P = 0.73, CFI = 1.00, NFI = 0.91, NNFI = 1.06, RMSEA = 0.00, RMSEA c.i. = 0.00-0.07, RMSEA 1<0.05 = 0.89). Other treatment variables did not give similar results.
Many alcohol-addicted women coming for treatment showed symptoms of depression, with related dysfunctional attitudes as well. Changes in alcohol consumption, symptoms of depression and dysfunctional attitudes were found over time, both at group and individual levels. Some of the strong reductions that were associated with high baseline levels could represent a regression to the mean effect (Finney, 2008;Gmel et al., 2007). Even though women with highest problem scores at the baseline also experienced most reduction over time, the results confirmed the findings of Dawson et al. (2007) that higher levels of alcohol consumption at the baseline were associated with higher alcohol consumption later on.
Our findings based on latent growth models have no direct counterparts in previous studies, as these studies have used other types of analyses, different ways of operationalization of the constructs and different time intervals. The number of time points may also be an important factor contributing to differences (Schutte et al., 1997). Nevertheless, the present findings support the hypothesis that both initial and prolonged elevated levels of depression is an important determinant of whether alcohol problems decrease over time or not (Bradizza et al., 2006;Charney et al., 1998;Greenfield et al., 1998;Kodl et al., 2008;Ramsey et al., 2005). The estimated model showed changes in depression symptoms more strongly to predict changes in alcohol consumption than did the baseline level in such depression symptoms. This confirms later alcohol consumption to be more strongly related to changes in depression than previously elevated levels of alcohol consumption and depression (Hartka et al., 1991). Thus, a prolonged level of depression may explain the apparent stability or even the increase in alcohol consumption and other alcohol-related problems over time (Driessen et al., 2001). These results may also suggest why alcohol abstinence at a later point in time is not necessarily predicted by baseline alcohol consumption and severity of dependence (Kodl et al., 2008). Even if no differences with regard to decrease in alcohol consumption between groups with or without baseline depression or anxiety were found (Burns et al., 2005), our results could indicate that this could be due to a decrease in depressive symptoms among subjects in the co-morbid group. Our results are also in agreement with a study showing that depressed patients are at risk for increased alcohol consumption over time (Holahan et al., 2004). These patients were at risk for increased negative events, less social support and greater levels of dysfunctional coping. Obtaining help for coping with depression and otherrelated factors may, therefore, also enhance the therapeutic effects on alcohol problems, for example, a reduction in harmful alcohol consumption. Our study showed that higher levels of alcohol consumption and depression at intake were associated with a greater decrease in alcohol consumption, given the expected changes in depression. This shows how changes in the outcome variable are moderated by the predictors, which indicates interaction effects (Bollen and Curran, 2006). The present investigation therefore extends earlier findings, as both the level of and change in depression contributed to the magnitude of change in alcohol consumption. These findings illustrate some of the shortcomings of traditional analyses of change, especially when using categorized variables. In sum, these findings underscore the importance of reducing both depression level and alcohol consumption among women who present with this co-morbidity.
Changes in dysfunctional attitudes were evident, even though these were more stable than the levels of depression and alcohol consumption. These findings were as expected (Elkin et al., 2006;Furlong and Oei, 2002;Hamilton and Dobson, 2002;Oei and Free, 1995). Earlier studies indicate that pretreatment dysfunctional attitudes are predictive of depression treatment response and depression relapse (Elkin et al., 2006;Hamilton and Dobson, 2002). In addition to the different analytic approaches used, the inclusion of the DAS change factor in addition to the baseline level could explain this difference in findings. Nevertheless, the baseline dysfunctional attitudes score is related to changes in depression indirectly via its relation to the depression baseline level. The growth curve analysis without a specified measurement error showed a statistically significant relation between the baseline level of dysfunctional attitudes and changes in depression. However, when measurement errors in the variables were accounted for, this direct relation was not supported. Reduction in dysfunctional attitudes was related to reduction in depression levels, a finding confirming previous findings (Beevers and Miller, 2004;Brown and Ramsey, 2000;Furlong and Oei, 2002;Kwon and Oei, 2003). Reduction in alcohol consumption was indirectly related to a lower baseline level of and reduction over time in dysfunctional attitudes. The DAS as a total scale did not predict alcohol consumption directly. However, some DAS items at follow-up predicted alcohol consumption at follow-up. This confirms that cognitions may be directly related to alcohol problems if relevant cognitions are measured (Heinz et al., 2009;Ramsey et al., 2002). These findings show how a possible direct relation between a DAS and alcohol consumption may not be found if this scale consists of heterogeneous items only partially relevant to drinking. Differences in the magnitude of the predictor relations were found to be dependent on which DAS sub-scales were analysed. This illustrates the importance of using conceptually homogeneous scales in analyses (Smith et al., 2009). Preliminary analyses also link increases in dysfunctional attitudes to mortality in this group of women, while no such results were found for changes in depression and alcohol consumption. Integrating different dimensions of dysfunctional attitudes into the model gives additional understanding of this co-morbidity between alcohol addiction and depression over time.
Compliance was predicted by baseline levels in dysfunctional attitudes and depression, and it predicted changes in alcohol consumption. The direct relation between the baseline level in depression and changes in alcohol consumption was not maintained when compliance was included in the model. However, this relation may be partially covered by the indirect relations between baseline depression and changes in alcohol consumption through the compliance variable. The relation between baseline dysfunctional attitudes and compliance was positive, whereas the relation with depression at the baseline was negative. Even if this finding is not easily interpreted, it is an interesting possibility that such dysfunctional attitudes may be related to the perceived need and motivation for treatment, while depression symptoms not covered by such attitudes might be associated with decreased motivation and passivity.
The generalizability of our findings is limited to a population of women being able to contribute at both measurements points, since missing data analysis indicated differences between those lost to follow-up and the rest of the group studied. Another sample limitation is linked to the model estimation with this sample size. Even if the results indicated close fit, some instability was also found. The model should therefore be replicated in samples of greater size. Although the Depression scale of SCL-90 has been widely used, its validity has been questioned (Cyr et al., 1985;Schmitz et al., 2000;Zack et al., 1998). Our analyses confirm that this scale is encumbered with some validity problems, which indicate that SCL-90 should be revised to achieve more homogeneous, distinct and meaningful dimensions. However, this measurement problem perhaps applies to diagnoses in general (Smith et al., 2009).
This growth curve analysis of the co-morbidity of level and changes in alcohol consumption and depression describes a clinical picture that both confirms and expands earlier findings, suggesting that both the baseline level of and changes in depression are related to changes in alcohol consumption over time. The inclusion of dysfunctional attitudes moderates the relation between depression and alcohol consumption; however, to different degrees depending on the individual dysfunctional attitude items included for the analysis. In addition, a DAS sub-scale directly related to alcohol consumption at follow-up. Lastly, intervention compliance was found as a mediating factor between the baseline level in these factors and changes in alcohol consumption. These findings have implications for treatment of co-morbid depression and alcohol addiction.
Acknowledgements -The assistance of
Funding -The project was funded by the
Background: Clustering is a widely used technique for analysis of gene expression data. Most clustering methods group genes based on the distances, while few methods group genes according to the similarities of the distributions of the gene expression levels. Furthermore, as the biological annotation resources accumulated, an increasing number of genes have been annotated into functional categories. As a result, evaluating the performance of clustering methods in terms of the functional consistency of the resulting clusters is of great interest. Results: In this paper, we proposed the WDCM (Weibull Distribution-based Clustering Method), a robust approach for clustering gene expression data, in which the gene expressions of individual genes are considered as the random variables following unique Weibull distributions. Our WDCM is based on the concept that the genes with similar expression profiles have similar distribution parameters, and thus the genes are clustered via the Weibull distribution parameters. We used the WDCM to cluster three cancer gene expression data sets from the lung cancer, B-cell follicular lymphoma and bladder carcinoma and obtained well-clustered results. We compared the performance of WDCM with k-means and Self Organizing Map (SOM) using functional annotation information given by the Gene Ontology (GO). The results showed that the functional annotation ratios of WDCM are higher than those of the other methods. We also utilized the external measure Adjusted Rand Index to validate the performance of the WDCM. The comparative results demonstrate that the WDCM provides the better clustering performance compared to k-means and SOM algorithms. The merit of the proposed WDCM is that it can be applied to cluster incomplete gene expression data without imputing the missing values. Moreover, the robustness of WDCM is also evaluated on the incomplete data sets.
The results demonstrate that our WDCM produces clusters with more consistent functional annotations than the other methods. The WDCM is also verified to be robust and is capable of clustering gene expression data containing a small quantity of missing values.
The changes of the gene expression levels are very common in the human complex diseases, such as cancers [1][2][3]. The advent of microarray technologies have made it possible to measure simultaneously the expression levels of many thousands of genes over different time points and/or under different experimental conditions [4][5][6]. Numerous computational techniques have been developed to analyze these gene expression data. Among them, clustering is a primary approach to group the genes with similar expression patterns across different conditions, which enables the identification of differentially expressed gene sets in cancerous tissues [7][8][9]. Clustering is an unsupervised learning technique which assigns a set of objects (genes) into subsets (called clusters) so that the objects in the same clusters are similar according to some similarity metric [10,11]. A cluster is therefore a collection of objects which are similar between them and are dissimilar to the objects belonging to other clusters.
Since clustering is proposed, an increasing number of clustering approaches have been developed and improved for the analyses of gene expression data. The common clustering methods include k-means [12,13], hierarchical clustering [8], and Self Organizing Map (SOM) [14,15], and so on. Each method has its own strengths and weaknesses. The k-means is an important clustering algorithm which partitions n objects into k clusters in which each object belongs to the cluster with the nearest mean. In k-means clustering, the number of clusters k is an input parameter, and an inappropriate choice of k may yield poor clustering results. The main advantages of this algorithm are its simplicity and computational speed which allows it to run on large datasets, however, it does not yield the same result with each run, since the resulting clusters depend on the initial random assignments. Besides, it conducts poorly with overlapping clusters and is sensitive for noisy data. The hierarchical clustering aims to create a hierarchy of clusters which may be represented by a tree structure called a dendrogram. The root of the tree consists of a single cluster containing all objects, and the leaves correspond to individual objects. The hierarchical technique requires relatively smooth data and the clusters themselves need to be well defined. Like k-means method, noisy data strongly affect the resulting clusters. SOM is a type of artificial neural network that is trained using unsupervised learning to produce a two-dimensional, discretized representation of the input space of observations. It requires the geometry of nodes as input, and the nodes are mapped into two-dimensional space, initially at random, and then iteratively adjusted. SOM imposes the structure on data, with neighboring nodes tending to define related clusters. SOM has good computational properties and is suited to clustering of large data sets. One major drawback of this algorithm is the "boundary effect" of nodes on the edges of the network, which may lead to less effective clustering results. Besides, these clustering methods mentioned above require a complete data set as an input, and therefore those gene rows containing the missing values are either removed or imputed using an imputation method on the missing entries prior to clustering analysis. Removing the missing gene rows may result in omitting some important genes, such as the genes related to diseases, whereas the badly estimated missing values even changes the quality of data, which could influence the accuracy of clustering results.
In this article, we propose a Weibull distributionbased clustering method called WDCM. The assumption of this method is that the gene expression of each gene can be considered as a random variable following unique Weibull distribution [16], and that a group of genes tend to be clustered together if the Weibull distributions of gene expressions of these genes have similar distribution parameters. Here, we use the gene expression values of each gene to construct its corresponding Weibull distribution and then group these genes by clustering their corresponding distribution parameters.
The following sections of this paper are organized as 'Results', 'Discussion and conclusion' and 'Methods'. In section 'Results', we first introduced three cancer gene expression data sets we used, and then visually demonstrated the clustering results obtained using the WDCM for the three data sets. Second, to assess the performance of the WDCM, we compared the functional consistency of the gene clusters produced by the WDCM to those of the k-means and SOM methods for the same data sets. We also used the external measure Adjusted Rand Index to establish the performance of the WDCM, and the comparisons with the other algorithms were conducted simultaneously. Finally, we tested the robustness of the WDCM on clustering the incomplete data sets. In section 'Discussion and conclusion', we first summarized the main work of this study, discussed the strength and limitation of the WDCM. In the end we briefly mentioned the improvement of the WDCM and the future study. In section 'Methods', we introduced the WDCM together with the algorithm used for clustering the Weibull distribution parameters, the functional consistency assessment method of the clustering result, and the external validation index Adjusted Rand Index of the clustering performance. Moreover, Robustness test of the WDCM on clustering the incomplete data set was also presented in this section.
In this section, the WDCM is described as follows: Given a m × n gene expression matrix, let g ij be the jth expression value of gene i, i = 1, ...,m, and j = 1, ...,n. We here treat one gene expression as a random variable, and construct the distribution of the gene expressions of gene i. We then choose a subset of genes whose distributions of the gene expressions belong to the common Weibull distribution [16]. Due to the consistent distribution function types, we consider that those genes with similar gene expression distribution parameters tend to share the similar expression patterns, and they are probably concerned with the same biological processes or functions together. We further cluster the genes in the selected subset by clustering their corresponding distribution parameters, as each gene corresponds to its unique distribution parameters. In the following we introduce the principle of the distribution function construction procedures.
First, we construct the empirical distribution of each gene expression [17], and then ascertain the precise distribution regarding the constructed empirical distribution using the Kolmogorov goodness of fit test [18][19][20].
The details as follows: assume that x i1 , x i2 , ..., x in are the gene expressions of gene gi, i = 1, ...,m, and sort them in ascending as x i1 < x i2 < • • • x in . For ∀ x (-∞,+∞), define the empirical distribution of g i as
Where I(•) is the indicator function.
We utilize the Weibull distribution type to fit F (i) n (x) , and then ascertain the distribution parameters which uniquely determine the distribution.
The probability density function of a Weibull distribution is defined as:
where a >0 is the scale parameter and b >0 is the shape parameter of the distribution. The scale parameter a determines the range of the distribution. The shape parameter b is what gives the Weibull distribution its flexibility. By changing the value of the shape parameter, the Weibull distribution can fit a wide variety of data.
Let F (i) (x) is a certain Weibull distribution with known parameters, and a Kolmogorov-Smirnov test is conducted to determine if the sample x i1 ,x i2 , ..., x in comes from the Weibull distribution F (i) (x). The null hypothesis is that the random sample of gene expressions of g i comes from the Weibull distribution F (i) (x). If the null hypothesis is true, the deviation of F (i) (x) and F (i) (x) is small. Construct the Kolmogorov-Smironov statistic
under the null hypothesis, √ nT (i) n converges to the Kolmogorov distribution [18]. The null hypothesis is rejected at significance level a if √ nT (i) n > K α , otherwise it is accepted, where K a is the critical value of the Kolmogorov distribution. Given a = 0.05, we here select the appropriate parameters for F (i) (x) in order to the null hypothesis is accepted (p -value > 0.05), that is, the random sample comes from the certain Weibull distribution F (i) (x), i = 1,2, ...,m. Following the above procedure, we can obtain the Weibull distributions of m gene expressions, denoted by F (1) (x),F (2) (x),...,F(m)(x).
Let θ i denotes the parameter of the Weibull distribution F (i) (x), j = 1, ...,m. Here θ i consists of double-parameter pair (a i ,b i ), we then cluster the m parameters θ 1 , θ 2 ,..., θ m using a certain clustering algorithm based on the hub points. This algorithm presented by Robert Clason designates a single point as a hub for each cluster and then finds the distance from each remaining point to each hub, as well as assigns this point to the hub to which it is closer [21]. The merit of it is to automatically ascertain the clusters number on the basis of the distances between data points. A detailed description of the algorithm is provided in Additional file 1.
In order to evaluate the performance of the proposed WDCM, we also apply the K-means and Self Organizing Map (SOM) clustering algorithms to the same gene subsets as the WDCM and obtain the gene clusters, respectively. We compare the functional consistency of the gene clusters produced by WDCM to those produced by the other methods. For this purpose, we consider the biological annotations of the gene clusters in terms of Gene Ontology (GO). The Gene Ontology (GO) project provides three structured, controlled vocabularies that describe the gene products in terms of their associated biological processes (BP), cellular components (CC) and molecular functions (MF) [22]. The annotation ratios of each gene cluster in three GO terms were calculated using the web-accessible DAVID 2008 tool [23]. For each of clusters found by one of three clustering methods, under the BP ontology, we search the just GO term in which the most genes in this cluster are enriched, and define the BP annotation ratio for this cluster as the number of genes in both the assigned GO term and this cluster divided by the number of genes in this cluster. After calculating the BP annotation ratios for all clusters, we treat the mean value of all annotation ratios as the final BP annotation ratio. We also define the CC and MF annotation ratios by the same manner. A higher annotation ratio represents that the corresponding clustering result is better than the other ones, that is, gene are better clustered by function, indicating a more functionally consistent clustering result.
The Adjusted Rand Index (ARI) is a measure of agreement between two partitions of the same set of objects [24,25]. One partition is given by the clustering method and the other is defined by the external criteria. For a gene expression data set, suppose X is the partition based on some external criteria and C is the clustering result obtained by some clustering method. Let a,b,c,d respectively denote the number of gene pairs that are in the same cluster in both X and C, the number of gene pairs that are in the same cluster in X and in different clusters in C, the number of gene pairs that are in different clusters in X and in the same cluster in C and the number of gene pairs that are in different clusters in both X and C. The Adjusted Rand Index ARI(X,C) is defined as follows:
The value of Adjusted Rand Index varies from 0 to 1 and higher value means that C is more similar to X.
Considering that the genes with similar expression patterns may be functionally related each other [26], we group the genes in the given data set according to functional similarity and define these gene clusters as X. The clustering results Cs are then given by the proposed WDCM, k-means and SOM. We compute and compare the values of Adjusted Rand Index between X and Cs to evaluate the performance of WDCM. To this end, we first use the Gene Functional Classification Tool of DAVID to group the genes into the highly functionally related gene clusters and then compute the values of ARI. The higher value indicates the corresponding clustering method performs better.
The WDCM can be applied to cluster the incomplete gene expression data set without imputing the missing values. To test the robustness of this approach, we compared the overlapped degree between the gene clusters for incomplete data sets and the ones for complete data sets. A higher overlapped degree represents a robust clustering method. To this end, we first randomly remove 5-25% of the complete data set in order to create the incomplete gene expression data sets, and then we apply the WDCM to cluster these complete and incomplete data sets and obtain the clustering results, respectively. Here, a Cluster Overlap Ratio (COR) index is introduced for assessing the overlapped degrees at individual missing percentages.
Suppose n gene clusters C 1 ,C 2 ,...,C n for the complete data set and m gene clusters I 1 ,I 2 , ... I m for the incomplete one. The Cluster Overlap Ratio (COR) index is then defined as follows:
where
We applied the WDCM to cluster the lung cancer data set. It consists of expression levels of 675 genes across 156 tissues, which include 17 normal and 139 carcinomas lung tissues [27]. Using the Kolmogorov-Smirnov goodness of fit test (see Methods), we tested whether the expression sample of each gene comes from the Weibull distribution. The results showed that the distributions of gene expressions of 402 genes belong to the common Weibull distribution, whereas the others whose distributions of gene expressions fail to be in the Weibull distribution are removed. The p-values produced by Kolmogoriv-Smirnov goodness of fit test for the 402 genes were reported in Additional file 2. We then used the hub node based clustering algorithm (see Methods) to cluster the 402 Weibull distribution parameters which consist of the shape parameters and scale parameters, and obtained 6 distribution parameter clusters, that is, 6 gene clusters. The clustered parameters scatter plots have been shown in Figure 1A. It is evident from Figure 1A that the distribution parameters of the genes of a cluster are close and compact to each other, which indicates the Weibull distribution parameters were clustered well. The expression profiles of the corresponding clustered genes plots have been shown in Figure 1B, from which it is also evident that the expression profiles of the genes within identical clusters are quite similar, whereas the profiles for the genes belonging to different clusters differ from each other.
We tested the WDCM on another follicular lymphoma data set consisting of expression levels of 798 genes in 19 B-cell follicular lymphoma specimens [28]. We utilized the Kolmogorov-Smirnov test to decide if the sample of individual gene on the follicular lymphoma data set comes from the Weibull distribution, and found 471 genes whose distributions of gene expressions belong to the common Weibull distribution. The p-values produced by Kolmogoriv-Smirnov goodness of fit test for the 471 genes were reported in Additional file 2. We then clustered the corresponding 471 distribution parameter pairs and determined 4 gene clusters. Figure 2 illustrates the clustered parameters scatter plots and the cluster profile plots of the clustering results.
From Figure 2A, the four parameters clusters are clearly distinguished from each other, meanwhile, the expression profiles of the genes within the same clusters are similar, whereas the ones of the genes across different clusters are distinct (see Figure 2B). The results indicate that the significantly distinct gene clusters were found using the WDCM on follicular lymphoma data set.
The bladder carcinoma data set contains 1203 genes measured over 40 different experimental conditions [29]. Using the Kolmogorov-Smirnov test, we found 1040 genes whose distributions of gene expressions belong to the common Weibull distribution. The p-values produced by Kolmogoriv-Smirnov goodness of fit test for the 1040 genes were reported in Additional file 2. Again, the hub node based clustering algorithm was employed to cluster the corresponding 1040 distribution parameter pairs. The number of clusters determined was 4. Figure 3 shows the clustered parameters scatter plots and the cluster profile plots of the clustering results.
To show the performance of the WDCM, we applied the K-means and Self Organizing Map (SOM) algorithms to the same gene subsets clustered by the WDCM and compared the functional consistency of the gene clusters produced by WDCM to those of the gene clusters produced by the other methods (see Methods). Simultaneously, the values of ARI for the WDCM, kmeans and SOM algorithms on these three data sets were also contrasted (see Methods).
Among these three tested algorithms, the WDCM show the highest functional annotation ratios on both lung cancer and follicular lymphoma data sets. The detailed comparisons for the lung cancer data set are given in Figure 4A, from which we found that the three final functional annotation ratios of the WDCM clusters all exceed the ones of the other methods clusters. Especially, the BP and MF annotation ratios of the WDCM clusters (91.57% and 92.16%) are much higher than those of the SOM clusters (82.76% and 83.96%). On Bcell follicular lymphoma data test, although the CC and MF annotation ratios of gene clusters found by each of three methods are asymptotically equal (see Figure 4B), the BP annotation ratio of WDCM clusters (84.9%) is much higher than those of K-means clusters (71.6%) and SOM clusters (74.8%). On bladder carcinoma data set, from Figure 4C, although the BP annotation ratio of WDCM clusters (59.82%) is less than those of SOM clusters (64.30%), it is still beyond that of K-means clusters (55.87%). Note that the CC and MF annotation ratios of the WDCM clusters are consistently superior to those of the K-means and SOM clusters.
Table 1 shows the values of ARI for algorithms WDCM, k-means and SOM on these three data sets. Note that among the three methods, WDCM provides the consistently best ARI values. Specifically, the ARI value for the proposed WDCM (0.5365) is much better than those for k-means and SOM (0.2478 and 0.3681) on lung cancer data set. Although these three ARI values (0.3991, 0.3481 and 0.2647) are close on B-cell follicular lymphoma data set, the ARI value for WDCM is better than the other values. For bladder carcinoma data set also, the proposed WDCM outperforms the other algorithms in terms of ARI. The values are reported in Table 1.
The above comparative analyses on the functional annotation ratios of the three algorithms have demonstrated that the genes in each cluster obtained using the WDCM show not only the similar expression patterns, but also more consistent functional annotations, which means these genes are more inclined to be involved in the same biological functions together. Also, the Adjusted Rand Index comparative results indicate the superiority of the performance of the proposed WDCM compared to the other algorithms.
To test the robustness with which the WDCM clusters the incomplete gene expression data, we applied the WDCM to cluster the above three gene expression data sets containing missing values and compared the overlapped degree between the gene clusters for incomplete data sets and the ones for complete data sets. These three data sets were preprocessed by randomly removing 5-25% of the data in order to create the incomplete gene expression data sets, and the WDCM then was applied to these data sets. Table 2 lists the average Cluster Overlap Ratio (COR) values with respect to the percentages of missing values (0-25%) achieved by WDCM over 100 runs for the lung cancer, B-cell follicular lymphoma and bladder carcinoma data sets, respectively. The WDCM provided the higher COR values regarding the smaller percentages of missing values for all three data sets. The COR values exceeded 0.9 at 5% missing value. At 10%, the COR value was also beyond 0.9 for the follicular lymphoma and bladder carcinoma data sets (0.9078 and 0.9702), and approximated 0.9 for the lung cancer data set (0.8654). For the bladder carcinoma data set, we see that the COR values were varied from 0.9823 to 0.9335, passing 0.9 at all missing values.
The results of the cluster overlapped degree comparison tests indicate that the WDCM gave a high overlapped degree of the gene clusters compared with those of complete data set at low missing value, highlighting the robustness and potential of the WDCM. We think that the results might stem from the fact that the missing gene expression values of individual genes have little influence on constructing their corresponding Weibull distribution parameters at low missing values.
In this article, we propose a robust approach based on Weibull distribution (WDCM) for clustering gene expression data. It is based on the idea that a group of genes tend to be clustered together if the distributions of gene expressions of these genes belong to the common Weibull distribution and have the similar distribution parameters. Consequently, we cluster the genes by clustering the distribution parameters of their gene expressions. A hub nodes-based dynamic clustering algorithm is utilized in the distributions clustering process. The clusters number in a gene expression data set is automatically determined in this clustering algorithm. The performance of the proposed WDCM has been compared with those of K-means and SOM clustering algorithms by the biological annotation ratios to show its effectiveness on three cancer gene expression data sets. The results show that the WDCM is more capable of grouping the genes with similar expression patterns and strong functional consistency together. We also used the external measure Adjusted Rand Index to validate the performance of the WDCM. The comparative results demonstrate that the WDCM provides the better clustering performance compared to k-means and SOM algorithms. Moreover, the WDCM can be applied to cluster the incomplete gene expression data set without imputing the missing values. The results have demonstrated that there is high overlap between the gene clusters for the incomplete data set and those for the complete data set, which illustrates the robustness of the WDCM on clustering the incomplete data set at low percentage of missing values. In general it is known that due to the complex nature of the gene expression data sets themselves and the experimental errors in detecting the gene expression data, it is difficult to discover an acknowledged best clustering approach. In clustering process, the WDCM disregards a few genes whose gene expression distributions fail to fit the Weibull distribution. In future study, we will consider replacing the single Weibull distribution with the mixture distribution in order to cluster the whole data set. Besides, we will also increase the robustness of this approach on clustering the incomplete gene expression data set containing the missing values of moderate percentage. For the gene clusters found by WDCM, we would like to investigate which gene clusters and genes are correlated with some cancer phenotype, and which biological processes or molecular functions these genes in the clusters are concerned with. Our study may be helpful to gain insights into the complex diseases.
Acknowledgements This work was supported in part by the
The authors declare that they have no competing interests.
Authors' contributions HKW and ZZW jointly proposed this approach and conducted the data experiments. XL gave the statistical idea of the method. BSG modified this paper. LXF partly wrote the program codes. Testing was done by YZ. All authors read and approved the final manuscript.
Additional file 1: A clustering algorithm based on "hub nodes". A clustering algorithm used to cluster the Weibull distribution parameters.
Additional file 2: P-values of tests for the three data sets. This file consists of three spreadsheets, each lists the gene numbers and p-values of Kolmogorov Smirnov test for one data set.
Non-alcoholic fatty liver disease (NAFLD) has become the most prevalent cause of liver disease in Western countries. The development of nonalcoholic steatohepatitis (NASH) and fibrosis identifies an at-risk group with increased risk of cardiovascular and liver-related deaths. The identification and management of this at-risk group remains a clinical challenge.
To perform a systematic review of the established and emerging strategies for the diagnosis and staging of NAFLD.
There has been a substantial development of non-invasive risk scores, biomarker panels and radiological modalities to identify at-risk patients with NAFLD without recourse to liver biopsy on a routine basis. These modalities and algorithms have improved significantly in their diagnosis and staging of fibrosis and NASH in patients with NAFLD, and will likely impact on the number of patients undergoing liver biopsy.
Staging for NAFLD can now be performed by a combination of radiological and laboratory techniques, greatly reducing the requirement for invasive liver biopsy.
Non-alcoholic fatty liver disease (NAFLD) encompasses a spectrum of disease ranging from simple steatosis, to inflammatory steatohepatitis (NASH) with increasing levels of fibrosis and ultimately cirrhosis. NAFLD is closely associated with obesity and insulin resistance, and is now recognised to represent the hepatic manifestation of the metabolic syndrome. Since the term NASH was first coined by Ludwig et al. in 1980, 1 the prevalence of NAFLD has risen rapidly in parallel with the dramatic rise in population levels of obesity and diabetes, 2 resulting in NAFLD now representing the most common cause of liver disease in the Western world. 3 Despite recent advances in elucidating the complex metabolic and inflammatory pathways involved in NAFLD, the pathogenesis of steatosis and progression to steatohepatitis and fibrosis ⁄ cirrhosis is not yet fully understood. 4,5 While steatosis alone appears to be associated with a relatively benign prognosis, 6 factors known to be involved in progression to more advanced and clinically relevant disease include inflammatory cytokines ⁄ adipokines, mitochondrial dysfunction and oxidative stress. 7 Insulin resistance causes impaired suppression of adipose tissue lipolysis, leading to increased efflux of free fatty acids (FFA) from adipose tissue to the liver. 8 Hyperinsulinaemia also promotes hepatic de novo lipogenesis, which is markedly increased in NAFLD patients compared with normal individuals. 9 It is now recognised that FFA promote insulin resistance, inflammation and oxidative stress, 10,11 and thus rather than being harmful, hepatic triglyceride accumulation may actually be protective by preventing the harmful effects of FFA. 12 The important role of oxidative stress mechanisms, pro-inflammatory cytokines such as TNFalpha and interleukin 6, and adipokines such as leptin (proinflammatory and pro-fibrotic), and adiponectin (anti-inflammatory and insulin-sensitising), in promoting NASH are also becoming increasingly delineated. 5 However, evidence that only a minority of patients with NAFLD progress to more advanced stages of NASH suggests that disease progression is likely to depend on a complex interplay between such factors and underlying genetic predisposition. 4,7 The causes, epidemiology and natural history of NA-FLD will be covered briefly, before discussing the established and emerging means of assessing and staging patients with NAFLD.
In the great majority of cases, NAFLD arises in association with one or more features of the metabolic syn-drome, namely insulin resistance, glucose intolerance or diabetes, central obesity, dyslipidaemia and hypertension. [13][14][15] However, after exclusion of a history of significant alcohol intake, which is conventionally <20 g ⁄ day, 16 other causes of steatosis which should be considered include nutritional causes, e.g. rapid weight loss and total parenteral nutrition, rare metabolic disorders and druginduced steatosis. Commonly implicated agents include glucocorticoids, amiodarone, synthetic oestrogens and highly active antiretroviral drugs (HAART). [16][17][18] Steatosis is also frequently associated with hepatitis C, particularly genotype 3, and endocrine disorders such as polycystic ovary syndrome (PCOS), 19,20 hypopituitarism 21 and hypothyroidism. 22
The prevalence of NAFLD is estimated to be between 20% and 30% in Western adults, 23,24 rising to 90% in the morbidly obese. 25 NASH, the more advanced and clinically important form of NAFLD, is less common, with an estimated prevalence of 2-3% in the general population 16 and 37% in the morbidly obese. 25 Of concern, NAFLD now affects 3% of the general paediatric population, rising to 53% in obese children, 26,27 with considerable implications for future disease burden. Steatosis was present in 70% of a large unselected cohort of patients with type 2 diabetes. 28 Non-alcoholic fatty liver disease affects all ethnic groups, although prevalence appears to be higher in Hispanic and European Americans compared with African-Americans. This difference remains after controlling for insulin resistance and obesity 23,29 and may be related to ethnic differences in lipid metabolism. 23,30
Patients with a diagnosis of NAFLD have been shown across several studies to have a worse outcome when compared with an age and sex-matched general population. 31 Of note, the excess mortality in this group is attributable to both cardiovascular and liver-related causes. 32,33 Since the description in 1999 of the prognostic relevance of different histological types of NAFLD, 34 several subsequent studies have demonstrated that the presence of just simple steatosis, with no inflammation or fibrosis, is associated with a similar overall and liverrelated mortality to that of an age and gender matched general population. This reinforces the need to stratify patients with NAFLD into simple steatosis or more advanced disease. More advanced disease can be defined as advancing levels of fibrosis and ⁄ or the presence ⁄ level of inflammation and hepatocyte ballooning. This distinction is pertinent as cohort studies thus far have only identified advanced fibrosis, and not inflammation, as a predictor of worse clinical outcome. 32 This may be a type 2 error reflecting small sample sizes, or it may be attributed to additional factors such as PNPLA3 polymorphisms 35 regulating the development of fibrosis.
A systematic literature search was performed to identify studies assessing methods for the diagnosis and staging of NAFLD ⁄ NASH. Relevant articles were identified by searching the PubMed database, MEDLINE and EMBASE, limited to articles published in the English language but not date-restricted. Search terms included fatty liver, NAFLD, NASH, steatosis, AND biomarkers, non-invasive, diagnosis, assessment, staging. Additional searches were also made for each of the individual methods described, e.g. NAFLD fibrosis score, transient elastography, Fibroscan, Fibrotest etc. Selected articles referenced in these publications were also examined.
Studies were included if:
(i) they were meta-analyses, systematic reviews or primary studies of one or more relevant diagnostic ⁄ staging tool;
(ii) they included at least 30 subjects, to reduce the risk of including underpowered studies;
(iii) liver biopsy was used as the reference standard;
(iv) the diagnosis of NAFLD had been established with exclusion of other causes of liver disease.
Studies were excluded if:
(i) publications were not in English;
(ii) data on disease stage e.g. fibrosis stage, was not identifiable;
(iii) they were only presented in abstract form. Using the search strategy described above, approximately 150 articles were considered. Following review, 68 articles met the selection criteria and were included in the analyses.
JD performed the data extraction, which was then checked by the remaining authors (PN and JT).
The diagnosis of NAFLD should be strongly suspected in the presence of features such as obesity, diabetes and obstructive sleep apnoea (OSA); however, other causes should always be considered before attributing abnormal liver function tests (LFTs) to NAFLD alone (Figure 1). Alternative diagnoses which should be excluded by history and serological testing include the viral hepatitides, excess alcohol consumption, haemochromatosis, autoimmune liver disease, alpha-1 antitrypsin deficiency, Wilson's disease and drug-induced liver dysfunction.
The majority of patients with NAFLD are asymptomatic and the diagnosis suspected after finding elevated transaminases on routine testing. Hepatic steatosis is also a frequent incidental finding on ultrasound scan (US) performed for other reasons such as suspected gallstone
Repeat LFTs to check if still abnormal ALT men >30 women >19
Liver screen and USS Hep B and C serology,ferritin/transferrin saturation, liver autoantibodies (AMA, ASMA, ANA), alpha-1-anti trypsin levels, ceruloplasmin if <40.
Liver screen negative Liver screen negative Manage as per condition Normal USS Echobright liver on USS NAFLD/NASH ? Cause ? Mild Steatosis disease. The most common symptoms are right upper quadrant discomfort and fatigue, although the latter may also be caused by OSA which is frequently observed in the typically obese population with NAFLD. Hepatomegaly is the most common clinical finding, with signs of chronic liver disease rarely present in the absence of cirrhosis. A recent study reported the novel finding that increased dorsocervical lipohypertrophy was the anthropometric parameter most strongly associated with severity of steatohepatitis. 36 Although NAFLD is often diagnosed after the finding of mildly abnormal LFTs, more than two thirds of patients have normal aminotransferase levels at any given time 37 and the entire histological spectrum of NAFLD can be observed in patients with normal alanine aminotransferase (ALT) values. 38,39 ALT is usually greater than aspartate aminotransferase (AST), and rarely more than three times the upper limit of normal. An AST:ALT ratio greater than 1.0 suggests the presence of more advanced disease. 40 Alkaline phosphatase can be slightly elevated but is rarely the only liver function test abnormality. 41 Gamma-glutamyltransferase (GGT) is frequently elevated and may also be a marker of increased mortality. 42,43 Low albumin and hyperbilirubinaemia indicate advanced liver disease and are not otherwise features of NAFLD. 44 Iron studies may show an elevated ferritin in up to 50% of patients and elevated transferrin saturation in approximately 10%. 40 However, such findings do not appear to correlate with elevated hepatic iron concentration, and the role of hepatic iron in the pathogenesis of NASH remains unclear. 45 The Fatty Liver Index (FLI) was developed as a simple algorithm to predict fatty liver on USS in the general population. 46 The FLI uses four variables of BMI, waist circumference, GGT and serum triglyceride levels, and achieved an accuracy of 0.84 in detecting fatty liver. 46 The FLI has since been utilised by several groups in population studies of NAFLD. [47][48][49] Ultrasound (USS) is a commonly used test in patients with suspected NAFLD, with steatosis typically appearing as a hyperechogenic liver. A recent study examined the accuracy of USS in 235 patients with suspected liver disease who underwent liver biopsy, and showed a sensitivity of 64% and specificity of 97%, rising to 91% and 93% respectively in patients with at least 30% steatosis. 50 However, the presence of morbid obesity considerably reduces sensitivity and specificity. 51 USS is unable to quantify the amount of fat present or provide any staging of disease, 52 and is operator-dependent with significant intra-and inter-observer variability. 53
Having made a diagnosis of NAFLD, the next step is to determine the severity, as that provides important information on prognosis. Historically this has required liver biopsy, although there have been many recent advances which allow non-invasive management for many patients. When staging patients with NAFLD, there are two aspects to consider; (i) the level of fibrosis and (ii) the level of inflammation ⁄ ballooning (Table 1).
The histological spectrum of NAFLD ranges from simple steatosis through steatohepatitis to fibrosis and cirrhosis. There are no pathological changes which can definitively distinguish NAFLD from alcoholic liver disease (ALD), thus an accurate alcohol history is essential to distinguish between these two common conditions. 54 The histological changes in NAFLD are mainly parenchymal and in a perivenular location, although portal and periportal lesions may occur. 54 Simple steatosis is usually macrovesicular resulting from accumulation of triglycerides within hepatocytes. 44 Features of steatohepatitis include hepatocellular injury, characterised by ballooned hepatocytes, with inflammation and fibrosis. 54 Mitochondrial abnormalities may occur in NASH, but rarely in simple steatosis, 11 supporting a role for mitochondrial defects in the pathogenesis of NAFLD-related liver injury. 54,55 The typical histological features of steatosis and inflammation often disappear in advanced disease, 56,57 thus many cases of 'cryptogenic' cirrhosis are likely caused by NASH. [56][57][58] Hepatocellular carcinoma is a well-recognised complication of NASH-related cirrhosis, 59,60 but can also be associated with precirrhotic NA-FLD. 61,62 Several systems have been proposed for the histological assessment of NAFLD, of which the Kleiner NAFLD activity score (NAS) 63 is probably the most well established. The NAS provides a composite score based on the degree of steatosis (0-3), lobular inflammation (0-3) and hepatocyte ballooning (0-2), with an additional score for fibrosis. A score of ‡5 suggests probable or definite NASH, and <3 indicates that NASH is unlikely. 63 However, although liver biopsy currently remains the gold standard for diagnosis of NASH, limitations of this technique include intra-observer variation 63,64 and sampling variability, 65,66 with features such as fibrosis often not uniformly distributed. 54
Such assessments can provide information on the amount of liver fibrosis and ⁄ or the presence of NASH, features which are usually, but not always, found together. The focus on fibrosis is based on cohort studies which demonstrate that fibrosis, rather than inflammation, predicts outcome. Several non-invasive diagnostic panels and scoring systems have been developed with varying diagnostic utility. The uneven distribution of fibrosis throughout the liver in NAFLD indicates that such scoring systems may potentially represent a more accurate reflection of global liver fibrosis severity than is permitted by the current gold standard liver biopsy, 67 which samples only 1 ⁄ 50 000th of the organ and is prone to significant sampling error. 65,66 Assessment of fibrosis. (i) Demographic factors and simple blood tests: Several diagnostic panels have been developed to facilitate the non-invasive assessment of NAFLD and differentiation between different stages of disease. These are generally based on a number of laboratory measurements, often in combination with clinical parameters such as age, sex and BMI. Such scoring systems have generally demonstrated greater utility in the detection of advanced fibrosis than intermediate and early stages of fibrosis, a group potentially more likely to benefit from therapeutic interventions. 37 The BARD score is a simple scoring system designed to identify NAFLD patients with a low risk of advanced disease. It combines three variables of BMI, AST ⁄ ALT ratio (AAR) and the presence of diabetes into a weighted sum (BMI ‡28 = 1 point, AAR of ‡0.8 = 2 points, DM = 1 point), to generate a score from 0 to 4. In the original study, a score of 2-4 was shown to be associated with an odds ratio for advanced fibrosis of 17 and a negative predictive value of 96%. 68 A further study of the BARD score in 138 patients with biopsy-proven NAFLD revealed an area under the receiver operating curve (AU-ROC) of 0.67 (95% CI, 0.56-0.77), with sensitivity, specificity, positive predictive value (PPV) and negative predictive value (NPV) of 51%, 77%, 45% and 81% respectively. 69 In a recent study including 145 Here the BARD score demonstrated an AUROC of 0.77, with sensitivity 89%, specificity 44%, NPV 95% and PPV 25%. 70 The BARD score was also validated in a Polish NAFLD cohort, where an NPV of 97% was demonstrated, 71 but appeared less useful in a Japanese cohort, where the AUROC was 0.73 with NPV 77%. 72 The BARD score is easily calculated and thus represents a simple tool for excluding the presence of advanced fibrosis in NAFLD patients.
The AST-to-platelet ratio index (APRI), 73 AST ⁄ ALT ratio, 74 and FIB-4 score 75 have previously demonstrated utility in the non-invasive assessment of fibrosis in a number of chronic liver diseases. Several recent studies have also examined the role of these markers in NAFLD, as will be described.
The APRI was originally developed for use in chronic hepatitis C, 73 but its utility in NAFLD has since been studied by a number of groups. Using this score, Cales et al. demonstrated an AUROC of 0.866 for significant fibrosis, 0.861 for severe fibrosis and 0.842 for cirrhosis in a study of 235 NAFLD subjects. 76 However, significantly lower values were obtained in other studies, where AUROCs of 0.564 for significant fibrosis, 0.568 for advanced fibrosis, 77 and 0.786 for predicting cirrhosis 78 were demonstrated. In their study of 145 NAFLD patients, McPherson et al. reported an AUROC of 0.67 for the diagnosis of advanced fibrosis. 70 The AST ⁄ ALT ratio (AAR) is calculated using two widely available laboratory liver function tests. In addition to its utility as an individual marker, the AAR is also a component of several other fibrosis scoring systems including the NAFLD Fibrosis score and BARD score. Despite its simplicity, using a cut-off of 0.8 McPherson et al. demonstrated an AUROC of 0.83, with sensitivity 74%, specificity 78% and NPV of 93% for the diagnosis of advanced fibrosis in NAFLD using the AAR. 70 The United States Nonalcoholic Steatohepatitis Clinical Research Network (NASH CRN) recently investigated the utility of readily available clinical and laboratory variables to predict histological severity of NASH in >600 patients with biopsy-proven NAFLD. In this study, a combination of serum AST, ALT and the AAR performed only modestly (AUROC 0.59) for predicting steatosis, but was able to predict cirrhosis with an AUROC of 0.81. However, the addition of demographic data, comorbidities and several other routinely measured laboratory tests increased the AUROCs to 0.79 for NASH and 0.96 for cirrhosis. 79 The FIB-4 test combines age with three standard biochemical values (platelets, ALT and AST) to assess fibrosis. In NAFLD FIB-4 has demonstrated similar results to the AST ⁄ ALT ratio where, using a cut-off of 1.3, an AU-ROC of 0.86, sensitivity 85%, specificity 65% and NPV of 95% were demonstrated for the diagnosis of advanced fibrosis. 70 In a US-based comparison of several non-invasive markers of fibrosis in 541 NAFLD patients, FIB-4 had the highest AUROC of 0.802, with PPV and NPV of 80% and 90% respectively for diagnosis of advanced fibrosis. In this study, AUROCs for the NAFLD fibrosis score, AAR, APRI, AST:platelet ratio and BARD score were 0.768, 0.742, 0.73, 0.72 and 0.70 respectively. 80 Increased serum GGT level has also been shown to be associated with advanced fibrosis in NAFLD, with a study of 50 NAFLD patients demonstrating an AUROC of 0.74 for the prediction of advanced fibrosis. Using a cut-off serum GGT value of 96.5 U ⁄ L, GGT predicted advanced fibrosis with 83% sensitivity and 69% specificity. 81 FibroMeter is a panel of serum markers which was originally developed for staging fibrosis in chronic HCV. 82 However, FibroMeter NAFLD has since been developed which has shown good diagnostic accuracy in staging NASH-related fibrosis. This panel combines seven variables (age, weight, fasting glucose, AST, ALT, ferritin and platelet count), and in a study of 235 NAFLD patients demonstrated AUROCs of 0.943 for significant fibrosis, 0.937 for severe fibrosis and 0.904 for cirrhosis respectively. The sensitivity, specificity, PPV and NPV of FibroMeter for diagnosing significant fibrosis were 78.5%, 95.9%, 87.9 and 92.1%. 76 The NAFLD fibrosis score (NFS) is a panel comprising six variables of age, hyperglycaemia, BMI, platelet count, albumin and AST ⁄ ALT ratio, which was constructed using a large panel of 733 biopsy-proven NA-FLD patients across several centres worldwide. Two cutoff scores were generated to predict the likelihood of the presence or absence of advanced fibrosis respectively. 67 In the original study, by applying the low cut-off score ()1.455), the NFS had an NPV of 93% and 88% in the estimation and validation groups respectively for excluding the presence of advanced fibrosis. By applying the high cut-off score (0.676), PPVs of 90% and 82% in the estimation and validation groups respectively were achieved for predicting the presence of advanced fibrosis. The AUROC was 0.84, and application of this model to the study population would have avoided liver biopsy in 75% of patients, with a correct prediction in 90%. 67 In the recent study by McPherson et al., the NFS demonstrated an AUROC of 0.81 with NPV 92% and PPV 72%, which was the highest PPV of the four tests examined. 70 Cales et al. demonstrated an AUROC of 0.884 for significant fibrosis, 0.932 for severe fibrosis and 0.902 for cirrhosis. 76 Studies in East Asian populations have also demonstrated good accuracy for excluding advanced fibrosis, with NPVs of 89% and 91% demonstrated in Japanese 72 and Chinese 83 NAFLD cohorts respectively. The NFS also demonstrated excellent accuracy at excluding fibrosis in morbidly obese subjects with NAFLD undergoing bariatric surgery, where NPVs of 98%, 87% and 88% for excluding advanced, significant and any fibrosis respectively were demonstrated. 84 In a recent meta-analyses, NFS achieved pooled AUROC, sensitivity and specificity of 0.85 (0.80-0.93), 0.90 (0.82-0.99) and 0.97 (0.94-0.99) for the identification of NASH with advanced fibrosis. 85 The NAFLD fibrosis score thus facilitates the identification of NAFLD patients with more advanced disease who require ongoing follow-up, and considerably reduces the requirement for liver biopsy in the minority of patients with an indeterminate score. 70 Of these various algorithms FIB-4 and the NAFLD fibrosis score (NFS) have been validated most widely with demonstrably superior test characteristics. 70,85 (ii) Fibrosis biomarkers: The Original ELF (European Liver Fibrosis) test is a panel of automated immunoassays to detect three markers of matrix turnover in serum: hyaluronic acid (HA), tissue inhibitor of metalloproteinase 1 (TIMP1) and aminoterminal peptide of pro-collagen III (P3NP), used in combination with age. 86 The simplified ELF panel excludes age but has a similar diagnostic performance. The addition of five simple markers -BMI, presence of diabetes ⁄ impaired fasting glucose, AST ⁄ ALT ratio, platelets and albumin -to the ELF test improved diagnostic accuracy further, with AUROCs of 0.98, 0.93 and 0.84 for the diagnosis of severe, moderate and no fibrosis respectively. 87 The ELF panel may also represent a useful prognostic tool, with a one unit change in ELF score shown to be associated with a doubling of the odds of significant liver-related mortality or morbidity at 6 year follow-up. 88 FibroTest is another validated marker for the quantitative assessment of fibrosis in NAFLD, ALD and chronic viral hepatitis. 89 Combining five biochemical markers of haptoglobin, a2-macroglobulin, apolipoprotein A1, total bilirubin and GGT, corrected for age and gender, a mean standardised AUROC of 0.84 for advanced fibrosis in NAFLD patients was demonstrated using FibroTest in one meta-analyses. Importantly the diagnostic value was found to be similar for the diagnosis of both intermediate and extreme fibrosis stages, with no significant difference between the AUROC of the intermediate adjacent stages F2 vs. F1 to that of the extreme stages F3 vs. F4 or F1 vs. F0. 89 FibroTest can also be combined with two other panels -SteatoTest and NASHTest -to form the FibroMax panel (BioPredictive, Paris, France), which provides a simultaneous and complete estimation of the liver injury in NAFLD. 89 FibroTest ⁄ FibroMax have now been widely adopted as a non-invasive alternative to liver biopsy. 90 Other potential fibrosis biomarkers include type VI collagen 7S domain and hyaluronic acid (HA), with the latter also representing a constituent of the ELF test. In a cohort of 112 NAFLD subjects, these two biomarkers were able to exclude advanced fibrosis with AUROCs of 0.82 and 0.80, and NPVs of 84 and 78% respectively. These biomarkers also demonstrated PPVs of 86% and 92% and AUROCs of 0.83 and 0.80 for discriminating NASH from simple fatty liver. 91 In a separate cohort of 72 patients, the AUROC curves for type IV collagen 7S domain and HA were 0.767 and 0.754 respectively for the detection of advanced fibrosis in NASH. However, after multiple regression analysis only type IV collagen 7S domain was independently associated with advanced fibrosis in this study. 92 Other studies of HA in NAFLD using varying cut-off levels have demonstrated AUROCs of between 0.89 and 0.97 for detecting advanced fibrosis. [93][94][95] Kaneda et al. demonstrated HA to have an AUROC, NPV, sensitivity and specificity of 0.97, 100%, 100% and 89% respectively for detecting severe fibrosis, with a lower AUROC of 0.87 demonstrated for type IV collagen. In this study, the platelet count alone was an independent predictor of cirrhosis, with an AUROC of 0.98, and sensitivity, specificity, PPV and NPV of 100%, 95%, 76% and 100% respectively using a cut-off value of 16 • 10 4 ⁄ lL. 94 Lesmana et al. also demonstrated the utility of levels of HA and type IV collagen to differentiate between mild (F1-2) and advanced fibrosis (F3-4), 96 and in a separate study of 80 NAFLD patients the combination of HA with AST, AAR, age, gender and BMI demonstrated an AUROC of 0.763 for distinguishing simple steatosis and NASH. 97 In a small study of serum extracellular matrix components in 30 NAFLD patients, serum laminin >282 ng ⁄ mL was shown to have an accuracy of 87%, sensitivity 82%, specificity 89%, PPV 82% and NPV 89% for identifying the presence of NASH with fibrosis. When combined with type IV collagen, both specificity and PPV increased to 100%, but with lower sensitivity and NPV of 64% and 83% respectively. 98 (iii) Radiological assessment: Although many imaging modalities have been evaluated in NAFLD, their major focus has been quantification of hepatic fat, with few allowing a reliable distinction between simple steatosis and steatohepatitis or fibrosis. 52,99 Conventional imaging techniques include ultrasound (US), computed tomography (CT) and magnetic resonance imaging (MRI).
Fibroscan. Transient elastography (Fibroscan, Echosens, Paris, France) is a non-invasive method of assessing liver fibrosis which can be performed at the bedside or in the out-patient clinic. It employs ultrasound-based technology to measure liver stiffness (LSM), and has been validated for use in chronic hepatitis C, HIV ⁄ HCV coinfection and cholestatic liver diseases. 100 Failure to obtain a reading occurs in only 5% of cases, but is more common in obese patients which has so far limited its use in the NAFLD cohort, 100 although a recently introduced XL probe may reduce this problem. 101 Although Fibroscan is less well validated in NAFLD, a stepwise increase in liver stiffness with increasing histological fibrosis was demonstrated in a study of 97 Japanese NAFLD patients, where AUROCs for the diagnosis of significant fibrosis, severe fibrosis and cirrhosis were 0.88, 0.91 and 0.99 respectively. 102 A larger study including 246 NAFLD patients from two ethnic groups demonstrated AUROCs for the diagnosis of moderate fibrosis ( ‡F2), bridging fibrosis ( ‡F3) or cirrhosis (F4) of 0.84, 0.93 and 0.95 respectively. 103 In this study, the best LSM cut-off scores for predicting F ‡ 2, F ‡ 3 and F4 were 7.0 kPa, 8.7 kPa and 10.3 kPa respectively. Cut-off values of 5.8 kPa and 9 kPa, and 7.9 kPa and 9.6 kPa had >90% sensitivity and specificity to rule out and rule in F2 and F3 fibrosis respectively. The cut-off of 10.3 kPa for F4 disease had 92% sensitivity and 88% specificity. 103 In this study, if liver biopsy was reserved for those patients with LSM of ‡8.7 kPa, 32% would require the procedure, which would miss only 4.8% patients with F3 and 0.6% patients with cirrhosis. However, if biopsy was performed only in those with scores between 7.9 and 9.6 kPa, only 16% patients would require biopsy. 103 In a 2010 metaanalyses of non-invasive assessment tools in NAFLD, transient elastography demonstrated pooled AUROC, sensitivity and specificity of 0.94 (0.90-0.99), 0.94 (0.88-0.99) and 0.95 (0.89-0.99). 85 Although further studies will undoubtedly add to the current evidence base, Fibroscan has now been validated in NAFLD, 104 and represents a useful tool for rapid, non-invasive assessment of liver fibrosis and determining need for biopsy.
An MR equivalent of transient elastography has recently demonstrated excellent diagnostic accuracy with sensitivity and specificity of 98% and 99% respectively for detecting all grades of fibrosis. 105 MR elastography was also associated with a higher technical success rate than US elastography, 106 and hepatic stiffness did not appear to be affected by steatosis using this technique 105 which had been a previous concern. 107 However, this technique remains experimental at the present time.
The combination of transient elastography with one or more of the serum marker panels described above represents a potential approach to the non-invasive measurement of fibrosis in NAFLD. 37 Assessment of NASH inflammation and steatosis. (i) Serum markers: SteatoTest combines 10 readily available blood tests with age, gender and BMI, and in a study of >2000 patients with viral hepatitis, NAFLD or ALD, demonstrated an AUROC of 0.8 for the diagnosis of steatosis, which was superior to that of GGT, ALT or ultrasound. 90,108 NASHTest combines 13 biochemical and clinical variables to predict the presence or absence of NASH, achieving specificity, sensitivity, PPV and NPV of 94%, 33%, 66% and 81% respectively. 109 Together with FibroTest, these three panels comprise the FibroMax panel described previously. 89 Many other potential serum biomarkers have been identified which are typically markers of the key mechanisms believed to be involved in NASH pathogenesis, such as inflammation, oxidative stress, apoptosis and insulin resistance. 44 Inflammation is associated with an increase in tumour necrosis factor-alpha (TNF-a) and decreased adiponectin expression (measured by ELISA) and this cytokine imbalance does appear to correlate with NASH, [110][111][112] although the accuracy or clinical usefulness of these markers is yet to be determined. 113
observed the serum adiponectin level to be significantly lower in patients with early-stage NASH (3.6 lg ⁄ mL) than in those with simple steatosis (6.0 lg ⁄ mL), with an AUROC of 0.765, sensitivity 68% and sensitivity 79% for distinguishing early-stage NASH. In this study, the combination of serum adiponectin level with HOMA-IR and type IV collagen 7S demonstrated a sensitivity of 94% and specificity of 74% for diagnosing NASH. 114 Hui et al. also reported significantly lower adiponectin levels and higher HOMA-IR in patients with NASH compared with subjects with simple steatosis, and demonstrated an AU-ROC of 0.79 for the combination of these markers for distinguishing between steatohepatitis and steatosis. 115 However, the relationship between adiponectin levels and severity of hepatic fibrosis remains to be established. 116,117 The inflammatory marker C-reactive protein (CRP), measured using immunometric assay, has demonstrated mixed results in NASH. While some studies have shown a significant increase in high-sensitivity CRP levels in NASH patients compared with controls, 118,119 another study demonstrated no significant difference. 120 CRP also lacks specificity for hepatic inflammation. Interleukin-6 (IL-6) is another marker of inflammation which has been shown to be elevated in NASH. 121 In a study comparing 43 NASH patients, 40 subjects with steatosis and 48 controls, normal levels of IL-6 were highly specific in confirming the absence of NASH, with an AUROC of 0.817 for distinguishing NASH from simple steatosis. The same study also demonstrated significantly elevated levels of vascular endothelial growth factor (VEGF) concentrations in NASH, although the AUROC of 0.678 was lower than for IL-6. 122 IL-6 levels were also independently associated with fibrosis in a study by Lemoine et al., where the combination of HOMA-IR with the adiponectin ⁄ leptin ratio demonstrated an AUROC of 0.82 for distinguishing between NASH and simple steatosis. 123 Indicators of oxidative stress, including lipid peroxidation products (measured using a spectrofluorometric method), vitamin E levels (measured using reverse phase high performance liquid chromatography), and copperto-zinc superoxide dismutase and glutathione peroxidase (GSH-Px) activity (measured using commercial antioxidant assay kits), have been investigated as surrogate markers of NASH. However, most studies to date have been small with mixed results, 113,124,125 and it is not yet clear whether or not oxidative stress in the liver is accurately reflected in the serum. 126 Thioredoxin (TRX) is stress-inducible thiol-containing protein which may represent a clinically useful indicator of oxidative stress. 127 In a small study of 57 patients, Sumida et al. demonstrated significantly elevated serum TRX levels in patients with NASH compared with those with simple steatosis and healthy controls, with an AU-ROC of 0.785 for distinguishing NASH from simple steatosis. 127 Similar findings have been reported elsewhere, with a correlation between serum TRX and ferritin levels also observed. 128 Apoptosis plays an important role in the liver injury observed in NAFLD, 129 Use mainly confined to specialist liver centres CK-18 Multi-centre validation study including 139 NAFLD patients demonstrated AUROC of 0.83 for NASH diagnosis. 130 Recently received independent validation for diagnosing NASH in meta-analyses demonstrating pooled AUROC, sensitivity and specificity for NASH of 0.82, 0.78 and 0.87 85 Use currently limited to clinical trials but holds promise for more widespread application b) Diagnostic technique Size of cohort used in studies ⁄ Comments Popularity
Validation study included 196 NAFLD patients. AUROCs of 0.84, 0.93 and 0.98 for detecting no fibrosis, moderate fibrosis and severe fibrosis respectively 87 Use mainly confined to specialist liver centres FibroMeter Study of 235 NAFLD patients demonstrated AUROCs of 0.943 for significant fibrosis, 0.937 for severe fibrosis and 0.904 for cirrhosis respectively. Sensitivity, specificity, PPV, NPV for diagnosis of significant fibrosis were 79%, 96%, 88% and 92% 76 Use mainly confined to specialist liver centres Fibroscan Largest study included 246 NAFLD patients, with AUROCs of 0.84, 0.93 and 0.95 for diagnosis of ‡F2, ‡F3 and ‡F4 fibrosis respectively. 103 Meta-analyses of non-invasive assessment tools in NAFLD demonstrated pooled AUROC, sensitivity and specificity of 0.94, 0.94 and 0.95 85 Use mainly confined to specialist liver centres. Equipment expensive
Validation study included 733 patients with biopsy-proven NAFLD: 480 in estimation group and 253 in validation group. PPVs of 90% and 82% for predicting advanced fibrosis and NPVs of 93% and 88% for excluding advanced fibrosis in estimation and validation groups respectively, with AUROC 0.84 67 For diagnosis of advanced fibrosis, AUROC of 0.768 demonstrated in study of 541 patients, 80 and AUROC of 0.81 in study of 145 NAFLD patients. 70 In study of 235 patients, AUROCs of 0.884 for significant fibrosis, 0.932 for severe fibrosis, and 0.902 for cirrhosis 76 Similar efficacy demonstrated in East Asian 72,83 and morbidly obese cohorts 84 Widely used in secondary care clinical practice. Accessibility increased by availability of on-line calculator. Use advocated in recent review of non-invasive scoring systems for exclusion of advanced fibrosis 70 FibroTest Meta-analyses including 267 NAFLD patients demonstrated mean AUROC of 0.84 for detecting advanced fibrosis. Can be combined with SteatoTest and NASHTest in Fibromax 89 Use mainly confined to specialist liver centres FIB-4 Study of 541 NAFLD patients demonstrated AUROC of 0.802, with PPV 80% and NPV 90% for diagnosis of advanced fibrosis. 80 Study of 145 NAFLD patients demonstrated AUROC of 0.86, sensitivity 85%, specificity 65% and NPV of 95% 70 Use advocated in recent review of non-invasive scoring systems for exclusion of advanced fibrosis 70
Study of 235 NAFLD patients demonstrated AUROCs of 0.866 for significant fibrosis, 0.861 for severe fibrosis and 0.842 for cirrhosis. 76 AUROCs of 0.73 for advanced fibrosis in study of 541 NAFLD patients, 80 and 0.67 in study of 145 patients 70 Developed for use in HCV. Not specific for NAFLD S Sy ys st te em ma at ti ic c r re ev vi ie ew w: : d di ia ag gn no os si is s a an nd d s st ta ag gi in ng g o of f N NA AF FL LD D plasma CK-18 levels measured using ELISA were significantly higher in patients with biopsy-proven NASH than in those with a borderline diagnosis and normal controls, with an AUROC of 0.83 for NASH diagnosis. CK-18 was an independent predictor of both NASH and severity of disease. 130 Other studies have corroborated these findings, [131][132][133][134][135][136][137] suggesting that CK-18 represents a potentially useful biomarker for the diagnosis and differentiation of NASH from simple steatosis. CK-18 recently received independent validation for diagnosing NASH in a 2010 meta-analyses, where pooled AUROC, sensitivity and specificity for NASH were 0.82 (0.78-0.88), 0.78 (0.64-0.92), and 0.87 (0.77-0.98) respectively. 85 An AU-ROC of 0.88 for diagnosis of NASH was also demonstrated in a morbidly obese population where, in addition, CK-18 levels were observed to fall significantly following bariatric surgery. 138 Younossi et al. evaluated the diagnostic utility of several ELISA-based assays in patients with biopsy-proven NASH. This study found that the levels of cleaved CK-18 (M30 antigen), and intact CK-18 (M65) predicted histological NASH with 70% sensitivity, 84% specificity, AUROC 0.711 and 64% sensitivity, 89% specificity and AUROC 0.814 respectively. Histological NASH was found to be predicted by a combination of four ELISA-based tests -cleaved CK-18, a product of the subtraction of cleaved CK-18 level from intact CK-18 level, serum adiponectin and serum resistin -with a sensitivity 96%, specificity of 70% and AUROC of 0.91. 139 CK18 fragment has also been combined with ALT levels and the presence of the metabolic syndrome in the 'Nice model', a composite model where AUROCs of 0.83-0-88 were demonstrated for the diagnosis of NASH, as defined by a NAS score ‡5, in a morbidly obese population. 140 Plasma homocysteine (Hcy) levels were shown to distinguish NASH from simple steatosis with good accuracy in a Turkish study of 71 NAFLD patients. Using a threshold of 11.935 ng ⁄ mL, sensitivity and specificity were 91.7% and 95.7% with an AUROC of 0.948 for predicting NASH. 141 Serum prolidase enzyme activity (SPEA) catalyses the final step of collagen breakdown by liberating free proline for collagen recycling, and is reported to be of hepatic origin. 142 In a Turkish study of 54 NAFLD patients, SPEA was significantly greater in patients with NASH than those with simple steatosis or controls, and positively correlated with the grade of liver fatty infiltration, lobular inflammation, stage of fibrosis and NAFLD activity score. 142 In this study, the SPEA was shown to be superior to AST, ALT or the AAR for distinguishing NASH from simple steatosis, with an AUROC of 0.85, and also the most useful of these tests for predicting lobular inflammation, NAFLD activity score and fibrosis. 142 Other novel biomarkers which may be of potential utility BARD score Original study included 827 NAFLD patients with NPV 96% for excluding advanced fibrosis. 68 AUROC 0.70 for advanced fibrosis in study of 541 NAFLD patients. 80 Study of 145 patients showed AUROC of 0.77, with sensitivity 89%, specificity 44%, NPV 95% and PPV 25% 70 Simple to use in primary or secondary care. Use advocated in recent review of non-invasive scoring systems for exclusion of advanced fibrosis 70 Techniques to non-invasively distinguish NASH from simple steatosis remain largely experimental, although show much promise for the future. Several methods are available for the non-invasive diagnosis ⁄ exclusion of fibrosis in patients with NAFLD. The BARD score is the simplest scoring system to calculate in the clinic, but the NAFLD fibrosis score can also be easily calculated by entering the relevant details into a freely available online calculator 70 (http://nafldscore.com). Currently FibroTest ⁄ FibroMax, FibroMeter and the ELF test are each only available from a single laboratory at significant cost, which has limited their use to mainly specialist liver centres.
in the diagnosis of NASH include plasma pentraxin 3 levels 143 and tissue polypeptide specific antigen. 144 Further study of these markers is warranted.
It should be highlighted that although some of the non-invasive serum markers described earlier are already widely used in the assessment and staging of NASH, fur-
Stage 1 NAFLD confirmed Stage 2 Calculate NAFLD fibrosis score <-2.5 -2.5 to -1.455 -1.455 to 0.676 >0.676 Low risk Indeterminate risk High risk Reassure Refer/Further investigation Repeat NFS in: Fibroscan 3-5 years 1 year <7.9 kPa 7.9-9.6 kPa NAS Indeterminate NASH (≥F3 fibrosis) NAS NASH (activity++ or ≥F2 fibrosis) Management of NASH and risk factors Discharge back to primary care >9.6 kPa Recommend liver biopsy Figure 2 | Proposed algorithm for the work-up of a patient with NAFLD. Patients with a NAFLD fibrosis score below the lower cut-off level have a low risk of significant fibrosis and subsequent disease progression and can be safely managed in primary care. Referral to specialist care is indicated if disease progression is suspected on clinical or biochemical grounds. A score in the indeterminate range or above merits further investigation by use of modalities such as specialist scans or blood tests. Liver biopsy should be considered for those patients in whom non-invasive tests are inconclusive. The use of Fibroscan in this algorithm may later be replaced by serum marker panels. NFS, NAFLD fibrosis score; NAS, non-alcoholic steatosis.
ther validation of their use in this setting will increase confidence in their utility (Table 2).
(ii) Radiological assessment: Contrast-enhanced ultrasound using Levovist is the first imaging technique to demonstrate efficacy in distinguishing between simple steatosis and NASH. In a study of 64 patients with either normal liver, NAFLD or NASH, this modality was able to diagnose NASH with an AUROC of 100%. 145 The accumulation of Levovist microbubbles in the liver parenchyma was shown to be decreased in NASH but not in NAFLD or chronic viral hepatitis, with the decrease seen in NASH correlating with fibrosis rather than steatosis. Severe decrease was seen in both NASH and ASH livers, especially in the presence of bridging fibrosis (F3); however, the same stage of fibrosis in chronic viral hepatitis showed only a mild decrease or normal uptake of Levovist. These differences may be explained both by changes in Kupffer cell function and differences between pericellular and periportal fibrosis in provoking disturbance of Levovist microbubble accumulation. 145 Although this remains an experimental technique at present, evolving imaging modalities may thus soon permit the non-invasive differentiation between steatosis and NASH.
The diagnostic yield using CT is similar to that of US in NAFLD. Unenhanced CT shows low attenuation of the steatotic liver in contrast to the spleen, and the severity of steatosis has been shown to correlate with the liver:spleen (L ⁄ S) attenuation ratio. 146 A CT L ⁄ S cut-off of value of 0.8 yielded 100% specificity and 82% sensitivity for diagnosing macrovesicular steatosis of 30% or greater, 147 although a CT L ⁄ S cut off of 1.0 for defining steatosis has been used in some studies. 148 Accuracy of unenhanced CT is greatly reduced with lesser degrees of steatosis. 147 Other pathologies, such as hepatic siderosis, may also alter attenuation values leading to misdiagnosis, 149,150 and the radiation exposure associated with CT limits its use in younger patients and in longitudinal studies. 149 Fibroscan has also recently demonstrated utility in detecting and quantifying steatosis, using a novel attenuation parameter termed 'Controlled Attenuation Parameter' (CAP), which was devised to specifically target the liver using a process based on Vibration Control Transient Elastography (VCTE, Echosens, Paris, France). A study of 115 patients using liver biopsy as reference demonstrated that CAP was able to accurately detect >10% (S1), >33% (S2) and >67% (S3) steatosis with AU-ROCs of 0.91, 0.95 and 0.89 respectively. CAP evaluated by the Fibroscan is not affected by fibrosis, and advantages over other imaging techniques include its ability to quantify and detect steatosis from only 10% of liver infiltration, and being non-ionising, relatively cheap and non-operator dependent. 151 There are at present no evidence-based guidelines for the work-up ⁄ staging of a patient with confirmed NAFLD. Although alternative approaches may be equally effective, a proposed algorithm for the work-up of patients with NAFLD is detailed in Figure 2. This combination of inexpensive non-invasive algorithms prior to more detailed examinations should hopefully permit a costeffective and efficient way of investigating what is likely to be a large number of potential patients.
Non-alcoholic fatty liver disease now represents the most common cause of liver disease in the Western world, and rising levels of obesity, diabetes and the metabolic syndrome render it an increasingly important cause of morbidity and mortality. While simple steatosis carries a relatively benign prognosis, a significant proportion of patients will progress to NASH and later cirrhosis with risk of hepatocellular carcinoma. Although liver biopsy remains the gold standard for disease assessment, the development of risk scores, biomarker panels and ultrasound modalities has resulted in much improved identification of at risk patients without recourse to use of liver biopsy on a routine basis.
Aliment Pharmacol Ther 2011; 33: 525-540
ª 2010 Blackwell Publishing Ltd
Aliment Pharmacol Ther 2011; 33: 525-540 ª 2010 Blackwell Publishing Ltd
Declaration of personal interests:
IgE antibodies specific to staphylococcal enterotoxin A or B (SEA or SEB) (8, 9). Lipoteichoic acid, cell wall component in S. aureus, has been detected in the lesions of AD patients and is correlated with AD severity (10). This finding together with the fact that about 20-40% of AD patients display intrinsic AD with no sensitization to any protein allergen suggests the importance of microbes (11).
Gram-negative bacteria secrete outer membrane vesicles (OMV) (12), which have pathogenic effects (13,14). Recently, we demonstrated for the first time that the Gram-positive bacterium S. aureus produces OMV-like vesicles called extracellular vesicles (EV) (15). The EV produced by S. aureus are spherical with a diameter of 20-100 nm and are shed from the bacterium's membranes. Proteomic analysis revealed that
Skin lesions in atopic dermatitis (AD) patients display histological alterations such as epidermal thickening and infiltration by eosinophils and mast cells (1). The pathogenesis of AD is believed to involve repeated abnormal innate and adaptive immune responses to environmental causative agents when the skin barrier is disrupted (2). Aeroallergens, including house dust mites and pollens, are well known to induce AD through IgE-mediated mechanisms (3). In addition, microbes represent another important group of extrinsic causative factors in the pathogenesis of AD (4). Staphylococcus aureus (S. aureus) appears to be particularly important because it colonizes almost all lesional skin in AD patients, and a reduction in its colonization has been shown to decrease disease severity (5)(6)(7). Some AD patients produce the protein expression pattern in these EV differs from that in whole bacteria, and that they contain various pathogenic molecules. Among these, a-hemolysin and cysteine protease have been linked with AD (16,17). These findings suggest that EV is a potent initiator of host immune responses.
In AD patients, S. aureus, rather than invading and infecting the skin, colonizes it. It is thus reasonable to assume that the skin is affected by secretory products from S. aureus. Given that staphylococcal secretory products are relevant to the pathogenesis of AD, we hypothesized that S. aureus EV, which are complexes of various pathogenic molecules secreted by the bacterium, are involved in the pathogenesis of AD. In this study, we found through in vitro and in vivo studies that S. aureus EV are causative agents in AD, and the clinical observation also supports this hypothesis.
SKH-HR1 (hairless) mice were purchased from Charles River Laboratories Japan, Inc. (Yokohama, Japan) and were bred in a pathogen-free facility at Pohang University of Science and Technology (POSTECH; Pohang, Korea). All animal experiments were approved by the POSTECH Ethics Committee.
Skin lavage fluids were obtained from two AD patients visiting Pediatric Clinic of Seoul Suncheonhyang Hospital (Seoul, Korea). Serum samples were obtained from 60 AD patients (30 patients aged 6-9 years and 30 patients aged over 10 years) and 20 healthy subjects aged 6-16 years, who were recruited from Seoul Samsung Hospital (Seoul, Korea). Skin lavage fluids and serum samples were isolated after written informed consent had been obtained. The study protocol was approved by the Ethics Committee of Seoul Suncheonhyang Hospital and Seoul Samsung Hospital, respectively.
Staphylococcus aureus EV were obtained as described previously (15). Briefly, S. aureus (ATCC14458) was cultured in nutrient broth (Merck, Darmstadt, Germany) and grown at 37°C to 1.0 of optical density (at 600 nm). Bacteria were removed by centrifugation, and the resulting supernatant was filtered through a 0.45-lm vacuum filter. The filtrate was concentrated by ultrafiltration using the QuixStand Benchtop System (Amersham Biosciences, Piscataway, NJ, USA) in conjunction with a 100-kD hollow-fiber membrane (Amersham Biosciences). The resulting concentrated filtrate was passed through a 0.22-lm vacuum filter. Extracellular vesicles were isolated from the resulting filtrate by ultracentrifugation at 150 000 g. The concentration of protein in the EV was measured by Bradford assays (Bio-Rad Laboratories, Hercules, CA, USA). Hereafter, the reported doses of EV refer to the amount of EV protein. Isolated EV were stored at )80°C before use. Bacteria and EV removed, and then <100 and >100-kD soluble fractions of bacterial culture media were concentrated. Protein concentration was measured by Bradford assays, and culture media were stored at )80°C before use.
The presence of SEA and SEB in EV and concentrated culture media was analyzed using SDS-PAGE and western blot. SEA and SEB were detected by monoclonal anti-SEA and anti-SEB antibody (Santa Cruz Biotechnology, Santa Cruz, CA, USA).
To create a mouse model of AD, S. aureus EV were applied to mouse skin according to the following protocol. To disrupt the cutaneous barrier, the dorsal skin of 6-week-old mice was stripped five to six times using Durapore surgical tape (3M Co., St. Paul, MN, USA). Gauze (1.5 • 1.5 cm) soaked with S. aureus EV in 100 ll of phosphate buffered saline (PBS) was then placed on the stripped skin and secured using Tegaderm bio-occlusive tape (3M Co.). For the evaluation of inflammation and immune dysfunction, the mice were euthanized 48 h after the final challenge.
Four-micrometer-thick sections of fixed skin tissues were stained with hematoxylin and eosin (H&E). Mast cells were stained with toluidine blue (TB). Cells were counted in 15-25 high-power fields at a magnification of •400.
Single cells of skin-draining lymph nodes (LN) were collected and stimulated with or without 0.1 lg/ml of S. aureus EV. Supernatants were harvested after 72 h, and the levels of cytokines measured by ELISA.
In vitro production of pro-inflammatory mediators from dermal fibroblasts Primary mouse dermal fibroblasts were isolated as described previously with some modification (18). Fibroblasts from passages 1-3 were used. Then, 2 • 10 4 cells of isolated cells were cultured in 24-well plates then treated with 1 or 10 lg/ml S. aureus EV or soluble fractions of bacterial culture media, or SEB (Toxin Technology, Sarasota, FL, USA). Supernatants were collected 24 h after stimulation, and mediator levels measured.
The cytokine and chemokine levels were measured by ELISA (R&D Systems, Mineapolis, MN, USA) according to the manufacturer's instructions.
Skin lavage fluids were obtained by rinsing patients' skin lesions three to four times with 50 ml of sterile PBS and were stored at )80°C. To remove bacteria and other debris, 40 ml of skin lavage fluids was centrifuged at 5000 and 10 000 g. After centrifugation, supernatants were filtered through 0.45 and 0.22 lm serially. Then, lavage fluids were concentrated to 1 ml by using Centriprep YM-50 (Millipore, Carringtwohill, Ireland). Same volume of PBS was added on concentrated lavage fluids, and they were ultracentrifuged at 150 000 g for 3 h at 4°C. The pellet was used as EV fraction.
Anti-S. aureus EV-specific polyclonal antibodies were coated on 96-well ELISA plate. Each well was blocked by 1% bovine serum albumin in PBS. After blocking, concentrated lavage fluids and EV fraction were added to each well. Then, biotinylated anti-S. aureus EV-specific polyclonal antibodies were added on each well, and then streptavidin-horseradish peroxidase (HRP) was added. After a final wash, chemiluminescence substrates (POD) were added to react with HRP.
Luminescence measured using a Wallac 1420 Victor luminometer (American Instrument Exchange, Inc., Haverville, MA, USA).
Serum samples were prepared from mouse blood for analysis of the total serum IgG1 and IgE levels by ELISA (Bethyl Laboratories, Montgomery, TX, USA) according to manufacturer's instructions.
Staphylococcus aureus EV-and SEB-specific IgG1 and IgE levels were measured by ELISA. The wells of 96-well ELISA plates were each coated with 0.1 lg of S. aureus EV or 1 lg of SEB. Then, wells were blocked with 3% bovine serum albumin in PBS. After blocking, diluted human serum was added to the wells. Then, HRP-conjugated anti-human IgG1 and IgE antibodies (Southern Biotech, Birmingham, IL, USA) were applied to each well. Chemiluminescence substrates were added, and luminescence was measured by the luminometer. Specific antibody levels were defined as elevated if a value is higher than mean +1 standard deviation of healthy subjects' values. tion with EV and >100-and <100-kD soluble fractions of bacterial culture media. (C) Western blotting to detect staphylococcal enterotoxin B (SEB) in EV and soluble (Sup) fractions of S. aureus culture media. (D) Levels of pro-inflammatory mediators from supernatants of dermal fibroblasts after stimulation with EV or SEB. Assays were performed in duplicate. *P < 0.05; **P < 0.01; ***P < 0.001.
Statistically significant differences between treatments were identified using Student's t-test, anova, or Wilcoxon's rank sum test. Multiple comparisons were initially made by anova.
Where significant differences were found, individual t-tests or Wilcoxon's rank sum tests were used to identify statistically significant differences between treatment group pairs. A P-value <0.05 was considered to be statistically significant.
In vitro production of pro-inflammatory mediators from mouse dermal fibroblasts treated with Staphylococcus aureus extracellular vesicles
Recently, we found that S. aureus EV contain 90 proteins, including proteins with pathological function, by proteomic analysis (15). Scanning electron microscopic images showed that S. aureus secreted EV (Fig. 1A). We evaluated whether EV or soluble fractions of S. aureus culture media induce the production of pro-inflammatory mediators from skin fibroblasts. We found that the production of IL-6, thymic stromal lymphopoietin (TSLP), macrophage inflammatory protein (MIP)-1a, and eotaxin by dermal fibroblasts was higher by stimulation with EV than with soluble fraction (Fig. 1B). It was reported that the presence of IgE antibodies to SEA and SEB was correlated with the severity of skin lesions in children with AD (8,9). Western blotting using anti-SEA and anti-SEB antibodies showed that SEB was present in soluble fraction, but not in EV, whereas SEA absent in both fractions (Fig. 1C). We compared in vitro activity between EV and SEB on the production of pro-inflammatory mediator. As shown in Fig. 1D, 1 or 10 lg/ml of SEB did not enhance the production of IL-6, TSLP, MIP-1a, and eotaxin, whereas •200 (upper panel) and •400 (lower panel)]. Arrowheads identify eosinophils. (C) Histological analysis of epidermal thickness and the numbers of eosinophils and mast cells infiltrating the dermis. *P < 0.05; **P < 0.01; ***P < 0.001.
1 lg/ml of EV upregulated the production of these mediators. Collectively, these findings suggest that S. aureus-derived EV are more potent compared to soluble components, in terms of production of pro-inflammatory mediators.
We evaluated the in vivo effects of S. aureus EV on the induction of skin inflammation. Different doses of EV were applied to tape-stripped mouse skin, and changes in skin inflammation were assessed 4 weeks after initial application (Fig. 2A). Histological analysis showed that the application of S. aureus EV to tape-stripped skin induced AD-like inflammation, including epidermal thickening and infiltration of the dermis by inflammatory cells (Fig. 2B). Moreover, S. aureus EV dose dependently induced epidermal thickening (Fig. 2C). As with dermal infiltration by inflammatory cells, significantly higher numbers of mast cells were found in the dermis of mice treated with S. aureus EV (at 5 or 10 lg, but not 0.1 lg) compared to PBS-treated controls. In addition, the numbers of eosinophils were significantly higher in EVtreated mice than in PBS-treated controls at all doses (Fig. 2 C). Together, these data suggest that the application of S. aureus EV to tape-stripped skin induces AD-like inflammation.
To test whether the application of S. aureus EV to the skin induces immunological dysfunction, 5 lg of S. aureus EV was applied to tape-stripped mouse skin three times per week for 3 weeks. Histological analysis showed that the application of S. aureus EV induced epidermal thickening and infiltration of the dermis by inflammatory cells (Fig. times a week for 3 weeks) to tape-stripped skin. *P < 0.05; **P < 0.01; ***P < 0.001. (A) Skin histology [H&E staining; magnification, •200]. (B) Histological analysis of epidermal thickness. (C) Levels of IFN-c and IL-17 in supernatants from S. aureus EV-treated cells from skin-draining LNs. (D) Levels of IFN-c, IL-17, IL-4, and IL-5 in skin tissue homogenates.
A). Tape stripping itself induced epidermal thickening (Fig. 3 B). To characterize the immune dysfunction induced by S. aureus EV, we measured the production of Th1, Th17, and Th2 cytokines by T cells from skin-draining LNs following in vitro stimulation with S. aureus EV. The production of IFN-c and IL-17 was significantly higher in cells from the EV-treated mice than in those from the PBS-treated animals (Fig. 3C). IL-4 and IL-5 were not detected in the supernatants from cells isolated from either treatment group (data not shown). In terms of the skin production of Th1, Th17, and Th2 cytokines, the levels of IFN-c, IL-17, IL-4, and IL-5 in skin homogenates from EV-treated mice were significantly higher than those from PBS-treated mice. Tape stripping did not itself enhance the production of these cytokines (Fig. 3D). With regard to the production of antibodies, we found the serum levels of total IgG1 and IgE to be similar in the EV-and PBS-treated mice (data not shown). Collectively, these data suggest that exposure to S. aureus EV for 3 weeks induces a mixed Th1-/Th17-/Th2-type inflammatory response in the skin, whereas a mixed Th1/Th17 cell response in skin-draining LNs.
To assess the effects of long-term exposure to S. aureus EV, S. aureus EV (5 lg) were applied to tape-stripped skin three times per week for 8 weeks. Histological analysis showed that prolonged exposure to S. aureus EV induced epidermal thickening and increased infiltration of the dermis by inflammatory cells (Fig. 4A,B). Infiltration of the dermis by mast cells and eosinophils was significantly increased in EV-treated mice relative to PBS-treated controls (Fig. 4A,B). Furthermore, long-term in vivo exposure to S. aureus EV enhanced the production of IL-17 by T cells in skin-draining LNs stimulated with S. aureus EV in vitro (Fig. 4C). However, it did not enhance the EV-induced production of IFN-c and IL-4 (data not shown). In terms of antibody production following long-term exposure to S. aureus EV, the serum total IgE levels were significantly higher in the EV-treated mice than in the PBS-treated controls, although the serum total IgG1 levels were similar between the two groups (Fig. 4D). Taken ing; magnification, •200 (upper panel) and •400 (lower panel)]. Arrowheads identify eosinophils. (B) Histological analysis of epidermal thickness and numbers of eosinophils and mast cells infiltrating the dermis. (C) Levels of IL-17 in supernatants from S. aureus EV-treated cells from skin-draining lymph nodes. (D) Serum levels of total IgG1 and total IgE.
together, these findings suggest that long-term exposure to S. aureus EV induces IgE production, whereas Th17-cell response in skin-draining LNs.
We evaluated the presence of S. aureus-derived EV in the skin lesion of AD patients. To test this objective, EV were isolated from the skin lesion of two AD patients and then evaluated whether EV isolated from the skin lesion of AD patients incorporate S. aureus EV-specific proteins using anti-S. aureus EV-specific polyclonal antibodies. As shown in Fig. 5, EV isolated from skin lavage fluids of AD patients contained S. aureus EV-specific proteins.
Finally, we measured serum levels of S. aureus EV-and SEBspecific antibodies in AD patients and healthy subjects. The serum levels of S. aureus EV-and SEB-specific IgG1 were comparable in AD patients and healthy subjects (Fig. 6A). By contrast, the serum S. aureus EV-specific IgE levels were significantly higher in both 6-9 year aged and >9 year aged AD patients than in age-matched healthy subjects; this IgE levels were elevated in 33.3% of 6-9 years aged and 40% of >9 years aged AD patients (Fig. 6B). In addition, the serum SEB-specific IgE levels were significantly higher in only >9 years aged AD patients than in age-matched healthy subjects; this IgE levels were elevated in 33.3% of 6-9 years aged and 33.3% of >9 years aged AD patients (Fig. 6C). However, neither total IgE nor SEB-specific IgE were correlated with S. aureus EV-specific IgE in AD patients (data not shown). These findings indicate that S. aureus EV-specific IgE may be a useful biomarker for identifying the cause of AD in individual cases.
Elucidating the pathogenesis of AD has been difficult because of complex interactions between the causative factors and host immune response. Among the known causative factors, microbes are thought to be particularly important in the pathogenesis of AD by elaborating secretary products, including soluble toxins. Recently, we found that S. aureus secretes EV, in which pathogenic proteins are incorporated (15). In the present study, experimental and clinical data support the hypothesis that S. aureus EV are involved in the pathogenesis of AD.
In terms of host immune responses of AD pathogenesis, recent studies showed that IL-17 producing cells are found in the blood of AD patients and that AD pathogenesis is related with IL-17-mediated immune responses (19,20). In the present study, we found that the cutaneous application of S. aureus EV induced Th17-cell response in skin-draining LNs, suggesting that Th17-cell response is important in the AD pathogenesis.
The present data showed that the cutaneous application of S. aureus EV induced skin inflammation characterized by infiltration of mast cells and eosinophils. This inflammatory response was associated with enhanced production of not only Th1/Th17-type cytokines, but also Th2-type cytokines in the skin. In addition, the present study showed that in vitro stimulation of fibroblasts with S. aureus EV increased the secretion of the Th2-type cytokines, such as TSLP and eotaxin (21). These findings suggest that Th2-type inflammation induced by S. aureus EV is mediated by local production of Th2-type cytokines from dermal fibroblasts.
Staphylococcus aureus can colonize in skin or nasal passage in human (22). In the present study, we firstly demonstrated that S. aureus-derived EV are present on the skin of AD patients. We also found that S. aureus EV-specific IgG1 were detected in serums from not only AD patients but also healthy subjects. These findings suggest that S. aureus secretes EV in skin, which induce systemic immune responses.
Previous studies showed that serum levels of IgE specific to staphylococcal enterotoxins were not only elevated in AD patients, but also correlated with disease severity (8,9). Our present data indicate that serum levels of both S. aureus EV-and SEB-specific IgE were elevated in the AD patients compared to healthy subjects. In addition, the present study showed that S. aureus EV did not contain enterotoxins and S. aureus EV-specific IgE levels did not correlate with SEB-specific IgE. These findings suggest that the production of EV-specific IgE is not influenced by staphylococcal enterotoxins.
Recently, it was reported that IL-17 can induce IgE production in B cells (23). The present study showed that the cutaneous application of S. aureus EV for 8 weeks induce Th17-cell response, but not Th2-cell response, along with the Recent evidence indicates that EV from Gram-negative bacteria induced systemic inflammatory response (13). We found that Gram-positive bacteria produced EV (15). This is the first report to show that EV obtained from Gram-positive bacteria can cause inflammatory disease. Given the abundance of Gram-positive bacteria in our environment, further research will be needed to elucidate the relationships between EV from Gram-positive bacteria and the pathogenesis of immune-based inflammatory diseases.
In summary, our present data indicate that S. aureusderived EV can induce AD-like inflammation in the skin, and that the physiological animal model we employed represents a useful tool for performing translational research into AD. Furthermore, our clinical findings provide strong evi- patients (right panel). (C) Levels of SEB-specific IgE in serum from AD patients and healthy subjects (left panel), and positive rate of elevated SEB-specific IgE in AD patients (right panel). Serum samples from AD patients (n = 30 in patients aged 6-9 years, and n = 30 in patients aged 9-16 years) and healthy subjects aged 6-16 years (n = 20); EV, S. aureus-derived EV; SEB, staphylococcal enterotoxin B. *P < 0.05; **P < 0.01.
dence of the importance of S. aureus EV in the pathogenesis of AD.
Allergy 66 (2011) 351-359 ª 2010 John Wiley & Sons A/S
We thank
The authors declare that there are no conflicts of interest.
Background: There are no published data on peanut sensitization in Egypt and the problem of peanut allergy seems underestimated. We sought to screen for peanut sensitization in a group of atopic Egyptian children in relation to their phenotypic manifestations.
Methods: We consecutively enrolled 100 allergic children; 2-10 years old (mean 6.5 yr). The study measurements included clinical evaluation for site of allergy, possible precipitating factors, consumption of peanuts (starting age and last consumption), duration of breast feeding, current treatment, and family history of allergy as well as skin prick testing with a commercial peanut extract, and serum peanut specific and total IgE estimation. Children who were found sensitized to peanuts were subjected to an open oral peanut challenge test taking all necessary precautions.
Results: Seven subjects (7%) were sensitized and three out of six of them had positive oral challenge denoting allergy to peanuts. The sensitization rates did not vary significantly with gender, age, family history of allergy, breast feeding duration, clinical form of allergy, serum total IgE, or absolute eosinophil count. All peanut sensitive subjects had skin with or without respiratory allergy.
Conclusions: Peanut allergy does not seem to be rare in atopic children in Egypt. Skin prick and specific IgE testing are effective screening tools to determine candidates for peanut oral challenging. Wider scale multicenter population-based studies are needed to assess the prevalence of peanut allergy and its clinical correlates in our country.
The prevalence of peanut allergy is not sufficiently studied in many developing countries including Egypt. Peanut allergy was estimated to affect 0.8% of children and 0.6% of adults in the US, showing a twofold increase over a 5-year period [1,2]. A recent 11 year follow up survey showed that the prevalence of peanut allergy in children in the US in 2008 was 1.4% compared to 0.8% in 2002 and 0.4% in 1997 [3]. In the United Kingdom, the total estimate for clinical peanut allergy was1.5% of 3-4 year-old children [4]. Relevant studies estimated the prevalence of peanut allergy to be 1.34% among primary school children in a Canadian province [5] and 1.15% among 3 year olds in the Australian capital territory with a trend for rise in prevalence between 1997 and 2005 [6]. Peanut was the third most common sensitizing allergen in an Asian community especially in young atopic children with multiple food hypersensitivities and a family history of atopic dermatitis [7].
Studies to address the reasons for increased prevalence and persistence of food allergies, focusing primarily on peanut, have included the hygiene hypothesis; changes in the components of the diet, including antioxidants, fats, and nutrients, such as vitamin D; the use of antacids, resulting in exposure to more intact protein; food processing, such as for peanut roasting and emulsification to produce peanut butter compared with fried or boiled peanut; and extensive delay of oral exposure, thus increasing topical (possibly sensitizing) rather than oral (possibly tolerizing) exposure to food allergens [8,9].
The evaluation of a child with suspected allergy to peanut should include a careful history taking, skinprick testing (SPT), measurement of serum-specific IgE, and, confirmation by an oral food challenge [10,11]. The prevalence of confirmed food allergy based on confirmatory tests is lower than perceived allergy which is based on self report [12,13]. Diagnostic cut-off values for SPT and specific IgE results have improved the diagnosis of food allergy and thereby reduced the need to perform oral food challenges [14]. Overall, a negative peanut SPT has a negative predictive value of more than 95% [15]. The positive predictive value, however, is significantly lower, reaching only 60% in patients with a convincing history of an allergic reaction [16]. Published values on positive SPT and specific IgE values vary from one series to another depending on several factors [12,[17][18][19][20]. Cut-off values do not always have general acceptability and sometimes need to be individualized in the context of clinical impression [21].
There is an impression that peanut allergy is uncommon in Egypt and there are no published data on its incidence or prevalence. We sought to investigate the frequency of peanut sensitization and allergy in a group of atopic Egyptian infants and children in a pilot attempt to uncover its importance as an allergen in our country.
This cross sectional study comprised 100 children diagnosed to have allergic diseases. They were enrolled consecutively after getting informed oral consent was obtained from the parents or care-givers. The study protocol gained approval from the local ethics committee.
Inclusion criteria: -Age at enrollment between one and 18 years.
-A physician made diagnosis of allergic diseases including asthma, allergic rhinitis, urticaria, and eczema.
-Patients who cannot stop antihistamine therapy.
-Extensive skin lesions, scars, positive dermatographism, and very dark skin.
-Treatment with systemic corticosteroids for more than 7 days.
-Other chronic or debilitating illness.
All patients included in the study were subjected to the following:
Detailed history was taken for the possible precipitating factors, peanut consumption (starting age and last consumption), duration of breast feeding, and family history of allergy. Patients were subjected to a general clinical examination, as well as chest, skin, and ENT examination to verify the diagnosis.
Serum total IgE was measured by enzyme linked immunosorbent assay (Genzyme Diagnostics, Medix Biotech Inc, San Carlos, CA, USA). A serum IgE level was considered elevated if it exceeded the highest reference value for age [22]. The value used in correlation analysis was the percentage from the highest normal value for age (patient's actual value/highest normal value for age multiplied by 100) Peanut Specific IgE was measured in children with positive peanut SPT results using the CLA allergenspecific IgE Assay according to the manufacturer's instructions (Hitachi Chemical Diagnostics, Mountain View, California 94043, USA). The concentration of ≥ 15 kUA/L was considered positive.
Complete blood counting was done using an automated cell counter (Coulter MicroDiff 18, Fullerton, CA, USA) and manual differential.
Skin prick test (SPT) was performed for each patient using a commercial peanut allergen extract, positive histamine control, and negative control (Omega Laboratories, Montréal, Canada). First generation short-acting antihistamines were avoided for at least 72 hours and second generation antihistamines were avoided for at least 5 days before testing. The test sites were marked and labeled at least three cm apart to avoid the overlapping of positive skin reactions. The marked site was dropped by the allergen and gently pricked by sterile skin test lancet. Positive and negative control solutions were similarly applied. The patient waited for at least 20 minutes before interpretation of the results. Largest and orthogonal diameters of any resultant wheal and flare were measured. A wheal diameter of 8 mm or greater was considered positive.
Children with proven peanut sensitization were subjected to open oral peanut challenges under close medical observation taking all the precautions needed to treat anaphylaxis. A second informed consent was obtained from the parents or care-givers prior to the challenge. Children were given gradually increasing amounts of roasted peanuts at 30 min intervals and symptoms and physical signs were closely monitored. The total dose of peanut before considering that the challenge was negative was 15 grams of roasted whole peanuts. An open feeding of a larger portion (ageappropriate serving) followed negative challenges and then children were kept under observation for 2 more hours. The cases with negative challenges were supplied with a contact phone number to report any reactions that might develop within the next 24 hours while cases with positive challenge were treated and kept in hospital for 12-24 hours under observation.
Data were analyzed by a standard computer program (SPSS version 13 for Windows, Chicago IL, USA). The mean, standard deviation (SD), median, and interquartile (IQ) range presented the descriptive data. Groups were compared using the students t-test for parametric and the Kruskal-Wallis and Mann-Whitney Z tests for nonparametric data. Fisher's Exact and Chi square (X 2 ) tests were used for comparison of categorical data. Pearson and Spearman coefficient tests were used to correlate the numeric data. For all tests, p values less than 0.05 were considered statistically significant.
The studied sample comprised 55 boys and 45 girls. Their ages ranged between two and 10 years [median (IQR) = 6.28 (4.0); mean (SD) = 6.51 (2.35) years]. The duration of exclusive breast feeding ranged from 3 to 8 months [median (IQR) = 6.00 (2.00); mean (SD) = 5.42 (1.32) months] and the age of stoppage of breast feeding ranged from 9 to 24 months [median (IQR) = 16.00 (7.00); mean (SD) = 16.30 (4.26) months]. None of the subjects gave a history suggestive of peanut allergy and all of them started consuming roasted peanuts before the age of two. The duration since last peanut consumption ranged between one and 15 days [median (IQR) = 4.00 (3.00); mean (SD) = 4.63 (8.94) days]. The diagnoses included bronchial asthma in 63 children, urticaria in 57, allergic rhinitis in 22, atopic dermatitis in five, and history of anaphylaxis due to unknown cause in one patient. Fifty six children had one, 40 had two, and four had three of the aforementioned allergic diseases.
Skin prick testing with peanut extract gave positive results (wheal diameter ≥ 8 mm) in seven children (7%). The specific IgE results of these children confirmed sensitization [range = 17 -24; median (IQR) = 21.0 (4.00); mean (SD) = 20.9 (2.41); 95% CI = 16.9 -24.5 kUA/L]. Six out of the 7 peanut sensitized patients consented for an open oral challenge with roasted whole peanuts. Three out of the six children showed immediate allergic manifestations after consumption (two developed urticaria and respiratory manifestations and one developed urticaria only) and were thus proven to have allergy to peanut. The remaining three developed no symptoms or signs upon oral challenge with peanuts (Table 1). Neither the positive nor negative OFC cases developed late phase reactions to peanut. Two of the peanut allergic children were brothers (Table 1; patients 1 and 2). They had bronchial asthma and attacks of urticaria and the elder one also suffered from allergic rhinitis. The peanut specific IgE levels could not be correlated to any of the clinical or laboratory data of the sensitized children. Children with positive oral challenge were statistically comparable to those with negative results as far as their clinical and laboratory data are concerned.
Seventy percent of the studied sample gave a positive family history of allergy with no statistically significant relation to peanut sensitization. However, the 7 children sensitized to peanut had positive family history of allergic diseases. None of the subjects gave a history suggestive of peanut allergy in the family. Peanut SPT results did not vary according to gender. Positive results were obtained in 5 out of 55 boys and 2 out of 45 girls. Peanut sensitization rates were not influenced by the duration of exclusive breast feeding, age at complete weaning from the breast, last peanut consumption, serum total IgE level, or peripheral blood eosinophil count (Table 2).
The SPT results were not influenced by the target organ affected whether respiratory or cutaneous. Also, peanut sensitization did not vary according to the number of target organs affected in the studied sample (X 2 = 2.714; p = 0.257). Ten children had confirmed allergy to other foods (egg allergy in two, fish in three, cow milk in two, sesame in one, banana in one, and prunes in one); 9 of them were not peanut sensitized while one was sensitized to peanut and allergic to bananas. The relation between peanut sensitization rates and the presence of other food allergies did not reach statistical significance (X 2 = 0.154; p = 0.695).
The wheal diameter of the peanut sensitized children was not correlated to age at complete weaning from the breast, days since last peanut consumption, days since last antihistamine consumption, serum total IgE, serum specific IgE, peripheral blood eosinophil count, or histamine wheal diameter.
The frequency of peanut sensitization in our series was 7%. The sensitized children were subjected to open oral challenges except for one child whose parents did not consent to the challenge due to a past history of anaphylaxis of unknown etiology. Only three patients were proven to be allergic to peanut by oral challenge giving a 3% rate of peanut allergy in the studied sample. The studied sample comprised children with physiciandiagnosed allergy and therefore does not represent the general population. The Peanut specific IgE was only estimated in the 7 children with positive SPT for confirmation of sensitization. It would be worthwhile to estimate its expression in children with negative SPT. However, this was not among the objectives of the current study.
Population based data from some other parts of the world show different rates of sensitization ranging between 3.3% in England [4] and 4.6% among children from the Netherlands [23]. Cross sectional random telephone surveys revealed self reported peanut allergy in 0.6-1.4% in the US [2,3], 0.93-1% in Canada, and 1.5% in the UK [12]. The prevalence of peanut allergy is relatively low in Asian children (0.43% -0.64%) [24].
We considered a SPT wheal size of 8 mm and specific IgE level of 15 kUA/L as our cut off values for predicting peanut allergy [17,25]. Nevertheless, only half of the sensitized children had clinical reactions upon oral challenge. We performed an oral challenge despite the fact that our cut off values may diagnose peanut allergy because none of the sensitized subjects gave a history suggestive of peanut allergy. It seems that the cut off values should be tailored to the levels of consumption in different geographical locations. A recent study from the UK reported a 22.4% prevalence of clinical peanut allergy among sensitized subjects. The authors used the same wheal diameter and specific IgE cut off values that we used [25]. The published peanut specific IgE cut off values are variable ranging between 5 up to 57 kUA/L [26,27]. A wheal diameter of 16 mm was considered to have a positive predictive value of 100% for allergy to peanuts [27].
Although a prior probability from the history is an important starting point in the diagnosis of peanut allergy, none of our subjects gave a history suggestive of peanut allergy and all of them started consuming roasted peanuts before the age of 2 years. Roasted peanut is a popular snack in Egypt and its allergy is not a public concern due to underestimation of its significance even from the health care workers' point of view. The history was reported to be notoriously poor (approximately 30% verified) in identifying causal foods for chronic disorders such as atopic dermatitis [16]. The early peanut consumption by children does not seem to be a risk factor for peanut allergy and was postulated to be even protective [28].
There was no significant relation between the duration of breast feeding and peanut sensitization in the current study. A relevant study noted that neither the maternal peanut consumption during pregnancy and lactation nor the duration of breast feeding was associated with the development of peanut allergy [29]. On the other hand, a recent case-control study assumed that exposure to peanut allergens in utero or through breast milk may increase the risk of developing peanut allergy [30].
A family history of atopy, especially of food allergy, is a good screening test to identify an individual at risk of food allergy and the rate of allergy in a sibling of an allergic person is known to be higher than the rate in the general population. Although we could not demonstrate a statistically significant relation between the family history of allergy and peanut sensitization, all peanut sensitive children came from atopic families. Also, two of the peanut allergic children were brothers.
The current study did not demonstrate a significant gender difference in peanut sensitization but the boys outnumbered girls in the whole sample. Higher prevalence of peanut allergy among male children was previously reported [31].
The peanut sensitization in the present study did not vary according to the site of allergy. However, the 7 peanut sensitized children had physician diagnosed skin allergy combined with respiratory allergy in five. A relevant study on Asian children revealed that most (89.5%) of first reactions featured skin changes and that respiratory and GI symptoms did not occur as the sole manifestation [32]. One child with positive oral challenge in our series had allergy in one target-organ (skin) and a past history of one attack of unexplained anaphylaxis and two children had symptoms in more than one system (skin and respiratory). According to a voluntary registry in the Isle of Wight, about half of all children with peanut allergy had allergic manifestations in one target-organ system, 30% in two, 10% -15% in three, and 1% had symptoms in four systems [33].
Self-reported food allergy is an independent risk factor for potentially fatal childhood asthma [34]. Peanut allergy was reported in 28.8% of children with asthma in a recent investigation [35]. One of our peanut sensitized children was allergic to bananas. Her peanut oral challenge yielded a negative result. The girl was not latex sensitive and the parents did not consent to any further investigations. Evaluation of pollen cross reactivity would have been worthwhile [36].
From this pilot study, it seems that peanut allergy in Egypt is underestimated and that the sensitization rates are even higher. Skin prick and specific IgE testing aided by history are good screening tools to determine candidates for peanut oral food challenge. Peanut allergy can be associated with any clinical form of allergy and the causal relationship needs extensive evaluation. The conclusions are limited by the sample size and study design which targeted physician-diagnosed allergy rather than the general population. Further wider-scale populationbased studies as well as a national registry are needed to be able to outline the real magnitude of peanut allergy and its clinical correlates in our country.
AEC: absolute eosinophil count; BF: breast feeding; IQR: interquartile range; Peanut +: peanut sensitive; Peanut -: not sensitized to peanut. * Peanut sensitization: SPT wheal diameter ≥ 8 and specific IgE ≥ 15 kUA/L
The authors declare that they have no competing interests.
Authors' contributions EH put the study design and coordination, performed the statistical analysis, and drafted the manuscript. GG participated in collection of the study sample and data analysis. AS performed the total and specific IgE assay and differential blood cell counting. AE collected the study sample and performed the skin prick test. All authors read and approved the final manuscript.
Various reactive dyes can elicit occupational asthma in exposed textile industry workers. To date, there has been no report of occupational asthma caused by the red dye Synozol Red-K 3BS (Red-K). Here, we report a 38-year-old male textile worker with occupational asthma and rhinitis induced by inhalation of Red-K. He showed positive responses to Red-K extract on skin-prick testing and serum specific IgE antibodies to Red-K-human serum albumin conjugate were detected using an enzyme-linked immunosorbent assay. A bronchoprovocation test with Red-K extract resulted in significant bronchoconstriction. These findings suggest that the inhalation of the reactive dye Red-K can induce IgE-mediated occupational asthma and rhinitis in exposed workers.
Reactive dyes form covalent bonds with the fibers in textiles. A reactive dye (RD) can act as a hapten 1 and elicit occupational asthma (OA) in exposed workers. 2 Since the first report of RDinduced OA in 1978, 3 the reported prevalence of RD-induced OA in Korea has been 2.5-5.9%. 4,5 RDs were among the most frequent causes of OA in Korea until the 1990s. 6 Of these dyes, Black GR is the most frequent sensitizer; others include Orange 3R, Blue GG, and Green 6B. 4,7,8 There has been no previous report of OA caused by Synozol Red-K 3BS (Red-K). Thus, we report the first case of OA induced by Red-K, confirmed by skinprick testing, bronchial provocation, and immunological testing in a dyer in the textile industry.
The patient was a 38-year-old male non-smoker. He had no history of allergic disease. He had worked in the textile industry for 15 years and had used reactive dyes, such as Synozol Black B 150 (Black B) and Red-K. Nine years after starting his job, he experienced rhinorrhea that progressed to a cough and dyspnea 1 year ago; these were aggravated at work.
At presentation, his blood differential count, serum biochemistry, total IgE (56 kU/L), and chest and paranasal sinus radiographs were normal. The sputum eosinophil count was elevat-
Hyun Jung Jin, Joo-Hee Kim, Jeong-Eun Kim, Young-Min Ye, Hae-Sim Park* Department of Allergy and Rheumatology, Ajou University School of Medicine, Suwon, Korea ed to 84%, while the blood and nasal eosinophil counts were normal (220/μL, 0%). Skin prick testing was positive for mugwort and nettle pollens. Baseline results included a FEV1 of 3.84 L (120.0%) with a FEV/FVC ratio of 90.61% and post-bronchodilator FEV1 of 3.92 L (122.3%) without reversibility. His initial methacholine PC20 was 1.25 mg/mL. We considered that his asthma and rhinitis symptoms may be related to his occupation, especially using RDs. Two RDs used in his factory were Black B (color index name Reactive Black 5) and Red-K. The structures of Black B and Red-K are shown in Fig. 1. Skin-prick testing was performed with extracts of these two RDs (10 mg/mL) and histamine. He had a positive result for Red-K (4 × 4/31 × 24 mm), but a negative result (0 mm) for Black B, when compared with histamine (3 × 3/35 × 29 mm). Serum specific IgE antibodies to two RD-human serum albumin (HSA) conjugates were measured using an enzyme-linked immunosorbent assay (ELISA), as described previously. 9 The positive cut-off value for a high serum specific IgE level was defined as the mean absorbance value plus three standard devia-tions for sera from 18 unexposed, non-atopic, healthy controls. He had a high level of IgE antibody specific to the Red-K-HSA conjugate, but no specific IgE binding was noted with the Black B-HSA conjugate (Fig. 2). To confirm OA, inhalation challenge tests with the two RD extracts were performed. No significant change in FEV1 was noted after the placebo inhalation, while significant bronchoconstriction (a 45.6% fall in FEV1 from the baseline value) with dyspnea and wheezing was noted 10 minutes after inhaling 10 mg/mL of Red-K extract; the response was negative after inhaling Black B extract up to 10 mg/mL (Fig. 3).
Based on these findings, he was diagnosed with OA and rhinitis caused by Red K. We recommended job relocation and the use of an inhaled corticosteroid/long acting β-2 agonist with regular follow-up.
Positive skin-prick tests to extracts of dyes and a positive radioallergosorbent test to RD-bound paper discs suggest that RDs can act as haptens, provoking IgE-mediated hypersensitivity reactions. 10 Both skin-prick testing and the detection of specific IgE to RD-HSA conjugate in serum are useful for screening, diagnosing, and monitoring OA resulting from exposure to RDs. 11 Our patient had positive responses to Red-K extracts on skin- Specific lgE level (OD×1,000) Patient NC 800 600 400 200 0 Specific lgE level (OD×1,000) Patient NC 300 200 100 0 A B prick testing and had high serum specific IgE antibody to Red-K-HSA conjugate. The bronchoprovocation test with Red-K extracts demonstrated immediate bronchoconstriction with the development of asthma symptoms. Thus, we confirmed the first case of OA caused by Red-K, not by Black B, in which an IgE-mediated response was suggested as a major pathogenic mechanism.
Reactive dyes contain a chromogen and reactive functional groups that form irreversible covalent bonds with the amino acid residues of cellulosic fibers. 12 There are many RDs with different reactive functional groups to which carrier proteins bind to induce immune responses. 2,11,13 The vinyl sulfone RDs are major causes of OA and have one or two vinyl sulfone reactive groups. 10,14,15 Black GR, the most frequent sensitizer, has two vinyl sulfone reactive groups, 9 while Red-K has one. However, the allergenicity of an RD can differ according to the carrier protein and the conditions under which the hapten conjugation process occurs. The chromogen components, neoantigenic determinants, and reactive components all contribute to allergenicity. 9 Our patient showed negative responses to skin-prick testing, ELISA, and bronchoprovocation testing for Black B, which is the most frequent sensitizer. Thus, we conclude that he was not sensitized to Black B. He was atopic, although atopy is not a predisposing factor for RD-induced OA. 7 He was a non-smoker, and smoking is a predisposing factor. 8 Park et al. 15 suggested that the duration from symptom onset to diagnosis was the only factor differentiating patients with normal and reduced lung function at diagnosis and that early diagnosis and treatment were essential for a good prognosis. Our patient's disease duration from asthma symptom onset to diagnosis was 1 year. Thus, his condition may improve if he moves to another workplace and takes anti-asthma medications.
In conclusion, we confirmed that Red-K powder inhalation can induce OA, and an IgE mediated response was suggested in the pathogenic mechanism. % Decrease of FEV1 10 0 -10 -20 -30 -40 -50 0 50 100 150 200 Time after challenge (min) Placebo Synozol Black B 150 Synozol Red K 3BS
This study was supported by a grant from the
Introduction: Levels of cerebrospinal fluid (CSF) β-amyloid (Aβ) and Tau proteins change in Alzheimer's disease (AD). We tested if the relationships of these biomarkers with cognitive impairment are linear or non-linear.
We assessed cognitive function and assayed CSF Aβ and Tau biomarkers in 95 non-demented volunteers and 97 AD patients. We then tested non-linearities in their inter-relations. Results: CSF biomarkers related to cognitive function in the non-demented range of cognition, but these relations were weak or absent in the patient range; Aβ 1-40 's relationship was biphasic.
Conclusions: Major biomarker changes precede clinical AD and index cognitive impairment in AD poorly, if at all.
The incidence and prevalence of Alzheimer's disease (AD) double every five years from age 65, to affect over one quarter of people aged over 85 [1,2]. AD pathology can develop long before clinical symptoms [3][4][5]. This means that, ideally, disease-modifying treatments should begin before diagnosis [6]. Consequently, there is much interest in finding biomarkers that can predict the onset of AD [7]. There is also much interest in finding biomarkers to assess treatment effects [8]. Leading candidates for these roles are β-amyloid (Aβ) and Tau proteins in the cerebrospinal fluid (CSF) [9][10][11][12][13][14][15]. To date, nearly all studies of CSF biomarkers have related them to diagnostic categories -AD and mild cognitive impairment (MCI). However, the boundaries of diagnostic categories, particularly MCI, are uncertain [16][17][18]. We, therefore, related CSF biomarkers directly to cognitive test scores.
The form of the relationship between biomarker levels and cognitive scores is important. To predict the onset of AD, a biomarker should relate to cognitive decline pre-clinically. Conversely, to monitor disease progression and treatment response, a biomarker should relate to cognitive level in the range of clinical dementia. These relationships may be such that alterations in biomarker expression may precede, coincide with, or lag behind changes in cognitive status. The simplest assumption would be that the relation of CSF amyloid and Tau biomarkers with cognitive function is linear from cognitive normality through MCI to AD. If so, then these biomarkers might be useful for both the prediction of AD and monitoring its progression. However, most studies that related CSF Aβ and Tau levels to cognitive scores [9,19,20] did not test this assumption of a linear relationship and one study [21] found no relationship in AD patients. Moreover, recent studies that used diagnostic categories have provided evidence of biphasic changes in levels of putative biomarkers between controls, MCI and AD patients in CSF [21,22] and blood [23,24]. These findings support Combrinck's hypothesis [25] that biomarkers may relate non-linearly to cognitive function. Here, we test this hypothesis further by analysing the forms of the relations between CSF biomarkers and cognitive scores.
We have previously reported a biphasic relation between CSF PGE 2 levels and cognitive scores, using an analysis of covariance with polynomial trends [25]. The present study uses change-point analyses with highly robust linear regression to test for non-linearity in cross-sectional relations between cognitive scores and CSF Aβ and Tau moieties.
Participants were volunteers in the Oxford Project To Investigate Memory and Ageing (OPTIMA), a naturalistic longitudinal study of memory and ageing. All OPTI-MA's protocols received prior ethical approval from the local research ethics committee (COREC #1656). OPTIMA is a convenience sample of patients with dementia and non-demented volunteers of similar age. We have described OPTIMA's recruitment and assessment protocol previously [26]. Briefly, at their initial assessment all participants underwent a physical examination, blood tests, CT scan and cognitive assessment using the Cambridge Cognitive examination (CAM-COG) [27]. We also obtained a detailed history from participants and an informant. We invited participants who could give valid consent to undergo a lumbar puncture (LP).
Clinical diagnoses of Alzheimer's disease or Other Dementia Syndromes used the National Institute of Neurological and Communicative Disorders and Stroke (NINCDS) criteria [28]. OPTIMA has a high rate of autopsy acceptance (over 80%). Details of OPTIMA's neuropathological examinations are available elsewhere [29]. We supplanted clinical diagnoses whenever possible by neuropathological diagnoses using Consortium to Establish a Registry for Alzheimer's Disease (CERAD) criteria [30]. The present analysis did not include consideration of Braak staging [3] in the determination of neuropathological diagnosis or its relationship to the variables studied.
The present report includes data from all non-demented controls and AD patients aged over 60 who underwent lumbar puncture and whose CAMCOG data were complete. It excludes patients with clinical or neuropathological diagnoses of non-Alzheimer dementias.
All LPs used standard clinical techniques [31]. Most took place in the late morning. LP has a low risk of adverse effects in our cohort [31]. We collected the CSF samples into polystyrene tubes.
We centrifuged CSF samples for 10 minutes at 4°C at 1,000 g to remove cells and stored the supernatant in aliquots of 0.5 ml in polypropylene tubes at -70°C. We assayed levels of CSF Aβ and Tau moieties using commercial kits. No sample underwent any freeze-thaw cycles between collection and these assays.
Aβ 1-40 was measured in the CSF with a human Aβ 1-40 Colorimetric solid phase sandwich Enzyme Linked Immuno-Sorbent Assay (ELISA) kit (catalogue # KHB3482, BioSource International, Camarillo, CA, USA) following the manufacturer's recommendations. This assay employs a mouse monoclonal antibody specific for the N-terminal half of Aβ 1-40 as capture and a rabbit anti-Aβ 1-40 neo-epitope (secondary antibody). The detection antibody consisted of a secondary anti-rabbit IgG:horse radish peroxidase (HRP) conjugate. HRP catalyzes the formation of a chromophore, tetramethylbenzidine (TMB), which was measured at 450 nm. A total of 100 μL of the sample (CSF diluted 1:20 in assay buffer) was used in this assay. The standards were provided in the BioSource assay kit and they ranged from 15.6 to 1000 pg/mL.
Aβ 1-42 was measured with Innotest™ Aβ 1-42 ELISA kit (Innogenetics Inc., Cat. #80040, Ghent, Belgium) following the manufacturer's recommendations with some modifications. Aβ 1-42 present in human CSF samples was first captured with a mouse monoclonal antibody specific for the C-terminal half of Aβ The detection system employs an N-terminal specific biotinylated mouse monoclonal antibody and a secondary conjugate made of HRP labeled strepavidin. The HRP is used to convert tetramethyl benzidine to a chromophore which is quantitatively measured at 450 nm. A total of 100 μL of the sample (CSF Diluted 1:3 with Sample Diluent) was used in each reaction. Aβ 1-42 standard was purchased from American Peptide (Sunnyvale, CA, USA) and the concentration was determined by amino-acid analysis. Standard concentrations in the assay ranged from 5.45 to 350 pg/mL.
Total Tau (t-Tau) expression was measured with a human Tau (hTAU AG Innotest™) ELISA kit (Innogenetics Inc., catalogue number 80226, Ghent, Belgium) following the manufacturer's recommendations. The analyte was first captured with a monoclonal antibody specific for all isoforms of Tau, and then subsequently bound by two biotinylated Tau-specific antibodies. The final detection was performed by peroxidase-labeled streptavidin. A total of 25 μL of the sample was tested undiluted. The standards were supplied with the kit and ranged from 37.5 to 1200 pg/mL.
Phosphorylated Tau-181 (pTau-181) was measured with the Phospho-TAU ( 181P ) Innotest™ ELISA kit (Innogenetics Inc., catalogue number 80062, Ghent, Belgium), following the manufacturer's recommendations.
The analyte was first captured with an antibody specific for all isoforms of Tau and then detected with a second detection antibody which specifically detects Tau molecules phosphorylated at threonine 181 (phospho-tau-181). A 75 μL sample was tested undiluted. The standards were supplied with the kit and ranged from 15.6 to 500 pg/mL.
All assays were analytically validated (inter-and intraassay precision, freeze/thaw stability, linearity, spike recovery, and sensitivity). In addition, quality control samples (low, medium, and high) were run on all plates and were used as part of the run acceptance criteria. All analytes were found to be stable (<20% change) after three freeze-thaw cycles. All sample analyses were performed in duplicate.
For the Aβ 1-40 assay, the intra-and inter-assay percent coefficient of variation (% CV) ranged from 4.1% to 7.6% and 9.4% to 12.5%, respectively. The spike recovery was determined to be 105 to 114% and the lower limit of reliable quantitation was 17.8 pg/mL.
For the Aβ 1-42 assay, the intra-and inter-assay % CV ranged from 3.7 to 4.7% and 5.9 to 7.6%, respectively. The spike recovery was determined to be 70 to 109% and the lower limit of reliable quantitation was 24.5 pg/mL.
For the t-Tau assay, the intra-and inter-assay % CV ranged from 3.8 to 9.0% and 6.2 to 7.2%, respectively. The spike recovery was determined to be 105 to 116% and the lower limit of detection was 37.5 pg/mL.
For the p-Tau assay, the intra-and inter-assay % CV ranged from 1.3 to 2.2% and from 4.9 to 5.3%, respectively. The spike recovery was determined to be 84 to 89% and the lower limit of detection was 15.6 pg/mL.
All statistical analyses used the open-source statistical programming language 'R' [32]. We used the Wilcoxon-Mann-Whitney signed rank test and Pearson's χ 2 to compare the demographic characteristics of the patient and non-demented groups.
We first performed omnibus tests for non-linearity using multivariate analysis of variance (MANOVA) [33]. The first MANOVA tested the dependence of all four biomarkers on age, gender, storage time, CAMCOG score and assay group. A second MANOVA included all the above variables and a second CAMCOG term representing an inflection in the relationship with cognitive function, corresponding with our previous work [25]. That study showed an inflection in the relation between CSF PGE 2 levels and CAMCOG learning sub-scale scores at a score of 11, the lowest limit of the non-demented range. In the present study, we first ascertained the total CAMCOG score that corresponded best with a CAMCOG learning sub-scale score of 11, then programmed an inflection at that total score (89).
The strategy of pre-defining an inflection point for all CSF biomarkers may be sub-optimal if different biomarkers show different inflection points. Therefore, having obtained evidence of a significant inflection in our MANOVA (see Results), we tested for change-points in the relations between each individual biomarker level and CAMCOG score. These change-point analyses were of two kinds. First, we tested for changes in the overall level of each biomarker in relation to CAMCOG. These analyses used the "Fstats" and "breakpoints" procedures for testing structural changes in linear regression models [34]. Second, we tested for change-points in the slope of the relation between each biomarker and CAMCOG. These analyses used the "linearSegmentation" procedure for piecewise linear segmentation of a time series [35]. This procedure requires data in a continuous equallyspaced series with one observation at each point. We converted the CAMCOG scores to a series of this kind by ranking them and randomly splitting ties. Since the order of the random splits may affect the result of the change-point analysis, we repeated the analysis 1,000 times with a different random seed to split ties in each repetition. Each analysis progressively increased the window size and tolerance angle in each repetition, until the procedure generated a single change-point for each biomarker. We recorded the change-points for each biomarker in each of the 1000 repetitions. Finally, we extracted the median of these 1000 estimates. Hence, for each biomarker we derived two change-points: 1) CAMCOG scores at which the biomarker changed its level, and 2) change in the slope of the relation between each biomarker and CAMCOG. These two changepoints were very similar for each biomarker, so we used their mean in further analyses.
We assessed the validity of the change-point for each biomarker in three ways. First, we visually compared the change-point and model-free robust Lowess fits (Figures 1, 2, 3, 4). Second, we compared the variances and medians of each biomarker on each side of its change-point using simple variance ratios and Wilcoxon-Mann-Whitney (WMW) tests. Third, we compared the relationship of the biomarker with CAMCOG on each side of its changepoint using Spearman's ρ and robust linear modelling (rlm) [36]. The rlms used highly robust M-estimation (via the 'MM' option) with a breakdown point of 0.5 and 95% relative efficiency at the normal.
A total of 192 participants fulfilled the inclusion criteria; 97 received a diagnosis of AD and 25 of the nondemented volunteers received a designation of MCI.
We obtained post mortem neuropathological confirmation of the diagnosis in 90% of the AD patients (88/97) and a third of the non-demented goup (30/95). All participants were white West Europeans. Table 1 shows their demographic and clinical characteristics. The age and gender distributions of the non-demented participants and AD patients did not differ (age: WMW P = 0.11; gender: Fisher exact P = 0.15). Compared with non-demented participants, AD patients had lower • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • cognitive impairment (CAMCOG score) CSF -Aβ 1-42 100 80 60 40 20 0 200 600 1,000 1,400 • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • type of fit diagnostic category linear regression change-point model robust Lowess line Figure 2 The dependence of CSF Ab 1-42 levels on CAMCOG score. All details as for Figure 1.
• • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • 50 100 150 200 cognitive impairment (CAMCOG score) P -τ 1 0 0 8 0 6 0 4 0 2 0 0 • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • • type of fit diagnostic category linear regression change-point model robust Lowess line Figure 4 The dependence of CSF phospho-Tau levels on CAMCOG score. All details as for Figure 1.
MMSE and CAMCOG scores and shorter follow-up (all WMW P-values < 0.001). We have provided further clinical details in Additional file 1. The role of the diagnostic procedures in the present study was to ensure that the study group did not include participants with non-Alzheimer dementias. Therefore, neither clinical nor neuropathological diagnostic categorisations played any part in the design or interpretation of our changepoint analyses, which concern only the forms of the relations between CSF biomarkers and CAMCOG scores.
The inflection point in the total CAMCOG score corresponding to our previous finding [an inflection at a CAMCOG learning sub-scale score of 11 -see 25] was 89. Biomarker levels showed different levels and their relations with CAMCOG scores showed different slopes above and below total scores of 89 (levels: F = 11.5, 4/167 df, P < 0.0001. Slopes: F = 3.31, 4/167 df, P = 0.012) (see Table 2 and Figures 1, 2, 3, 4).
CSF Aβ 1-40 levels showed biphasic dependence on CAMCOG scores, with the change-point at a CAMCOG of 90 (Figure 1). The change-point model fitted the data better than the simpler linear model (F = 3.27, 2/187df, P = 0.04). There was a positive relation of Aβ 1-40 with CAMCOG scores above the change point (rlm: t = 2.65, 187 df, P = 0.009) but a negative relation below it (rlm: t = -3.25, 187 df, P = 0.001; Table 2). Congruent with this, the Spearman's ρs on each side of the change-point were opposite in sign and both differed from zero (Table 2; Figure 1). The median and variance of Aβ 1-40 levels were greater above the change-point than below it (median: WMW P < 0.001; variance: F = 1.95; 96/96 df; P < 0.0001; see Table 2). In summary, the levels and variability of CSF Aβ 1 to 0 and the sign of its relation with CAMCOG differed above and below the changepoint.
CSF Aβ 1-42 levels showed monotonic but non-linear dependence on CAMCOG scores with the change-point at a CAMCOG of 90 (Figure 1b). The change-point model fitted better than the simpler linear model (F = 11.2, 2/178 df, P < 0.001). There was a marked negative relation of Aβ 1-42 with CAMCOG scores above the change point (rlm: t = -4.29, 178 df, P < 0.001), but no relation below it (rlm: t = -1.17, 178 df, P = 0.24) (Table 2; Figure 2). Congruent with this, Spearman's ρs showed a negative relation only with CAMCOG scores above the change point (see Table 2). The median and variance of Aβ 1-42 levels above the change-point were greater than below it (median: WMW P < 0.001; variance: F = 2.11; 92/100 df; P < 0.001). In summary, the levels and variability of CSF Aβ 1-42 and the slope of its relation with CAMCOG differed above and below the change-point.
CSF Tau levels showed monotonic but non-linear dependence on CAMCOG scores with the change point at a CAMCOG of 94. The change-point model fitted better than the simple linear model (F = 5.39, 2/187 df, P = 0.005). Tau levels related to CAMCOG scores more strongly above the change-point (rlm: t = 2.01, 187 df, The values for Age, MMSE, CAMCOG and Follow-up are medians and inter-quartile ranges.
Table 2 Parameters of the CSF biomarkers and their relations to cognitive impairment, above and below their change-points
CSF Biomarker Ab 1-40 Ab 1-42 Tau Phospho-Tau Cf change-point Above Below Above Below Above Below Above Below Median 5,052 3,937** 682 350** 259 568** 50 80** Variance 3,857,482 1,854,983** 36,846 71,482** 176,946 73,418** 1,700 1,000** Spearman's ρ 0.21* -0.23* -0.24* -0.12 0.34** 0.28** 0.32** 0.16 Robust slope 93** -15** -0.048** -0.002 0.049* 0.004** 0.041** 0.004**
The values in the first row are the raw biomarker levels (pg/mL); the second row shows the variance of the median-standardised levels; the third row shows Spearman's ρs for the rank correlations of biomarker levels with CAMCOG scores; the fourth row shows the slopes from the robust regressions of medianstandardised biomarker levels. For each biomarker, the "Above" and "Below" columns shows values for CAMCOG scores above and below the change-point, respectively. Note that the correlations and slopes take cognitive impairment as the x-axis, which is the reverse of CAMCOG scores. Key:-( * ) : 0.10 >P > 0.05; *: P < 0.01; **: P < 0.01. For medians and variances, probabilities refer to the comparison of values above and below the change-point; for Spearman's ρ and slopes they refer to comparison of values with zero. P = 0.046) than below it (t = 4.46, 187 df, P < 0.001) (Table 2; Figure 3). Spearman's ρs showed positive relations with cognitive impairment both above and below the change point (see Table 2). The median and variance of CSF Tau levels in the CAMCOG range below the change-point were much greater than above it (median: WMW P < 0.001; variance: F = 7.06; 119/73 df; P < 0.001). In summary, the levels and variability of CSF Tau and the slope of its relation with CAMCOG differed above and below the change-point.
CSF phospho-Tau levels showed monotonic but nonlinear dependence on CAMCOG scores, with the change point at a CAMCOG of 93. The change-point model fitted better than the simple linear model (F = 10.2, 2/187 df, P < 0.001). phospho-Tau levels related directly to CAMCOG scores above the change-point (t = 2.39, 187 df, P = 0.018), and below it (t = 2.48, 187 df, P = 0.014) (Table 2; Figure 4). The Spearman's ρs showed a positive relation with cognitive impairment only in the CAMCOG range above the change point (see Table 2). The median and variance of CSF phospho-Tau levels in the CAMCOG range below the change-point were greater than above it (median: WMW P < 0.001; variance: F = 2.82; 113/79 df; P < 0.001). In summary, the levels and variability of CSF phospho-Tau and the slope of its relation with CAM-COG differed above and below the change-point.
CSF β-amyloid and Tau peptides showed distinctive non-linear relationships with cognitive scores. These relationships were strongest in the "normal" CAMCOG range and every biomarker's change-point was near the lower boundary of this range. These non-linearities are inconsistent with a simple linear relationship between central amyloid and Tau levels and cognitive impairment which underlies the view that modulating central amyloid and Tau may prevent cognitive decline in dementia [37,38]. Instead, they suggest that most changes in central amyloid and Tau precede the development of clinical dementia [39].
The findings of our change-point analyses are consistent with published data. First, our re-analysis of data shown in Figure 2 in Maruyama [19] shows a change-point in the relationship between MMSE scores and CSF Aβ 1-42 . This resembles our present finding (see our Figure 2) in that there is no relation between MMSE and Aβ 1-42 across MMSE scores below 20 (t = 0.21, NS), but a significant relation across MMSE scores above 20 (t = 2.08, 76 df, P = 0.041). Second, our findings that relations between CSF biomarkers and CAMCOG were weak or absent below the change-points parallel the observations of Vemuri [40] in AD patients; the stronger relations that we found above the change-points may also parallel Vemuri's report [40] that biomarker levels differed in their Controls and people with MCI. Third, our finding of an inverted-U relation between Aβ 1-40 and CAMCOG scores parallels the report that Aβ 1-40 levels were high in MCI patients who progressed to AD [21] The change-points that we defined are near the bottom of the range of CAMCOG scores in our nondemented volunteers. Hence, our findings are consistent with many previous reports of higher levels of CSF Tau moieties and lower levels of CSF amyloid moieties in AD [22,25,41,42]. However, we prefer our change-point categories to the diagnostic classification because they derive from a model-free method, whereas the definition of Alzheimer's disease can vary [43,44]. Moreover, our results indicate that neither the diagnostic classification, nor the diagnostically agnostic linear regression fully describes the relations between cognitive function and CSF biomarkers. Figures 1, 2, 3, 4 each show the four models that assume that the relations between CSF biomarker levels and cognitive function are either (a) simply linear (green dotted lines) or (b) simply categorical (blue dot-dash lines). In each case these simple relations fit the data quite well, but the non-linear relations based on the change-point analyses (c -black solid lines) fit better and more closely resemble the model-free robust Lowess fits (d -red dashed lines). These improvements in fit are not large, but their theoretical importance outweighs their actual size, as we discuss below.
CSF Aβ 1-40 showed biphasic dependence on cognitive scores (Figure 1), with maximal levels near the lower limit of our non-demented participants' CAMCOG scores. This is consistent with the findings that MCI patients who progressed to AD had high CSF Aβ 1-40 [21] and CSF BACE activity is higher in MCI than in either controls or AD patients [22]. It also fits with findings of biphasic changes in plasma amyloid levels [24]. Zhong et al. [22] interpreted their findings as supporting the amyloid cascade hypothesis. However, the biphasic relationship of Aβ 1-40 admits at least two other overlapping interpretations. First, it is consistent with Combrinck's hypothesis [25] that a pathogenic mechanism may 'burn out' early in AD. Second, it is consistent with the hypothesis that Aβ may be protective [45,46], and AD occurs only when pathogenic mechanisms overwhelm the amyloid response. Further studies are necessary to explore these three possibilities.
CSF Aβ 1-42 showed an overall inverse relationship with the degree of cognitive impairment (Figure 2). This fits with many reports of low Aβ 1-42 in AD patients [40][41][42]. Our findings extend those reports by showing that Aβ 1-42 depends on cognitive level only in the range of nondemented volunteers' CAMCOG scores, not in the patients' range. The absence of any relation between CSF Aβ 1-42 and CAMCOG scores below the change-point (in the patient range) mirrors the recent report of Vemuri et al. [40]. Together, these results indicate that low Aβ 1-42 may provide an early marker of likely progression to AD, rather than an index of the severity of pathology in established AD. The contrast between the monotonic relation of Aβ 1-42 with cognitive level and the biphasic relation of Aβ 1-40 is striking (compare Figures 1 and 2). This contrast fits with observations that a low ratio of Aβ 1-42 /Aβ 1-40 associates with imminent risk of mild cognitive impairment and Alzheimer's disease [47][48][49]. Together, these findings highlight a need for further studies to explain why the ratio of Aβ 1-42 /Aβ 1-40 may vary [50].
We found high CSF Tau moieties in AD patients [51]. Our results extend earlier reports by showing that, like Aβ 1-42 , Tau and phospho-Tau levels relate to CAMCOG mainly above the change-point (in the non-demented range of scores). Again, paralleling Vemuri et al.'s report [40], we found no simple relation between CSF phospho-Tau and CAMCOG scores below the change-point, in the patient range (Figures 3, 4). Together, these results reinforce the view that phospho-Tau may provide an early marker of likely progression to AD, but not an index of the severity of pathology in established AD. The findings that Aβ 1-42 and phospho-Tau show opposite relations with cognitive scores within the non-demented range fits with reports that the ratio of Aβ 1-42 /Tau may be a sensitive indicator of progression to AD [41,52,53].
The variances of the CSF biomarkers differed above and below their change-points. The simplest explanation of this is that the variance related to the mean, as often occurs in biological variables. It may also represent individual differences [54], or variations in pathogenic processes or cognitive profiles [55]. For example, the heteroscedasticity may possibly relate to ApoE status. We did not test relations of ApoE alleles with changepoints, since we had no a priori hypothesis about this. Further studies should address this potentially important possibility. Whatever its explanation, the heteroscedasticity in every biomarker indicates a need for caution when interpreting biomarker levels in patient and control groups, or in individuals [56]. For now, we note that our use of highly robust linear modelling with a breakdown point of 0.5 [36] guards against potential bias in our analyses due to heteroscedasticity.
We collected CSF samples in polystyrene tubes and stored them in polypropylene tubes. Both kinds of tube can adsorb large amounts of biomarker molecules [57,58]. If such adsorption saturates asymptotically, and so depends non-linearly on initial biomarker concentration, then this might possibly contribute to both the heteroscedasticity and the non-linear relations between biomarker levels and cognitive function that we found. However, we think this unlikely, as follows. If adsorption of biomarkers tends to cause floor effects in their apparent levels, then cognitive function might relate to biomarker levels only when they exceed this artefactual floor. Such floor effects could possibly explain our finding that CAMCOG scores did not relate to Aβ 1-42 in the patient range of CAMCOG scores (because Aβ 1-42 levels are lowest in this range and so most susceptible to adsorption-mediated floor effects). However, it cannot simultaneously explain the inverted-U biphasic relation of CAMCOG with Aβ 1-40 ; nor can it explain the direct relationships of the CAMCOG scores with tau and p-tau in the non-demented range of CAMCOG scores (where lower Tau/p-Tau levels are potentially more susceptible to adsorption-mediated floor effects). Even for Aβ 1-42 , this account appears tenuous, in view of Bjerke's observation that detergent treatment released similar percentages of Aβ 1-42 from adsorption in both control and patient samples [57]. In summary, we think it unlikely that measurement technicalities can account for the non-linear dependence of CSF biomarkers on cognitive function that we observed.
We related CSF biomarker levels to raw CAMCOG scores. Hence, apparent non-linearities in the relationships may in part reflect non-linearities in the metric properties of the CAMCOG (for example [59]). In particular, the narrow range of CAMCOG scores above the change-points may contribute to the differences in the slopes of their relations with biomarkers, since CAM-COG scores here may relate less strongly to true cognitive ability. However, this cannot account for (a) the inverted-U relation of Aβ 1-40 with CAMCOG scores; nor for (b) the differences in Spearman's rank correlation coefficients (which is independent of the metric) for biomarkers and CAMCOG scores above and below the change-points; nor for (c) the differences in variance of biomarkers above and below the change-points. Therefore, while the slopes of the relationships that we found on each side of the change-points may vary under non-linear transformations of the CAMCOG scores, it seems unlikely that our use of raw CAMCOG scores can account for the existence of the change-points.
The main limitation of our change-point analyses is that they used only cross-sectional data. Their use of CAM-COG scores as the metric for dementia removes time from the analysis of progression. The absolute cognitive level is clinically meaningful regardless of age or the duration of symptoms. Therefore, using it as the metric for progression of pathological mechanisms may be preferable to using time. We know that individuals' CAM-COG scores can decline over time through the normal and patient ranges [59]. Hence, it is tempting to view the cross-sectional dependence of CSF biomarkers on CAMCOG scores as a model of an individual's likely progression over time. However, the heteroscedasticity that we observed (see above) means that individuals might show important variations from this model. Consequently, it would be inappropriate at this stage to conclude that the non-linear cross-sectional relationships that we observed can indicate which non-demented people will progress to AD, or which AD patients will decline faster. Conversely, the cognitive stability of many non-demented participants implies that the crosssectional relations of CSF biomarkers with their cognitive function may index long-term adaptations to factors that pre-dispose to AD [60]. Alternatively, these crosssectional relationships may reflect Vemuri et al.'s [40] observation that levels of biomarkers differed between controls and MCI groups, since we did not distinguish these. MCI may be stable, remit, or progress to AD [17,18]. Hence, the cross-sectional relations of biomarkers with CAMCOG scores in the non-demented range may reflect a mix of long-term adaptations and of vulnerabilities to progression. Further longitudinal studies relating CSF biomarker levels to cognitive function are necessary to define more precisely the links between CSF biomarkers and the putative primary pathogenic processes of AD.
The generalizabilty of our study may be comparable with other reports of CSF biomarkers. Participants in all such studies are partly self-selected, both for entry to the cohort and for consenting to LP. OPTIMA is a convenience cohort from a relatively small geographical area in and around Oxford. Our study group was relatively homogeneous and most non-demented volunteers were cognitively stable, with few converting to AD, despite our long follow-up. Together, these two considerations indicate that our study group is unlikely to be representative of the general population. Even so, the consistency of our results with previous reports, both with regard to differences in CSF biomarkers between patients and non-demented controls and to non-linearity [cf. [21][22][23][24]], even in a Japanese sample [19], implies that our findings may reflect general phenomena. Two further aspects of our study may improve its generalizability. First, we confirmed the diagnosis of AD and excluded non-Alzheimer dementias via neuropathological examination in most patients (though people who consent to autopsy are non-representative of the general population [61]). Second, we related CSF biomarker levels directly to cognitive scores. Neuropathological designations and cognitive test scores may provide a firmer basis for generalization than clinical diagnoses, whose boundaries are uncertain [16][17][18]43,44]. Overall, then, our study compares favourably with other reports in this field.
The change-points in relation between cognitive function and all biomarkers were in the "normal" range of CAMCOG scores. This is consistent with increasing evidence that many non-demented older people have some AD pathology [3][4][5], which implies that diseasemodifying treatments may be maximally effective for prophylaxis before clinical dementia occurs [6,62]. Hence our results could help explain recent reports that anti-amyloid treatments, such as tarenflurbil and anti-amyloid immunotherapy, had no effect in established AD [63][64][65], assuming that these treatments had anti-amyloid effects at the doses tested. The biphasic relationship we observed for Aβ 1-40 is also consistent with the possibility that it may contribute to neuronal protection or maintenance [45,60,66,67]. If this were ultimately to prove to be the case, then BACE inhibition at critical stages might be detrimental. Whereas the field focuses most on Aβ 1-42 , this possibility suggests that more attention should be given to understanding the physiological role(s) of Aβ 1-40 .
We are grateful to the participants and nursing staff at OPTIMA for undertaking the lumbar punctures.
Competing interests JS, AD, OL, and WP are employees of and own stock in Merck Sharp & Dohme Corp. and GKW is co-Editor-in-Chief of Alzheimer's Research & Therapy. The other authors declare that they have no competing interests.
Authors' contributions JHW conceived and performed the modelling and analysis. JS, AD, OL and WP were responsible for the biomarker studies. All authors contributed to the interpretation of the results and drafting of the final report.
Additional file 1: Supplementary methods. A document outlining additional diagnostic considerations.
Background-Nonobstructive hypertrophic cardiomyopathy (nHCM) is often associated with reduced exercise capacity despite hyperdynamic systolic function as measured by left ventricular ejection fraction. We sought to examine the importance of left ventricular strain, twist, and untwist as predictors of exercise capacity in nHCM patients.
Methods-Fifty-six nHCM patients (31 male and mean age of 52 years) and 43 age-and gendermatched controls were enrolled. We measured peak oxygen consumption (peak VO 2 ) and acquired standard echocardiographic images in all participants. Two-dimensional speckle tracking was applied to measure rotation, twist, untwist rate, strain, and strain rate.
Results-The nHCM patients exhibited marked exercise limitation compared with controls (peak VO 2 23.28 ± 6.31 vs 37.70 ± 7.99 mL/[kg min], P < .0001). Left ventricular ejection fraction in nHCM patients and controls was similar (62.76% ± 9.05% vs 62.48% ± 5.82%, P = .86). Longitudinal, radial, and circumferential strain and strain rate were all significantly reduced in nHCM patients compared with controls. There was a significant delay in 25% of untwist in nHCM compared with controls. Both systolic and diastolic apical rotation rates were lower in nHCM patients. Longitudinal systolic and diastolic strain rate correlated significantly with peak VO 2
Conclusions-In nHCM patients, there are widespread abnormalities of both systolic and diastolic function. Reduced strain and delayed untwist contribute significantly to exercise limitation in nHCM patients.
Patients with nonobstructive hypertrophic cardiomyopathy (nHCM) often complain of breathlessness and/or fatigue, and most have reduced exercise capacity. Patients with nHCM typically exhibit diastolic dysfunction, which is thought to contribute significantly to the genesis of breathlessness in these patients. Active relaxation is slowed and filling rate is typically diminished at rest in nHCM patients. Furthermore, we previously showed that there is frequently a failure of the normal increase in rate of active relaxation (as measured by time to peak filling) on exercise; indeed, in some patients, active relaxation paradoxically slowed on exertion.
Resting echocardiographic Doppler parameters of left ventricle (LV) filling have been used to measure diastolic function. However, these parameters are dependent on loading conditions. The LV twists during systole as a result of counterclockwise rotation of the apex and clockwise rotation of the base with corresponding untwisting during diastole. The resulting untwisting may represent a useful marker of diastolic function. The velocity of LV untwisting has been correlated with invasive measurements of LV relaxation under a variety of loading and inotropic conditions in animal models. Dong et al confirmed that LV untwisting is a measure of LV relaxation that is independent of preload as reflected in left atrial (LA) pressure.
The development of speckle tracking echocardiography (STE) has allowed LV strain, twisting, and untwisting rates to be measured relatively easily with reasonably high temporal resolution. In this study, we used STE to measure LV strain, strain rate, and twist and untwist rates in nHCM patients and controls; and we sought to assess their relationship with exercise capacity.
This study was funded by a British Heart Foundation project grant (PG/05/087). The authors are solely responsible for the design and conduct of this study, all study analyses, the drafting and editing of the paper, and its final contents.
Fifty-six nHCM patients and 43 age-and gender-matched controls were included in this study, all of whom provided written informed consent. The study was approved by the local ethics committee, and the investigations conform to the principles outlined in the Declaration of Helsinki. All study participants had electrocardiogram (ECG), echocardiogram, and cardiopulmonary exercise test.
All HCM patients were in sinus rhythm and fulfilled conventional echocardiographic criteria for the diagnosis of HCM (LV wall thickness >1.5 cm in the absence of another cause for hypertrophy and without cavity dilatation). All HCM patients underwent assessment of LV outflow tract obstruction gradient, and those with a resting or provocable gradient (on Valsalva maneuver) >30 mm Hg were excluded. Those patients were excluded because LV outflow tract obstruction is an important cause of symptoms and exercise limitation. All controls had no history or symptoms of any cardiovascular disease, with normal ECG and echocardiogram (LV ejection fraction [LVEF] ≥55%).
This was performed using a Schiller (Baar, Switzerland) CS-200 Ergo-Spiro exercise machine that was calibrated before every study. Subjects underwent spirometry, and this was followed by symptom-limited erect treadmill exercise testing using a standard ramp protocol with simultaneous respiratory gas analysis. Samplings of expired gases were performed continuously, and data were expressed as 30-second means. Minute ventilation, oxygen consumption, carbon dioxide production, and respiratory exchange ratio were obtained. Peak oxygen consumption (peak VO 2 ) was defined as the highest VO 2 achieved during exercise and was expressed in milliliters per minute per kilogram. Blood pressure and ECG were monitored throughout. Participants were encouraged to exercise to exhaustion with a minimal requirement of respiratory exchange ratio >1.
Echocardiography was performed with participants in the left lateral decubitus position with a GE (Horten, Norway) Vivid 7 echocardiographic machine and a 2.5-MHz transducer. Measurements were averaged for 3 beats and were stored digitally and analyzed off-line. Resting scans were acquired in standard long and short parasternal basal level (at mitral valve), papillary muscle level, and apical level (distal LV cavity where papillary muscle is not visible) as previously described as well as apical 4-chamber and apical 2-chamber axis. Left ventricular volumes were obtained by biplane echocardiography, and LVEF was derived from a modified Simpson formula. A pulse wave Doppler sample volume was placed at the mitral valve tips to record 3 cardiac cycles. Mitral annulus velocities (pulse wave tissue Doppler imaging [PW-TDI]) were recorded from basal anterolateral and basal inferoseptal segment in apical 4chamber view. Left atrial volumes were measured by area length method from apical 2 and 4 chambers as previously described and indexed to body surface area to derive LA volume index (LAVI). Left ventricular mass was measured by area length method and indexed to body surface area to derive LV mass index as previously described. The greatest thickness measured in the LV wall at any short-axis parasternal view was considered to represent maximal LV wall thickness (MWT) in nHCM patients.
Speckle tracking echocardiography (STE) was measured using Echopac workstation (version 4.2.0) (GE Vingmed Ultrasound AS, Horten, Norway). In this speckle tracking method, the displacement of speckles of myocardium in each spot was analyzed and tracked from frame to frame. We selected the best-quality digital 2-dimensional image cardiac cycle, and the LV endocardium was traced at end systole. The region of interest width was adjusted as required to fit the wall thickness. The software package then automatically tracked the motion through the rest of the cardiac cycle. Adequate tracking was verified in real time. A tracking score of ≤2.5 was accepted with frame rate between 60 and 100 Hz. Left ventricular peak strain and strain rate were measured in all segments and then averaged, representing the entire myocardium. The basal and apical LV global rotation (rot) and LV rotation rate (rot-r) STE data were measured from basal and apical short-axis view, respectively; exported to a spreadsheet program (Excel 2003, Microsoft Corp, Seattle, WA); and then exported to DPlot Graph Software (2001-2008 by HydeSoft Computing, Inc, Vicksburg, MS) to calculate LV twisting and untwisting rate. Definitions of twist and untwist are shown in (Figure 1). Interobserver and intraobserver reproducibility and variability of STE twist, untwist, and strain measurements were previously validated by our group.
Data were analyzed using SPSS version 15.0 for Windows (SPSS Inc, Chicago, IL) and Microsoft Office Excel 2007, and expressed as mean ± SD. Comparison of variables between nHCM patients and controls was by the unpaired Student t test (2-tailed) if variables were normally distributed and the Mann-Whitney U test if the data were nonnormally distributed. A difference of P < .05 was taken to indicate statistical significance. Pearson correlation coefficient was used to examine correlations between exercise capacity and echocardiographic parameters in nHCM patients. Multivariate linear regression stepwise analysis was then performed by entering into model variables that were considered significant on univariate analysis to determine independent predictors of exercise capacity.
The clinical characteristics and cardiopulmonary exercise test results of nHCM patients and controls are shown in Table I. Both groups were well matched with respect to age and gender. The nHCM patients exhibited marked exercise limitation as compared with controls, whereas mean EF was comparable in both groups. Heart rate was lower in the nHCM group as compared with controls because of rate-limiting medication (β-blockers or/and calcium-channel blockers).
Standard, Doppler, and PW-TDI echocardiographic findings of nHCM and controls groups are listed in Table II. Transmitral Doppler measurements showed no significant difference in early diastolic peak (peak E), late diastolic peak (peak A), and E/A ratio in nHCM versus controls (Table II). However, PW-TDI measurements of myocardial peak systolic velocity (Sm), myocardial peak early diastolic velocity (Em), myocardial peak late diastolic velocity (Am), and E/Em of the inferoseptal, anterolateral, and average basal segments were significantly lower in nHCM patients compared with controls (Table II).
The STE measurements of strain and strain rate are shown in Table III. Longitudinal strain and strain rate (SrL) during both systole and diastole were all significantly reduced in nHCM patients as compared with controls. Similarly, radial strain and strain rate were also significantly lower in nHCM patients than in controls. Circumferential strain and diastolic circumferential strain rate (SrC) were also significantly lower in nHCM patients as compared with controls. There was a nonsignificant trend to lower systolic SrC in nHCM patients compared with controls.
The STE measurements of LV rotation, twist, and untwist in both groups are shown in Table IV. There was no significant difference in the LV ability to twist in systole and untwist in diastole between nHCM and controls. However, 25% untwist was significantly delayed in nHCM versus controls. Unlike the base, there was a significant reduction in apical rotation rates during systole and diastole in nHCM patients as compared with controls.
In nHCM patients, exercise capacity (peak VO 2 ) was independent of LV mass index (r = 0.01, P = .97) (Figure 2, A), MWT (r = 0.04, P = .8) (Figure 2, B), or LAVI (r = -0.08, P = .6) (Figure 2, C). However, it significantly correlated with the following STE parameters: SrL during systole (systolic peak [peak S]) (r = -0.34, P = .01) (Figure 3, A), during early diastole (peak E) (r = 0.36, P = .006) (Figure 3, B), and during late diastole (peak A) (r = 0.31, P = .02) (Figure 3, C); and it also correlated with 25% untwist (r = 0.36, P = .006) (Figure 3, D). The 25% untwist and SrL peak E were independent predictors of exercise capacity (P = .003 and P = . 01, respectively).
In this study, we assessed myocardial strain using an ultrasound speckle tracking technique that avoids the angle dependency of Doppler-based techniques. Longitudinal, radial, and circumferential strain and strain rates, and twist and untwist rates were measured in nHCM patients and in age-and gender-matched controls. Although nHCM patients typically exhibit hyperdynamic systolic function as assessed by conventional echocardiography, our STE findings of reduced SrL and delayed 25% untwist as significant predictors of exercise capacity are to the best of our knowledge novel.
The STE tracks characteristic speckle patterns created by interference of ultrasound beams in the myocardium that is based on gray scale B-mode images, and is angle independent. In humans, measurement of LV long-axis strain by STE correlated well with magnetic resonance imaging tagging (r = 0.87). We are now able to characterize myocardial function in greater detail, breaking down global function into the individual components that make up the whole. This can now be done for both the diastolic phase and systolic phase of myocardial function quickly and reproducibly using a relatively new technique of speckle tracking. Our findings of reduced longitudinal, circumferential, and radial strains are in agreement with previously published works. However, the clinical relevance of these alterations in function has not been previously assessed.
The function of the LV is often described as hyperdynamic in HCM. This is based on the frequent observation of high normal or supranormal EF. In this study, we showed that HCM patients had reduced systolic STE-derived longitudinal, circumferential, and radial strain and strain rates despite normal LVEF. These results are consistent with TDI-derived data that have shown that systolic long-axis function is reduced in patients with HCM. Carasso et al showed reduced longitudinal but higher circumferential STE-derived strain and strain rate in HCM patients compared with controls. Our findings of reduced circumferential strain are consistent with previous tagged magnetic resonance imaging studies.
The majority of HCM patients have some degree of diastolic dysfunction independent of symptoms, the presence of ventricular outflow tract obstruction, the magnitude of hypertrophy, or a combination of these. The mechanisms responsible for the diastolic dysfunction are probably multifactorial. The hypertrophy, myocyte disarray, and fibrosis may increase passive LV stiffness. Several studies had shown abnormalities of active relaxation in HCM patients; and these may relate to abnormal calcium handling, to microvascular myocardial ischemia and mechanical dyssynchrony, or to impaired myocardial energy utilization. There are several echocardiographic measurements that can quantify diastolic dysfunction in HCM patients, but there remain some controversies about their accuracy. Our data showed no significant difference in transmitral Doppler-derived parameters between HCM patients and controls. However, there were significant abnormalities of other echocardiographic measurements of diastolic function (such as LA volume, LAVI, and TDI-derived indexes) in HCM patients. More importantly, untwist (a marker of diastolic function) was significantly delayed in these patients, which is in agreement with a previous study.
Traditionally, it has been believed that LV diastolic dysfunction is the predominant cause of breathlessness and exercise limitation in HCM patients. Interestingly, we found that markers of both LV systolic and diastolic function significantly correlated with exercise capacity.
We have previously demonstrated that exercise capacity in patients with HCM is limited by the failure to increase stroke volume on peak exercise. Furthermore, we showed an abnormal lengthening of time to peak LV filling on exercise in these patients. The mechanism of this limitation of LV filling has not been fully elucidated. Left ventricular untwist is increasingly being recognized as an important component of diastolic function, contributing significantly to suction. Knudtson et al have shown that myocardial ischemia impairs untwist. Untwisting, a recoil phenomenon, occurs during the isovolumetric relaxation period, tending toward geometric restoration of the zero twist point (end diastole) with concurrent maximum LV volume. Twist and untwist rates appear to be closely related. During systole, myocardial deformation (twist) results in elastic energy being stored in compressed titin and transmural shear between myofibril sheets. This energy is spent during early diastole, before the onset of LV filling, which leads to untwisting of the LV that contributes to the generation of suction. The basis of twist and untwist abnormalities in HCM patients is possibly due to regional abnormalities of systolic and diastolic function. Peak ventricular untwist rate has been correlated with the time constant of LV pressure decay and intraventricular pressure gradients, both markers of diastolic function. Reduced untwist rate may lead to impairment in diastolic filling, which becomes more notable with increased heart rates due to short filling period. Notomi et al showed that greater untwisting rate produced an increase in suction during exercise in healthy subjects, which is able to fill the ventricle more despite a shorter filling period. This suggests that untwist rate is an important mechanism augmenting LV filling. The delayed untwist in HCM patients observed in this study may lead to reduced LV suction and filling, which will reduce the ability to use the Starling mechanism, leading to a failure to increase cardiac output, and as a result will limit exercise capacity.
The relevance of changes in twist and untwist during exercise would be interesting to examine; however, we here demonstrate that even resting measurements of strain and untwist were significant predictors of exercise capacity. We have been unable to reliably measure these parameters during exercise because of several technical difficulties. First of all, it was hard to maintain the same heart rate during exercise while acquiring apical and basal short-axis echocardiographic images that are crucial for the measurement of twist and untwist. Furthermore, exercise exaggerates the issue of through-plane motion, particularly at the basal short-axis level that is again vital for twist measurement. In addition, the current STE analysis software has insufficient temporal resolution for higher heart rate during exercise. It optimally operates at frame rate of 40 to 100 frames per second that is adequate only for resting studies. Moreover, excessive respiratory movement may degrade the quality of images as compared with the resting images. Further advances in the STE technique (eg, 3-dimensional speckle tracking with higher temporal resolution) are required to resolve these issues.
In this study, we report that there are no significant differences in twist between nHCM patients and controls. However, we showed a significant reduction in apical rotation rates during both systole and diastole. Previous study showed that apical rotation represents the major contributions to LV twist. The absence of significant differences in LV twist observed in this study could be explained by the increased through-plane motion at the base of the heart.
The heart rate is slower in the nHCM group (Table I) probably because of medications (ie, βblockers and/or calcium-channel blockers). Such therapy might have affected the findings in these patients. All pertinent physiologic parameters were corrected for heart rate. Untwist timing parameters were corrected for heart rate by converting R-R interval to 100% as previously described. All longitudinal, radial, and circumferential strain parameters were independent of heart rate because we measured peak values rather than timing.
The presence of asymmetric hypertrophy in nHCM patients renders geometric assumptions inaccurate in LV mass measurement. Hence, we measured MWT, which had been widely used and validated as a marker of LV hypertrophy in these patients. Furthermore, we showed that there is no relationship between exercise capacity and extent of LV hypertrophy as measured by both LV mass index and MWT (Figure 2).
Although nHCM patients often had normal/supranormal LVEF as measured by conventional echocardiography, STE imaging demonstrated a marked impairment of both systolic and diastolic LV function; and they were both significant determinants of exercise capacity. Abozguia et al. Page 11 Published as: Am Heart J. 2010 May ; 159(5): 825-832.
Sponsored Document Sponsored Document Abozguia et al. Page 12 Published as: Am Heart J. 2010 May ; 159(5): 825-832. Sponsored Document Sponsored Document Sponsored Document Sponsored Document Sponsored Document Sponsored Document Abozguia et al. Page 13
Table I The clinical characteristics and cardiopulmonary exercise test results of nHCM patients and controls Parameter nHCM Controls P value Age (y) 52 ± 11 51 ± 16 .74 n (female) 56 (16) 43 (12) .64 Heart rate (beat/min) 69 ± 13 80 ± 16 .0002 ⁎ Systolic BP (mm Hg) 126 ± 19 125 ± 20 .78 Diastolic BP (mm Hg) 78 ± 11 78 ± 11 .80 Peak VO 2 (mL/[kg min]) 23.28 ± 6.31 37.70 ± 7.99 <.0001 ⁎ Drug therapy, n (%) Diuretic 10 (18) 0 -ACE inhibitor 6 (11) 0 -ARB 1 (2) 0 β-Blocker 17 (31) 0 -Calcium-channel blocker 26 (47) 0 -Aspirin 13 (24) 0 -Warfarin 5 (9) 0 -Nitrate 2 (4) 0 -Statin 15 (27) 0 -BP, Blood pressure; ACE, angiotensin-converting enzyme; ARB, angiotensin receptor blockers. ⁎ Indicates statistical significance. Published as: Am Heart J. 2010 May ; 159(5): 825-832. Sponsored Document Sponsored Document Sponsored Document Abozguia et al. Page 14 Table II Standard, Doppler, and PW-TDI echocardiographic characteristics of nHCM versus controls Parameter nHCM Controls P value EF biplane (%) 62.76 ± 9.05 62.48 ± 5.82 .86 Biplane LA volume (mL) 72.69 ± 29.22 35.09 ± 9.83 <.0001 ⁎ LAVI (mL/m 2 ) 33.92 ± 14.62 11.76 ± 10.17 <.0001 ⁎ LV mass index (g/m 2 ) 181.92 ± 84.09 71.67 ± 23.82 <.0001 ⁎ MV E (m/s) 0.69 ± 0.16 0.67 ± 0.15 .49 MV A (m/s) 0.64 ± 0.19 0.59 ± 0.15 .17 MV E/A 1.18 ± 0.48 1.20 ± 0.38 .83 MV dec time (ms) 231.12 ± 71.66 260.10 ± 68.24 .06 Antlat Sm (cm/s) 0.06 ± 0.02 0.10 ± 0.03 <.0001 ⁎ Antlat Em (cm/s) 0.08 ± 0.03 0.11 ± 0.04 <.0001 ⁎ Antlat Am (cm/s) 0.07 ± 0.03 0.09 ± 0.02 <.0001 ⁎ Infsep Sm (cm/s) 0.06 ± 0.02 0.08 ± 0.02 .0001 ⁎ Infsep Em (cm/s) 0.05 ± 0.02 0.08 ± 0.03 <.0001 ⁎ Infsep Am (cm/s) 0.07 ± 0.02 0.09 ± 0.02 <.0001 ⁎ Average Sm (cm/s) 0.06 ± 0.02 0.09 ± 0.02 <.0001 ⁎ Average Em (cm/s) 0.06 ± 0.02 0.09 ± 0.04 <.0001 ⁎ E/Em antlat 9.70 ± 3.86 6.71 ± 2.35 .0001 ⁎ E/Em infsep 17.05 ± 7.38 9.64 ± 3.34 <.0001 ⁎ E/Em average 11.81 ± 4.28 8.04 ± 2.92 <.0001 ⁎ MV E, Transmitral early diastole velocity; MV A, transmitral late diastole velocity; dec, deceleration.
⁎ Indicates statistical significance.
Published as: Am Heart J. 2010 May ; 159(5): 825-832.
Published as: Am Heart J. 2010 May ; 159(5): 825-832.Sponsored DocumentSponsored DocumentSponsored Document
Published as: Am Heart J. 2010 May ; 159(5): 825-832.
We are grateful for the help and advice on statistics given by
Table III STE results of longitudinal, radial, and circumferential strain and strain rates in nHCM patients versus controls
Parameter nHCM Controls P value Longitudinal SL peak S (%) -12.74 ± 4.21 -18.38 ± 2.99 <.0001 ⁎ SrL peak S (1/s) -0.92 ± 0.25 -1.14 ± 0.2 <.0001 ⁎ SrL peak E (1/s) 0.94 ± 0.35 1.33 ± 0.3 <.0001 ⁎ SrL peak A (1/s) 0.78 ± 0.28 1.07 ± 0.29 <.0001 ⁎ Radial SR peak S (%) 17.52 ± 7.86 29.63 ± 13.26 <.0001 ⁎ SrR peak S (1/s) 1.22 ± 0.34 1.39 ± 0.34 .02 ⁎ SrR peak E (1/s) -1.0 ± 0.39 -1.46 ± 0.52 <.0001 ⁎ SrR peak A (1/s) -0.77 ± 0.38 -1.09 ± 0.61 <.01 ⁎ Circumferential SC peak S (%) -16.50 ± 4.29 -19.99 ± 5.53 <.001 ⁎ SrC peak S (1/s) -1.4 ± 0.35 -1.54 ± 0.34 .05 SrC peak E (1/s) 1.34 ± 0.42 1.8 ± 0.65 <.001 ⁎ SrC peak A (1/s) 0.87 ± 0.34 1.09 ± 0.47 .02 ⁎ SL, Longitudinal strain; SR, radial strain, SrR, radial strain rate; SC, circumferential strain.
⁎ Indicates statistical significance.
Published as: Am Heart J. 2010 May ; 159(5): 825-832. Table IV STE measurement of LV rotation, twist, and untwist in nHCM and controls Parameter nHCM Controls P value Apical rot and rot-r Peak rot (°) 7.65 ± 6.90 9.48 ± 5.40 .18 Peak S rot-r (°/s) 49.77 ± 49.49 74.56 ± 42.25 .01 ⁎ Peak E rot-r (°/s) -41.19 ± 39.14 -72.96 ± 38.00 <.001 ⁎ Peak A rot-r (°/s) -26.24 ± 33.94 -39.87 ± 27.01 <.05 ⁎ Basal rot and rot-r Peak rot (°) -6.25 ± 3.73 -5.80 ± 2.81 .55 Peak S rot-r (°/s) -57.89 ± 28.94 -61.24 ± 17.82 .54 Peak E rot-r (°/s) 53.07 ± 25.96 48.88 ± 26.42 .47 Peak S rot-r (°/s) 42.39 ± 24.52 44.67 ± 20.68 .65 Twist and untwist Peak twist (°) 13.37 ± 7.14 13.69 ± 6.55 .83 25% Untwist (°) 9.85 ± 5.47 10.27 ± 4.91 .71 Time 25% untwist (R-R%) 11.34 ± 5.63 7.21 ± 8.01 .01 ⁎
⁎ Indicates statistical significance.
Published as: Am Heart J. 2010 May ; 159(5): 825-832.
We can learn about people's conceptions of the ideal life by looking at what they imagine heaven to be like. Although voluptuous and gendered (and even sexist) accounts of the afterlife are familiar, more reflective views grow ever more distant from our actual human form of life-many Christians believe that in heaven there will be no marriage, sexual intercourse, or procreation. For those of us who think of the human frame not as the creation of a divine designer but as a contingent product of blind natural selection, it is simply a truism that our biology falls far short of perfection. If we were to engineer ex nihilo a new form of intelligent life that would be maximally flourishing, it would bear little resemblance to actual human beings. Nor is it likely to be divided into male and female, or to engage in sexual intercourse for reproductive purposes-sexual dimorphism was after all not selected because it reflects some deep intrinsic value, but for familiar evolutionary reasons. Awareness that we are mere products of blind chance, that there is no special necessity that intelligent beings would be divided into male and female, or walk on two legs, or enjoy music or dance, might be disturbing to some. But we mustn't confuse pressure in the gut with a reductio.
Maximally flourishing rational beings are unlikely to be sexually dimorphic. This unremarkable point, however, doesn't by itself say anything about what kind of children we humans should have. We have defended procreative beneficence: the claim that parents have reasons to create children with the best chance of the best life (Savulescu 2001;Savulescu and Kahane 2009). There are three points to elaborate in relation to this claim.
First, evaluating embryos should be individualistic based on maximum available information about that particular embryo. As we have pointed out, it will be possible to test for thousands of genes, or other biological states. Sex will only be one feature of a predicted embryo. But male and female embryos will each have different psychological, cognitive, and physical capacities. An individual female embryo might have a disposition to borderline personality disorder, be of lower intelligence, or be infertile. A male embryo may have a higher predicted range of intelligence, greater resistance to disease, and exceptional musical abilities. Faced with these facts, it would be rational to choose the male embryo. Sparrow (2010) claims that female sex is better because of longer life and capacity to experience pregnancy and childbirth. But pregnancy and childbirth are associated with mortality and permanent damage-it is not clear these are an overall benefit (Sparrow seems to mistakenly identify the best expected life with the life with the most open future, two distinct if partly overlapping goals). And men can have vastly more biologically related children (Genghis Khan had apparently thousands) and can have sex in more ways. Even if there were some overall advantages to female sex, such advantages are contingent and not necessary. For example, men could be modified to carry pregnancies or live longer. And women could avoid the pain and risks of morbidity and mortality of pregnancy and childbirth by ectogenesis or xenogestation, the genetic modification of nonhuman animals, such as pigs, to carry human fetuses to term. Australia has just recognized the first person to have neither female nor male sex legally. It is only a matter of time before enhancement technologies allow people to be both male and female, or neither. So sex will become a matter of personal choice (Savulescu 2010).
Even if there were advantages to being female, these would have to be weighed against individual predictive indices relating to a particular embryo. Other differences in genetic value are likely to swamp the contribution of sex in any particular evaluation. We would have as much reason to select female embryos as we would to select African American, Asian, or Caucasian embryos, depending on which racial origin embryos were likely overall to statistically have the best life. Sparrow mistakenly imports a group characteristic into an individual multifactorial evaluation.
Second, the reasons to select embryos are radically context dependent. They apply to the expected lives of particular individuals in highly specific social conditions (Kahane and Savulescu 2009)-conditions that include the impact of tradition and prejudice, and the choices of other parents. Sparrow needs to argue that in these conditions, girls can be expected to have clearly better lives. This is partly an empirical question. We find Sparrow's arguments for this claim unpersuasive, but we will leave it to others to debate the details.
Third, even if girls could be expected to have better lives than boys, and parents thus had a reason to prefer girls, these reasons would still not dictate a procreative choicethey would first need to overcome competing reasons, a point we have also emphasized (Savulescu and Kahane 2009), and which Sparrow overlooks in the rush to a reductio. In Savulescu (2001), it was argued that we could even have reasons to have a disabled child, such as a child with Down syndrome, to make a political statement or further a political goal. Procreative beneficence is not the only principle or supplier of reasons, as we have been at pains to repeat. These other reasons might include the good of the parents, and the social good. It might be that a couple could have most reason to create a boy with a prospect of a good life, even if they could create a girl with a better life, for reasons of sex balancing, or just because this boy will make the lives of other women even better. When these other reasons are taken into account, then it seems to us that, even if Sparrow's basic argument was correct, and assuming a social environment with minimal prejudice against women, parents might at best have weak reasons to prefer a girl to a boy if this is their first child, and if sexual selection is not widely practiced.
We therefore doubt that Sparrow's argument succeeds. There is, however, a truth in the discomfort that he expresses about the prospect of a future without sexual dimorphism. As our powers of biological intervention increase, we will be repeatedly faced with the choice of whether to hold on to our imperfect human biology and the forms of life that are shaped around it, or to overcome these in favor of an alternative with greater potential for well-being. Sparrow ends by suggesting that in response to this discomfort, we should conclude that the biologically normal has deep moral value. This conclusion is puzzling, given that Sparrow had earlier cited some of the many critics of this view, and has carefully outlined the dangers of status quo bias. To think of our actual biology as the product of blind evolution is precisely to give up the idea that the normal features of human biology have any inherent value (Kitcher 1999;Dorsey 2010;Kahane and Savulescu 2009). The arguments for this conclusion seem to us unanswerable, and Sparrow says nothing to address them.
Might there be a better way to ground this discomfort? In many contexts we have reasons to be partial: to give greater weight to ourselves and our family, or to care about some country or institution, even when there are better alternatives, and when doing so doesn't maximally promote possible good from the "point of view of the universe." These reasons are not generated by the intrinsic betterness of the people and things we care about, but by our relations to them, and our contingent history. Might we have such reasons to be partial to the human way of life, including sexual bimorphism and the role it plays in human life? It seems perfectly plausible to hold we have such "conservative" reasons (Cohen unpublished). If a shared history can generate such reasons, then our reasons to preserve features of human life that have existed from time immemorial should be very strong. And contrary to Williams (2006), such partiality is not a 'human prejudice' based in our subjective attitudes, but can be objectively justified.
Needless to say, such conservatism has its limits. We don't need to tear down an antique house or abandon venerable traditions simply because we could replace them with something better-but it doesn't follow we can always justifiably hold on to the way things are and give up the opportunity to promote very great goods. This is why we hold that, all things considered, we have strong reasons to enhance cognition and mood beyond the biologically normal. But we might have reason to preserve the distinction between men and women precisely because-unlike biological limits on human cognition, emotion, motivation, and longevityit does not present some deep obstacle to the promotion of human flourishing. Preserving sexual dimorphism is thus perfectly compatible with the project of human enhancement.
Thomas Marino, Temple University School of Medicine Robert Sparrow's article (2010) attempts to defend the position that humans are sexually dimorphic. However, human embryology clearly shows this is not the case. Human embryos start out as indifferent and only the correct set of genes and hormones appearing in the correct order and at the appropriate time can lead to male and female phenotypes. So variations in sexual morphology are not uncommon. Historically this has been acknowledged and debated. In the early medical literature, the notion of sexual intermediaries was discussed. The medicalization of sex and gender, however, added complexity to this discussion. In cultures where sexual intermediaries are acknowledged and accepted, individuals are given their own position in the society. In these instances, sexual dimorphism clearly is not accepted. It seems difficult to acknowledge all this information and still defend sexual dimorphism. In 2000, Blackless and colleagues examined the literature and suggested that 0.2-2.0% of the population have variations in sexual development that affect either chromosomes, gonads, or the phenotype of the individual. Other studies have shown that 1 in 5,000 infants are born with ambiguous genitalia (Thyen et al. 2006;Hamerton et al. 1975). In addition, 1 in 400 male and 1 in 700 female newborns had sex chromosome abnormalities (Hamerton et al. 1975). A major question is whether this many individuals with variations in sexual morphology should be viewed as nonexistent. From several perspectives it appears they should not be.
As early as the 1800s, it was known that developing human embryos go through a period of development called the indifferent stage (Marino 2010). From conception until about 6 weeks of gestation, the embryo has indifferent gonads capable of becoming ovaries or testes. Depending on the genes expressed and other developmental events, the gonads can develop or remain in a primitive stage. In Address correspondence to Thomas Marino, Temple University School of Medicine, 3400 N. Broad St., Philadelphia, PA 19140, USA. E-mail: marino@temple.edu some individuals they can remain as ovotestes or streak gonads.
The genital duct system also goes through the indifferent stage, and early in development, ducts are present that can give rise to the male or female genital duct system. Depending on the factors expressed, the mesonephric ducts can become the male genital duct system and the paramesonephric ducts degenerate. Conversely, the female genital ducts can develop from the paramesonephric ducts and the mesonephric ducts can degenerate because of the lack of testosterone. However, if testosterone receptors are not present but Mullerian inhibiting hormone is present, the individual can be born without either duct system.
External genital ambiguity persists until the end of the first trimester. With dihydrotestosterone or excess androgens, the external genitalia will masculinize. With estrogens present, the external genitalia will have the female morphology. In XX individuals, too much testosterone or androgens will cause masculinization of the external genitalia, and in XY individuals the lack of dihydrotestosterone and testosterone will cause feminization.
Therefore, depending on the factors present, the embryo can take any number of developmental pathways. The majority of embryos become either male or female, but the pathways leading to one or the other are not fully understood. Recent reviews have shown there is a remarkable complexity to understanding the pathway from a fertilized egg to sex differentiation (Wilhelm et al. 2007). In addition, very little is known about the development of the ovary. As Wilhelm and colleagues state (2007), "The early development of the vertebrate ovary is poorly understood at the histological, cellular, and molecular levels, a rather amazing situation given the importance of this organ" (20). Nonetheless, one could argue there are at least three genders: male, female, and indifferent. And other
July, Volume 10, Number 7, 2010
Acknowledgment: We are grateful to the
Intracoronary testosterone infusions induce coronary vasodilatation and increase coronary blood flow. Longer term testosterone supplementation favorably affected signs of myocardial ischemia in men with low plasma testosterone and coronary heart disease. However, the effects on myocardial perfusion are unknown. Effects of longer term testosterone treatment on myocardial perfusion and vascular function were investigated in men with CHD and low plasma testosterone. Twenty-two men (mean age 57 ± 9 [SD] years) were randomly assigned to oral testosterone undecanoate (TU; 80 mg twice daily) or placebo in a crossover study design. After each 8-week period, subjects underwent at rest and adenosine-stress first-pass myocardial perfusion cardiovascular magnetic resonance, pulse-wave analysis, and endothelial function measurements using radial artery tonometry, blood sampling, anthropomorphic measurements, and quality-of-life assessment. Although no difference was found in global myocardial perfusion after TU compared with placebo, myocardium supplied by unobstructed coronary arteries showed increased perfusion (1.83 ± 0.9 vs 1.52 ± 0.65; p = 0.037). TU decreased basal radial and aortic augmentation indexes (p = 0.03 and p = 0.02, respectively), indicating decreased arterial stiffness, but there was no effect on endothelial function. TU significantly decreased high-density lipoprotein cholesterol and increased hip circumference, but had no effect on hemostatic factors, quality of life, and angina symptoms. In conclusion, oral TU had selective and modest enhancing effects on perfusion in myocardium supplied by unobstructed coronary arteries, in line with previous intracoronary findings. The TU-related decrease in basal arterial stiffness may partly explain previously shown effects of exogenous testosterone on signs of exercise-induced myocardial ischemia.
Cross-sectional studies suggest an inverse association between coronary heart disease (CHD) incidence and/or severity and endogenous testosterone levels irrespective of age, and
testosterone treatment increases coronary blood flow and improved signs on myocardial ischemia in men with CHD. The aim of the present study was to investigate the hypothesis that longer term oral testosterone undecanoate (TU) treatment beneficially affects myocardial perfusion in men with established CHD and low testosterone. Second, we investigated effects on vascular and endothelial function, selected metabolic risk factors for CHD, quality of life, and angina symptoms.
Men aged 40 to 75 years with angiographically proven CHD (≥70% lesion in ≥1 major coronary artery or major branch) were recruited from the outpatient departments, cardiac catheterization lists, and hospital databases of the Royal Brompton and Harefield Hospitals, London, United Kingdom. All patients had plasma testosterone ≤12 nmol/L and normal prostate-specific antigen (normal range 0 to 4 μg/L). Exclusion criteria included myocardial infarction, coronary intervention, or thoracic or abdominal surgery in the previous 3 months; hypertension; history of testosterone treatment or similar hormone therapy; hemoglobin >16 g/dl; hematocrit >50%; history of hormone-dependent cancer; permanent pacemaker; intolerance of confined spaces; or participation in another research study within the previous 60 days. The study complied with the Declaration of Helsinki, the protocol was approved by the local research ethics committee, and all patients gave written informed consent.
The study used a randomized, placebo-controlled, crossover design. After consenting to the study, a screening blood sample was obtained to measure serum testosterone, prostate-specific antigen, and full blood count. Patients fulfilling the inclusion criteria then returned for a randomization visit, at which anthropomorphic measurements were obtained, quality-of-life questionnaires were completed, and study drug was dispensed (oral TU, 80 mg twice daily [Andriol Testocaps, Organon, Oss, The Netherlands] or identical placebo). After 8 weeks, an evaluation visit was conducted that included measurement of myocardial perfusion, endothelial function, quality-of-life assessment, and blood sampling for measurement of hormone, lipid, and hemostatic profiles. Patients then crossed over to the opposite treatment and returned for identical repeated testing after an additional 8 weeks at the same time of day as the first evaluation visit. Patients completed a symptom diary throughout the study.
Patients were permitted to use current antianginal therapy throughout the randomized treatment period, but stopped vasoactive therapy (including statins) for 1 week before each evaluation visit. Sublingual glyceryl trinitrate was permitted for treatment of anginal attacks during this time. However, if used on the day of the evaluation visit, it was rescheduled. Patients fasted for ≥6 hours, and caffeine-containing beverages were prohibited for ≥12 hours before each evaluation visit.
All cardiovascular magnetic resonance (CMR) studies were performed on a 1.5-Tesla scanner (Siemens Sonata, Siemens, Germany) with a 4-channel body-array coil. For each study, firstpass myocardial perfusion at rest and with adenosine stress was assessed. A high-resolution saturation-recovery fast low-angle shot sequence was used (field of view read 320 to 400 mm, field of view phase 75% to 100% of the field of view read, acquired voxel size 1.3 to 1.6 × 1.3 to 1.6, slice thickness 10 mm, flip angle 10°, acquisition duration 295 ms, echo time 1.11 ms, and time from saturation pulse to beginning of imaging 63 ms). Gadoliniumdiethylenetriaminepentaacetic acid was injected as a bolus at 7 ml/s using a power injector (Medrad, Cambridgeshire, United Kingdom) through an 18 G intravenous cannula. Images were acquired over 50 cardiac cycles, and patients were asked to breath-hold in end-expiration for as long as was comfortable. One midventricular short-axis slice, representative of the whole ventricle, was acquired during each cardiac cycle in diastole. The stress study was performed ≥20 minutes after the study at rest. Adenosine was infused for 4 minutes at 140 μg/kg/min before stress images were acquired. On the second visit, the short-axis slice planned to be imaged was compared with the short-axis slice from the first visit to ensure an equivalent slice was studied.
Because studies were analyzed quantitatively, a dual-bolus protocol was used. For both studies at rest and with stress, patients received a low-dose (0.01 mmol/kg at 0.05 mol/L) bolus of gadolinium (Omniscan; Nycomed, Lidingö, Sweden) approximately 2 minutes before a highdose bolus of gadolinium (0.1 mmol/kg at 0.5 mol/L). This allowed an accurate arterial input function and optimized signal in the myocardium. Blood pressure (BP) and heart rate were monitored throughout each study.
On both visits, left ventricular (LV) volume, function, and mass measurements were acquired using breath-hold segmented true fast imaging with steady-state precession cines according to a standard protocol.
Each study was analyzed by a blinded observer using customized software (CMRtools; Cardiovascular Imaging Solutions, London, United Kingdom). For perfusion analysis, subendocardial and subepicardial myocardial borders were drawn and adjusted if necessary to compensate for cardiac and respiratory motion. Arterial input function and myocardial tissue response curves were corrected using baseline subtraction, then assumed to be linearly proportional to gadolinium concentration. An index of myocardial perfusion (a unitless measure) was calculated at rest and at stress from the respective first-pass myocardial signal intensity-time curves using Fermi deconvolution. Myocardial perfusion index (the ratio between stress and rest myocardial perfusion) was calculated using model-based deconvolution using the Fermi function. Each CMR image was divided into 4 (anterior, lateral, inferior, and septum), and areas of myocardium supplied by coronary arteries with and without significant (≥70% lesion) angiographically documented coronary stenosis were analyzed separately. LV mass, end-diastolic and end-systolic volumes, and ejection fraction were determined using standard manual contouring techniques.
Peripheral pressure waveforms were recorded from the right radial artery using applanation tonometry (SphgmoCor PX, version 6.31; AtCor, Sydney, Australia). BP was measured noninvasively every 5 minutes throughout the tonometry study on the contralateral arm (Dinamap, GE Healthcare, Chalfont St. Giles, Buckinghamshire, United Kingdom). Immediately after each BP measurement, radial artery pulse recordings were acquired with an averaged peripheral waveform generated from 20 sequential waveforms. Corresponding central waveforms were generated using an online validated transfer function to determine central augmentation index, central pressure, and heart rate. Augmentation index was calculated as the ratio of pulse pressure at the second systolic arterial pressure waveform peak to that of the first systolic peak and is a measure of arterial stiffness. Patients underwent measurement of endothelium-dependent and -independent responses using radial artery applanation tonometry before and after salbutamol (400 μg) and sublingual glyceryl trinitrate (250 μg) as previously described.
Different aspects of quality of life were measured using validated questionnaires: the 36-Item Short Form Health Survey questionnaire, which measures general health (a high score indicates better functioning or more pain); Cardiac Health Profile, which measures cardiac well-being (a low score indicates a positive response); and the Aging Male Symptom Scale.
Testosterone, 17β-estradiol, sex hormone-binding globulin, follicular-stimulating hormone, luteinizing hormone, and insulin were measured using radioimmunoassay (Abbott IMX System; Abbott Diagnostics, Berkshire, United Kingdom), and glucose was measured using a standard Beckman Delta analyzer (Beckman Coulter UK Ltd, United Kingdom). Total cholesterol, high-density lipoprotein cholesterol, and triglycerides were measured using a Beckman CX7 analyzer. Low-density lipoprotein cholesterol was estimated using the formula of Friedewald. Full blood count was measured using a standard technique with a Technicon 8*1 analyzer (Bayer Technicon, United Kingdom). Fibrinogen and factor VII were measured using an automated coagulometer (Instrumentation Laboratory, United Kingdom), and plasminogen activator inhibitor-1, using chromogenic substrate assays. Sample-size calculation was determined by the primary end point of myocardial perfusion. Our intracoronary testosterone study showed an increase in blood flow to testosterone of approximately 10% to 15%. A mean difference of 15% in myocardial perfusion index with an SD of 0.5 requires 22 patients, with assumptions of a 5% significance level, 80% power to detect a difference between groups, and randomization to the 2 groups in equal proportions. Allowing for a 20% drop-out rate, 28 patients were required for enrollment.
First, the size of the period-by-treatment interaction (carryover effect) was examined, and no carryover effect was found for any variable. Second, the difference between treatment groups was analyzed by subtracting results of the second period from results from the first period, and a t test was performed to compare these differences. An advantage of this method compared with a paired t test is that it removes the effect of period from the analysis. When data were not normally distributed, analysis was performed using Mann-Whitney test. Significance was set at p ≤0.05. Data are presented as mean ± SD or mean difference (95% confidence interval).
For myocardial perfusion data, each CMR image was divided into 4 (as described). Because this violates the assumption of independence, data were analyzed in 2 ways. First, a t test was used (which assumes independence), then linear regression with robust SEs to allow for repeated observations from the same subjects.
Approximately 4,000 patients were contacted by post, and 50% of patients expressed interest in participating in the study. Of these, approximately 80% were excluded because of hypertension. One hundred nine patients attended for a consent and screening appointment. Half those screened had testosterone greater than the inclusion criteria.
Twenty-eight patients were enrolled in the study, and 23 completed the protocol. Two patients were randomly assigned, but withdrew consent before taking the study drug (1 moved away and 1 repeatedly did not attend the randomization visit), and 1 patient was withdrawn after a few days on study medication (TU) subsequent to hospitalization for a suspected stroke. One patient withdrew consent after the first treatment phase (placebo) because of depression, and 1 patient withdrew for personal reasons. Characteristics of the 25 patients who had at least 1 evaluation visit are listed in Table 1. Although not actively questioned about symptoms of hypogonadism, 1 patient was known to have previously used testosterone treatment (stopped 3 months before randomization). In addition, prerandomization Aging Male Symptom Scale scores indicated that 22 patients (88%) had moderate or severe impairment of sexual function. This score was not changed by either testosterone or placebo.
Hormone levels are listed in Table 2. TU did not significantly change serum total testosterone or 17β-estradiol; however, percentages of free and bioavailable testosterone and dihydrotestosterone levels increased (all p <0.001). Conversely, TU treatment was associated with a significant decrease in sex hormone-binding globulin, follicular-stimulating hormone, and luteinizing hormone (all p <0.001).
Twenty-two patients had assessable CMR data for both evaluation visits. Global myocardial perfusion was not significantly different between treatments transmurally (1.8 ± 0.7 vs 1.6 ± 0.7, TU vs placebo; p = 0.28) or in the subendocardium (1.8 ± 0.7 vs 1.5 ± 0.6, TU vs placebo; p = 0.17) or subepicardium (1.8 ± 0.7 vs 1.6 ± 0.7, TU vs placebo; p = 0.33). When analyzed by segments, no significant difference in perfusion was shown in the subendocardium (1.83 ± 0.92 vs 1.54 ± 0.88, TU vs placebo, p = 0.2) or subepicardium (1.86 ± 0.83 vs 1.65 ± 0.89, TU vs placebo, p = 0.54) after adjustment for repeated measures. However, myocardial segments supplied by coronary arteries without significant obstruction showed significant improvement in the myocardial perfusion index of the subendocardium (1.83 ± 0.9 vs 1.52 ± 0.65, TU vs placebo, p = 0.037; Figure 1) and borderline improvement in the subepicardium (1.85 ± 0.82 vs 1.59 ± 0.69, TU vs placebo, p = 0.058; Figure 1) in patients administered TU versus placebo. Myocardial segments supplied by coronary arteries with significant coronary atherosclerosis showed no difference in myocardial perfusion indexes in the subendocardium (1.83 ± 0.94 vs 1.55 ± 1.03, TU vs placebo respectively; p = 0.11) or subepicardium (1.86 ± 0.84 vs 1.69 ± 1.02, TU vs placebo; p = 0.29). Adenosine did not decrease heart rate or cause atrioventricular block in any patient.
Sixteen patients had LV data for both visits. LV ejection fraction was significantly higher after testosterone treatment compared with placebo (67 ± 8% vs 65 ± 8%; p = 0.05). However, stroke volume (89 ± 12 vs 88 ± 14 ml; p = 0.61), end-systolic volume (45 ± 14 vs 47 ± 14 ml; p = 0.1), end-diastolic volume (135 ± 19 vs 134 ± 18 ml; p = 0.97), and mass (170 ± 33 vs 169 ± 34 g; p = 0.73) were not different.
Seventeen patients completed both endothelial function assessments. TU treatment was associated with significantly lower baseline radial augmentation (p = 0.03; Figure 2) and aortic augmentation indexes (141 ± 12 vs 147 ± 9, TU vs placebo, respectively; p = 0.02) compared with placebo, indicating a testosterone-induced decrease in basal arterial stiffness. However, testosterone did not affect salbutamol-induced changes in radial augmentation indexes compared with placebo (p = 0.48; Figure 2), signifying no effect on endothelium-dependent responses. Glyceryl trinitrate-induced decreases in radial augmentation indexes were similar after both treatments (p = 0.15; Figure 2). Time to return of reflected waves, an estimate of pulse-wave velocity, was prolonged after TU treatment compared with placebo (142.4 ± 16.6 vs 137.2 ± 9.3 ms; p = 0.04). Brachial systolic BP was not altered throughout the tonometry study (148 ± 16 vs 147 ± 17 mm Hg, TU vs placebo; p = 0.71). Similarly, there was no significant difference in central pulse pressure (52 ± 10 vs 52 ± 13 mm Hg, TU vs placebo, respectively; p = 0.88) or heart rate at rest (62 ± 7 vs 60 ± 7 beats/min, TU vs placebo, respectively; p = 0.08).
TU treatment was associated with significantly lower high-density lipoprotein cholesterol and apolipoprotein-A1 and higher high-density lipoprotein-low-density lipoprotein ratio compared with placebo (Table 3). Apolipoprotein-A1: apolipoprotein-B ratio decreased with borderline statistical significance, and glucose and insulin were not different between treatments (Table 3).
There was no difference in plasminogen activator inhibitor-1, fibrinogen, or factor VII after treatment with TU or placebo (Table 3). Hematocrit slightly but significantly increased after TU treatment (44.5 ± 1.8% vs 43.5 ± 1.9%, TU vs placebo, respectively; p = 0.002). However, hemoglobin was not affected (15 ± 0.6 vs 14.9 ± 0.7 g/dl, TU vs placebo, respectively; p = 0.22).
There was a statistically significant increase in hip circumference after TU compared with placebo (108 ± 6 vs 107 ± 6 cm, respectively; p = 0.03), but no difference in body weight (92 ± 11 vs 92 ± 13 kg; p = 0.1), waist circumference (107 ± 8 vs 108 ± 10 cm; p = 0.77), waisthip ratio (0.98 ± 0.05 vs 1 ± 0.06; p = 0.14), body surface area (2.1 ± 0.2 vs 2.1 ± 0.2 m 2 ; p = 0.08), or body mass index (30.5 ± 3.2 vs 30.7 ± 3.6 kg/m 2 ; p = 0.9). There was a trend toward a slight TU-induced increase in seated systolic BP (141 ± 17 vs 137 ± 16 mm Hg; p = 0.07); however, there was no difference in diastolic BP (83 ± 8 vs 83 ± 7 mm Hg; p = 0.97) or heart rate (62 ± 8 vs 64 ± 9 beats/min; p = 0.37).
There was no significant difference between treatments for any of the 36-Item Short Form Health Survey, Cardiac Health Profile, or Aging Male Symptom Scale measures examined and no difference in incidence of angina symptoms (2 ± 5 vs 1 ± 2 per 8-week treatment period, TU vs placebo, respectively). There were no other symptoms considered related to study drug. One patient who experienced "almost a depression" was administered placebo and did not cross over to TU treatment.
Our results show that 8 weeks of oral TU in men with CHD and low testosterone modestly increased myocardial perfusion in myocardium supplied by unobstructed coronary arteries, whereas perfusion in areas of myocardium supplied by coronary arteries with significant coronary atheroma was not affected. To our knowledge, this is the first investigation of blood flow effects of testosterone in territories associated with atherosclerotic coronary disease in humans. TU treatment decreased basal peripheral and central arterial stiffness, and concurrently, there was a comparative increase in LV ejection fraction. There was a null effect on overall myocardial perfusion, global endothelial function, quality of life, and angina symptoms compared with placebo. We report for the first time the effects of testosterone on myocardial perfusion in myocardial territories supplied by atherosclerotic coronary arteries. A previous study showed that acute intracoronary testosterone induced coronary vasodilation and increased coronary blood flow in unobstructed coronary arteries in patients with CHD in other major epicardial coronary arteries, but flow responses were not measured in significantly diseased arteries. In line with the intracoronary study, the present longer term study showed testosterone-related enhancement of myocardial blood flow in normally perfused regions of myocardium that may be mediated through effects on the epicardial coronary arteries or/and microvasculature. Our previous short-term intracoronary testosterone study showed no effect on coronary microvessels or coronary flow reserve. However, the different testosterone preparations and durations of treatment used may have had differing effects. A previous study of perfusion CMR in patients with cardiac syndrome X showed subendocardial hypoperfusion in response to adenosine, but no such effect was shown in the present study. It is perhaps surprising that we were unable to identify enhanced perfusion in myocardial segments associated with significant atherosclerosis because testosterone therapy in the longer term has beneficial effects on signs of exercise-induced myocardial ischemia. TU-related decreases in basal arterial stiffness imply that the mechanism involved in this anti-ischemic effect may at least in part involve effects on vascular tone (with no effect on BP or heart rate), thereby decreasing afterload. However, our myocardial perfusion data correspond with the null effect of testosterone treatment on angina symptoms shown in both the present study and other reports.
Puberty in males is associated with an increase in large-vessel arterial stiffness relative to prepubertal males, related at least in part to changes in sex-steroid milieu. However, a recent longitudinal study in older men (mean age 68 years) reported an independent inverse relation between endogenous serum testosterone and arterial stiffness. Our results provide evidence that oral TU treatment has a beneficial effect on arterial stiffness in older men with testosterone at or less than the lower limit of normal range. In light of recent research showing that arterial stiffness was a predictor of cardiovascular outcomes, our results may suggest that in the longer term, oral TU could potentially confer cardiovascular benefit through its effects on arterial stiffness. The endothelium did not appear to mediate effects of oral TU on arterial stiffness or vascular tone in vivo in humans, shown by the present study and others, although animal studies show that nitric oxide, prostacyclin, endothelium-derived hyperpolarizing factor, endothelin, and thromboxane-A2 mediate testosterone-induced vascular effects. The lower augmentation index measured after TU treatment could be caused by a combination of decreased wave reflection (possibly from an effect on small vessels) and delayed wave reflection suggestive of lower pulse-wave velocity from an effect on larger vessels (in the absence of a BP change).
Oral TU treatment resulted in a small but significant increase in LV ejection fraction compared with placebo, with no associated difference in LV mass, stroke volume, or end-systolic ordiastolic volumes. This is interesting in light of a report of decreased LV ejection fraction in patients with low endogenous free testosterone and may be caused by a direct action of TU on myocardial contractility. Castration was associated with decreased cardiac performance in male rats with slightly improved functioning in castrated testosterone-replaced animals. These effects may be explained by testosterone-induced actions on calcium-myosin adenosine triphosphatase activity. However, data are scarce and additional studies are needed, particularly investigating the actions of physiologic testosterone replacement on cardiac function in humans.
Plasma total testosterone levels did not increase after TU treatment because TU is absorbed into the lymphatic system with newly formed cholymicrons, then subsequently absorbed into blood from the intestine and rapidly converted into dihydrotestosterone. In this way, hepatic first pass is avoided. Dihydrotestosterone significantly increased and sex hormone-binding globulin, luteinizing hormone, and follicular-stimulating hormone decreased after TU administration.
The rationale for the 8-week treatment period was based on our knowledge of the possible longer term mechanisms of action of ovarian and related steroids on the vascular system after long-term exposure. We therefore wished to expose patients to several weeks of therapy, and the available data suggested a minimum period of 4 to 8 weeks. It is possible that testosterone therapy for longer periods may show different results. Long-term testosterone treatment is likely to affect aspects of human cardiovascular physiology that were not measured in the present study, but that may influence myocardial perfusion and vascular function. CMR perfusion techniques are rapidly developing, and although newer methods may not change the overall results, they may give more comprehensive information, for example, single versus multiple slices. The anticipated changes in myocardial perfusion and those shown were relatively small and it may be of interest to perform similar studies using different perfusion techniques.
Published as: Am J Cardiol. 2008 March 01; 101(5): 618-624.
Published as: Am J Cardiol. 2008 March 01; 101(5): 618-624.
Published as: Am J Cardiol. 2008 March 01; 101(5): 618-624.
Published as: Am J Cardiol. 2008 March 01; 101(5): 618-624.Sponsored DocumentSponsored DocumentSponsored Document
Background: Little is known about associations of gestational weight gain (GWG) with long-term maternal health. Objective: We aimed to examine associations of prepregnancy weight and GWG with maternal body mass index (BMI; in kg/m 2 ), waist circumference (WC), systolic blood pressure (SBP), and diastolic blood pressure (DBP) 16 y after pregnancy. Design: This is a prospective study in 2356 mothers from the Avon Longitudinal Study of Parents and Children (ALSPAC)-a populationbased pregnancy cohort. Results: Women with low GWG by Institute of Medicine recommendations had a lower mean BMI (21.56; 95% CI: 22.12, 21.00) and WC (23.37 cm; 24.91, 21.83 cm) than did women who gained weight as recommended. Women with a high GWG had a greater mean BMI (2.90; 2.27, 3.52), WC (5.84 cm; 4.15, 7.54 cm), SBP (2.87 mm Hg; 1.22, 4.52 mm Hg), and DBP (1.00 mm Hg; 20.02, 2.01 mm Hg). Analyses were adjusted for age, offspring sex, social class, parity, smoking, physical activity and diet in pregnancy, mode of delivery, and breastfeeding. Women with a high GWG had 3-fold increased odds of overweight and central adiposity. On the basis of estimates from random-effects multilevel models, prepregnancy weight was positively associated with all outcomes. GWG in all stages of pregnancy was positively associated with later BMI, WC, increased odds of overweight or obesity, and central adiposity. GWG in midpregnancy (19-28 wk) was associated with later greater SBP, DBP, and central adiposity but only in women with a normal prepregnancy BMI. Conclusions: Results support initiatives aimed at optimizing prepregnancy weight. Recommendations on optimal GWG need to balance contrasting associations with different outcomes in both mothers and offspring.
Few studies have examined the long-term effects of gestational weight gain (GWG) on maternal health. A systematic review examining child and maternal outcomes of GWG identified 5 studies that looked at the association of GWG with long-term ( 3 y) weight retention, 4 studies of the association of GWG with interpregnancy weight retention, and 1 study of the association of GWG with premenopausal breast cancer (1). Both this systematic review (1) and the 2009 US Institute of Medicine (IOM) guidelines on GWG (2) highlight the need for further high-quality research regarding associations of GWG with longterm maternal outcomes (2). Similarly, the new National Institute for Health and Clinical Excellence (NICE) guidelines on weight management before, during, and after pregnancy note the lack of evidence on whether adhering to the IOM guidelines is associated with benefit to mothers and offspring and whether these guidelines are applicable to the United Kingdom (3).
Since publication of the aforementioned systematic review (1), Mamun et al (4) have reported a positive association of GWG with weight retention 21 y postpartum in a cohort of 2055 Australian women. That study used just 2 measurements of gestational weight and thus was unable to explore different patterns of GWG in relation to later weight retention. Further indirect evidence of an association between GWG and long-term weight gain stems from studies examining the association between parity and postpartum weight gain (5,6). To the best of our knowledge, associations between GWG and cardiovascular disease risk factors or disease in later life have not been studied previously but are plausible given the existing evidence (even if scant and indirect) of positive independent associations of GWG with postpartum weight retention. The aim of this study was to examine the associations of prepregnancy weight and GWG with maternal body mass index (BMI), waist circumference (WC), and blood pressure (BP) measured some 16 y after pregnancy.
The Avon Longitudinal Study of Parents and Children (ALSPAC) is a prospective population-based birth cohort study that recruited 14,541 pregnant women resident in Avon, United Kingdom, with expected dates of delivery between 1 April 1991 and 31 December 1992 (http://www.alspac.bris.ac.uk). There were 13,617 mother-offspring pairs from singleton live births who survived to 1 y of age; only singleton pregnancies are considered in this article. We further restricted the analyses to women with term deliveries (between 37 and 44 wk of gestation; n = 12,976). Ethical approval for this study was obtained from the ALSPAC Law and Ethics Committee and the Local Research Ethics Committee.
Six trained research midwives abstracted data from obstetric medical records. No between-midwife variation in mean values of abstracted data and repeat data entry checks demonstrated error rates consistently ,1%. Obstetric data abstractions included every measurement of weight entered into the medical records [median number of repeat measurements per woman: 10, interquartile range (IQR): 8, 11] and the corresponding gestational age and date.
From age 7 y, surviving offspring, with parental consent, were invited to regular follow-up clinics. While offspring attended the 15-y follow-up clinic (n = 5509), clinic staff measured the accompanying adult's weight, height, and BP if time permitted. None of the accompanying biological mothers who were asked to participate declined consent. Of 4279 biological mothers who accompanied their child to the clinic, BP was measured in 3877, height and weight in 2401, and WC in 1619.
Weight and height were measured while the subjects were wearing light clothing and no shoes. Weight was measured to the nearest 0.1 kg by using Tanita scales (Tanita Europe BV, Amsterdam, Netherlands). Height was measured to the nearest 0.1 cm by using a Harpenden stadiometer (Holtain Ltd, Crymych, United Kingdom). WC was measured to the nearest 1 mm at the midpoint between the lower ribs and the pelvic bone with a flexible tape. Seated BP was measured by using a Dinamap 9301 Vital Signs Monitor (Morton Medical, London, United Kingdom). Two readings of SBP and DBP were recorded, and the mean is used here.
Maternal age, parity, mode of delivery (cesarean or vaginal delivery), diagnosis of diabetes, BP, and proteinuria at each antenatal visit and the child's sex were obtained from the obstetric records. On the basis of questionnaire responses, the highest pa-rental occupation was used to allocate the children to family social class groups [classes I (professional/managerial) to V (unskilled manual workers)] by using the 1991 British Office of Population and Census Statistics classification. Information on height, prepregnancy weight, maternal smoking in pregnancy, physical activity and diet in pregnancy, duration of breastfeeding, and current smoking was obtained from questionnaire responses. Maternal smoking in pregnancy was categorized as never smoked, smoked before pregnancy or in the first trimester and then stopped, or smoked throughout pregnancy. Physical activity in pregnancy was assessed at 18 wk of gestation, expressed in average metabolic equivalents (METs) (7) and categorized into fifths. Energy intake was assessed by using a food-frequency questionnaire at 32 wk of gestation and adjusted for underreporting as previously reported (8). Duration of breastfeeding (at 15 mo) was categorized as never, 0-3 mo, 3-5 mo, or 6 mo. Current smoking (assessed 12 y after pregnancy) was categorized as none, and the smokers' distribution was divided into quartiles on the basis of number of cigarettes smoked per week. The index pregnancy was considered to be the last pregnancy if women reported no additional pregnancies by 134 mo after the index pregnancy.
Associations of GWG with outcomes were studied by using the new 2009 IOM definitions of recommended GWG (2) and serial measurements of maternal weight that allowed us to study the effect of the timing of GWG on outcomes. To allocate women to IOM categories of lower than, recommended, and higher than recommended GWG, we used weight measurements from the obstetric notes and subtracted the first from the last weight measurement in pregnancy to derive absolute weight gain. Prepregnancy BMI was based on the predicted prepregnancy weight from the multilevel models (see below) and maternal report of height. Selfreported and predicted prepregnancy weight were highly correlated (Pearson's r = 0.92).
As previously described, all pregnancy weight measurements [median number of repeat measurements per woman: 10; interquartile range (IQR): 8, 11] were used to develop a linear spline multilevel model (with 2 levels: woman and measurement occasion) relating weight (outcome) to gestational age (exposure), with knots at 18 and 28 wk (9). The knots were placed to best reflect the observed data. Here, additional data were available compared with a previous publication examining offspring outcomes (9); hence, the difference in the placement of the knots. This multilevel model was then used to predict for each woman her weight at 0 wk gestation (referred to as "prepregnancy weight") and GWG (per week) from 0 to 18 wk (early-pregnancy GWG), 19 to 28 wk (midpregnancy GWG), and 29 wk to delivery (late-pregnancy GWG). We scaled maternal prepregnancy weight and gestational weight change to facilitate clinical interpretation, examining the variation in outcomes per additional 1 kg of maternal weight at conception and per 400-g gain per week of gestation for GWG (the recommended rate of GWG in women with a normal prepregnancy BMI) (2).
Complete data on GWG, outcomes, and confounders were available for 1397 women for associations with later BMI, 978 women for WC, and 2200 women for BP. We compared the characteristics of women in our subsample (women with data on any of the outcomes and all confounders; n = 2,356) with women excluded from the subsample because of either missing outcome or confounder data but for whom GWG data were available when linear or logistic regression was used as appropriate. Associations of outcomes with the IOM categories and with the individual estimates of maternal prepregnancy weight and early-, mid-, and latepregnancy GWG, estimated from the multilevel models, were undertaken by using linear and logistic regression. In the basic model we adjusted for the potential confounders age at outcome measurement and offspring sex. For the multilevel model exposures only, we adjusted for prepregnancy weight and GWG in previous periods in a second model. In the fully adjusted model, we also adjusted for the following potential confounders: parity, pregnancy, smoking, total caloric intake, physical activity, social class, and current smoking (via its association with prepregnancy smoking status), mode of delivery, and duration of breastfeeding as a potential mediator. In models using IOM categories, we also adjusted for prepregnancy BMI and gestational age (this is taken account of in the multilevel models). We also examined whether there were interactions between GWG and being normal weight (BMI , 25) or overweight/obese during prepregnancy (BMI 25) (10). For our main analyses, we examined BMI and WC after pregnancy as continuously measured variables. We also examined binary outcomes of overweight/obesity and central adiposity (WC 80 cm) (11). Associations of absolute GWG with outcomes were tested for linearity by using models with fractional polynomials.
We repeated the analyses including only women for whom the index pregnancy was their last (n = 966). We also repeated the analyses excluding women who did not gain weight during pregnancy (n = 7) and women who gained .30 kg (n = 3), because these values may indicate an underlying pathology, and excluding women with diabetes (n = 23) and preeclampsia (n = 43).
We used multivariate multiple imputation to impute missing variables (mostly outcome data) for participants with measures of weight in pregnancy (n = 12,447), including all exposures, covariables, outcomes, and potential predictors of missing data in the imputation equations (see online supplement) (12). The multiple multivariate imputation approach creates many copies of the data (in this case, 30 copies), each of which has missing values imputed, with an appropriate level of randomness, by chained equations. We repeated the analyses using the multiple imputation data sets. The results were obtained by averaging estimates across the 30 imputed data sets by using Rubin's rules, and the procedure takes into account uncertainty in the imputation (12). All analyses were conducted by using Stata version 11.0 (StataCorp, College Station, TX).
GWG measured in the study cohort was compared with the 2009 IOM recommendations (Table 1). In our cohort, women who were underweight before pregnancy and women with normal prepregnancy BMI had a mean GWG within the recommended range, whereas overweight and obese women gained more than recommended on average (Table 1). Women with less than the recommended GWG (n = 825), the recommended GWG (n = 957), and higher than the recommended GWG (n = 574) had mean weight gains of 8.4 (range: 26.9 to 12.4), 13.1 (5.0-18.0), and 17.7 (9.1-33.5) kg, respectively. Other characteristics according to IOM categories are presented in Table 2.
The mean difference in BMI, WC, SBP, and DBP and the risk of overweight/obesity and central adiposity 16 y after the index pregnancy are shown in Table 3 by IOM categories of GWG. Women with lower than recommended GWG had lower BMI, lower WC, and a reduced risk of overweight/obesity than did women with recommended GWG in both age and offspring sex and in fully adjusted models. Conversely, women who gained more than recommended had a higher mean BMI, WC, SBP, and DBP and a greater risk of overweight/obesity and central adiposity than did women with the recommended GWG. The odds of overweight/obesity and central adiposity in women with higher than recommended GWG were 3 times those in women who gained as recommended.
In Table 4, associations of prepregnancy weight and GWG with outcomes are examined by using the estimates from the multilevel models. Interactions between being overweight/obese (BMI 25) before pregnancy and GWG were noted for mid-pregnancy (weeks 19-28) GWG in relation to BMI, SBP, DBP, and central adiposity and between being overweight/obese before pregnancy and GWG in late pregnancy (weeks 29) in relation to DBP and WC. Hence, results stratified by whether women were normal or overweight/obese before pregnancy are presented for these associations. Prepregnancy weight was positively associated with all outcomes in both the age-and offspring sex-adjusted model (model 1) and in the fully adjusted model (model 3). GWG in all periods of gestation were positively associated with BMI; however, in midpregnancy (19-28 wk gestation), this association was only present in normal-weight women. GWG in all periods of gestation were positively associated with WC when adjusted for potential confounders and mediators. With the exception of a positive association of GWG in midpregnancy (19-28 wk) with BP in women who were normal weight, no strong evidence indicated that GWG in any time period was associated with SBP or DBP. GWG in all 3 periods of gestation was associated with greater odds of overweight/obesity and central adiposity, although the association of midpregnancy GWG with central adiposity was only apparent in women with normal weight. A stronger association of GWG in late pregnancy (29 wk) with central adiposity was observed in women who were overweight/ obese before pregnancy than in women with a normal prepregnancy BMI. We found no strong evidence (all P 0.05 for the linear compared with the best-fitting nonlinear model) for nonlinearity of associations of GWG with outcomes.
The results were minimally changed when women with diabetes and women with preeclampsia were excluded from the analyses, when women with extreme GWG values (0 and .30 kg) were excluded, or when the analyses were limited to women for whom the index pregnancy was the last (results available from authors on request).
A comparison of the subgroup of women included in our analyses with those who were excluded, mainly because of missing outcome data, and the differences between the 2 groups are shown elsewhere (see Supplemental Table 1 under "Supplemental data" in the online issue). However, the distributions of BMI, WC, SBP, and DBP in our subsample and in the imputed data sets were very similar (see Supplemental Table 2 under "Supplemental data" in the online issue). Moreover, when we repeated the analyses using the imputed data sets (see Supplemental Tables 3 and 4 under "Supplemental data" in the online issue), the results were essentially the same as those presented here, except for one difference. In overweight women, we found a strong (odds ratio: 4.70; 95% CI: 2.60, 8.50) association between GWG in midpregnancy and central adiposity in model 3, which was not noted in the complete case analysis.
In this contemporary cohort, we have shown that women with lower than recommended GWG (according to the 2009 IOM guidelines) had a lower mean BMI and WC, 16 y after the index pregnancy, than did women who gained as recommended. However, women with a higher than recommended GWG had a greater mean BMI and WC and a higher mean SBP and DBP than did women with the recommended GWG after adjustment for potential confounders. In more detailed analyses, we found positive associations of prepregnancy weight with all outcomes and that GWG in early and late pregnancy were positively associated with overweight/obesity and central adiposity, and mid-pregnancy GWG was positively associated with overweight/obesity, central adiposity, SBP, and DBP in women who were normal weight before pregnancy.
Two previous studies examined the association between GWG using IOM categories and weight (13)/BMI (14) 15 y after pregnancy and another with BMI 21 y after pregnancy (4). Similarly to our results, all studies found that women with higher than recommended GWG [by 1990 IOM categories (13, 14) and 2009 IOM categories (4)] subsequently had a higher BMI/weight than did women who gained as recommended. Conversely, and similarly to us, 2 studies (13, 14) found that women with a lower than recommended GWG had lower BMI/weight 15 y after pregnancy than did women with the recommended GWG. No strong evidence of this was found by Mamun et al (4) for BMI measured 21 y after pregnancy. None of these studies had the detailed repeat measurements of gestational weight that we were able to examine in our study, and none examined associations with postpregnancy WC or BP.
Because the IOM criteria combine prepregnancy BMI and GWG and recommend lower levels of weight gain in women who are already overweight or obese before pregnancy, it is possible that the associations of IOM categories with outcomes are driven by either prepregnancy BMI or GWG. In particular, the category of higher than recommended GWG may include more women who are overweight or obese before pregnancy because these have a lower threshold to reach to enter this category. In our study, more overweight women were indeed classified as having a higher than recommended GWG than were normal-weight women (51% compared with 19%); however, in detailed analyses (Table 4) there were positive associations of GWG in mid-pregnancy with some outcomes only in women with a normal prepregnancy weight, which existed after adjustment for prepregnancy weight. Hence, GWG is associated with outcomes in addition to prepregnancy weight.
Several potentially non-mutually exclusive mechanisms could explain our findings. First, they could reflect the tracking in individual size across the life course, with greater BMI associated with a susceptibility to greater GWG. However, prepregnancy BMI is inversely associated with GWG (2), and we found evidence of independent associations of prepregnancy weight and GWG with outcomes in later life. Alternatively, women with greater prepregnancy BMI may continue to engage in lifestyles (high-energy diet and low levels of physical activity) during and/or after their pregnancy that promote greater GWG and higher BMI and WC in the longer term. However, we did adjust for caloric intake and physical activity in pregnancy in our analyses. Our finding that GWG in midpregnancy was more detrimental to women who were not overweight before pregnancy was perhaps counterintuitive because thinner women are "allowed" to gain more weight during pregnancy than are heavier women according to IOM recommendations. However, this GWG may reflect greater maternal fat accretion: several (15)(16)(17), although not all (18), studies have shown inverse associations between gestational fat accretion and maternal prepregnancy obesity. Therefore, GWG by women who were overweight before pregnancy may be easier to lose. Consistent with this, in a prospective study of 405 Brazilian women, each 1-unit greater prepregnancy BMI was associated with a 0.5-kg lower weight retention at 9 mo postpartum (15). The results were obtained from linear or logistic regression models. SBP, systolic blood pressure; DBP, diastolic blood pressure.
A 400-g/wk change is the recommended rate of GWG in women with a normal prepregnancy BMI.
Adjusted for maternal age at outcome assessment and offspring sex.
Adjusted for maternal age at outcome assessment, offspring sex, prepregnancy weight, and GWG in previous exposure periods. 5 Adjusted for maternal age at outcome assessment, offspring sex, prepregnancy weight, GWG in previous exposure periods, head of household, social class, parity, smoking, energy intake and physical activity in pregnancy, mode of delivery, duration of breastfeeding, and current smoking. The availability of repeat measurements of weight in pregnancy allowed us to examine associations of GWG with outcomes in 3 different periods of gestation. Our results suggest that the strongest and most consistent associations of GWG with outcomes are in early and midpregnancy (0-28 wk). This may reflect that the maternal component of GWG is largely complete by 28 wk (19,20).
Our observed differences in GWG associations with later adiposity by prepregnancy overweight/obesity may reflect deposition and persistence of differential fat depots between normal and overweight women, because analyses of skinfold thicknesses have shown that obese women put on more fat in the upper body compartment, but that lean women put on more fat in the lower body compartment (17,21,22).
This study, which incorporated repeat measurements during pregnancy, allowed detailed temporal analysis of the long-term associations of GWG with several cardiovascular disease risk factors. The main limitation was that we used an opportunistic subsample that differed from women not included in the analyses (see Supplemental Table 1 under "Supplemental data" in the online issue). However, there is no reason to believe that associations would be different between these 2 groups or that whether women were included in the sample or not (ie, accompanied their child to the clinic or not) was related to their outcome measures. Moreover, results using imputed data sets were substantially the same as those presented here. Prepregnancy height was self-reported, and prepregnancy weight was estimated from models of weight gain during pregnancy; however, predicted and self-reported weight were highly correlated. We are unable to identify the separate maternal and fetal contributions to overall GWG, but our ability to divide GWG into 3 separate periods of gestation helped with the interpretation of our findings. Moreover, it would have been advantageous to have more detailed repeat measurements of postnatal weight in these women to assess whether the observed association of GWG with BMI later in life is due to weight gained later in life or to GWG never being lost. Unfortunately, such data are unavailable here and to the best of our knowledge in any other study. Because of sample size limitations, we grouped obese and overweight women into one category. However, sensitivity analyses showed that results for this group were not driven solely by the obese women. Finally, we did not have data on GWG and other pregnancy characteristics of subsequent pregnancies occurring between the index used here and the outcome assessment. Yet, results were essentially the same when we restricted our analyses only to the subgroup of women for whom this was their last pregnancy.
In summary, our finding that greater prepregnancy weight is associated with greater adiposity and higher BP 16 y after pregnancy, together with our recent findings from this cohort of increased obesity and cardiovascular disease risk factors in the offspring of women who had greater prepregnancy weight (9), support initiatives aimed at optimizing prepregnancy weight. If associations of higher than recommended GWG with long-term maternal and offspring health (9) are replicated in further studies, regular monitoring of weight in pregnancy in the United Kingdom may need to be reconsidered because it may provide a window of opportunity to intervene to prevent adverse health outcomes later in life. That said, whereas lower GWG may be beneficial for some outcomes, it is detrimental to others, particularly in offspring (2).
Hence, it is important to recognize that identifying an ideal GWG has to reflect these competing risks.
n in brackets. ALSPAC, Avon Longitudinal Study of Parents and Children.GESTATIONAL WEIGHT GAIN AND LONG-TERM MATERNAL HEALTH
P values for continuous variables are from a Bonferonni post hoc test and for categorical variables are from a chi-square test.
Mean 6 SD (all such values).
Results adjusted for underreporting, and P values obtained from regression analysis.
Mean; range in parentheses (all such values).
The results were obtained from linear or logistic regression models. SBP, systolic blood pressure; DBP, diastolic blood pressure.
Adjusted for maternal age at outcome assessment, gestational age, and offspring sex.
Adjusted for maternal age at outcome assessment, gestational age, offspring sex, head of household, social class, parity, smoking, energy intake and physical activity in pregnancy, mode of delivery, duration of breastfeeding, and current smoking.
Values are odds ratios (95% CIs); the reference (null) is 1, with values .1 indicating a greater risk and values ,1 indicating a reduced risk.
We are grateful to all the families who took part in this study, the midwives for their help in recruiting them, and the whole ALSPAC team, which includes interviewers, computer and laboratory technicians, clerical workers, research scientists, volunteers, managers, receptionists, and nurses.
The authors' responsibilities were as follows-AF and DAL: conceived and designed the study; AF,
From the
Late childhood and early adolescence represent a critical transition in the developmental and academic trajectory of youth, a time in which there is an upsurge in academic disengagement and psychopathology. PAR projects that can promote youth's sense of meaningful engagement in school and a sense of efficacy and mattering can be particularly powerful given the challenges of this developmental stage. In the present study, we draw on data from our own collaborative implementation of PAR projects in secondary schools to consider two central questions: (1) How do features of middle school settings and the developmental characteristics of the youth promote or inhibit the processes, outcomes, and sustainability of the PAR endeavor? and (2) How can the broad principles and concepts of PAR be effectively translated into specific intervention activities in schools, both within and outside of the classroom? In particular, we discuss a participatory research project conducted with 6th and 7th graders at an urban middle school as a means of highlighting the opportunities, constraints, and lessons learned in our efforts to contribute to the high-quality implementation and evaluation of PAR in diverse urban public schools.
Participatory Action Research (PAR) can be broadly characterized as a ''theoretical standpoint and collaborative methodology that is designed to ensure that those who are affected by the research project have a voice in that project'' (Langhout and Thomas, this issue, p. xx). With longstanding practice in community psychology, public health, adult education, and international development as a means of engaging marginalized populations in projects that address conditions of oppression, PAR approaches are becoming more common in other fields such as youth development. Other terminologies are used across disciplines, such as community-based participatory research (CBPR) in public health, Freireian models in adult education, and youth-led participatory research in youth development (Checkoway et al. 2003;Freire 1970;Ginwright et al. 2006;Israel et al. 1994Israel et al. , 2003;;London et al. 2003;Minkler and Wallerstein 2003;Schensul et al. 2004).
Common across these approaches is a focus on engaging in a cooperative, iterative process of research and action in which non-professional community members are trained as researchers and change agents, and power over decisions affecting all phases of the research and action are shared equitably among the partners in the collaboration (Israel et al. 1994(Israel et al. , 2003)). This process intends to provide opportunities for (typically disenfranchised) community members to work together to solve problems of concern to them, develop relevant skills, increase their understanding of their sociopolitical environment, and create mutual support systems (Israel et al. 1994(Israel et al. , 2003;;Zimmerman 1995).
There has been recent attention to PAR as a promising practice in the youth development field. The majority of case reports and other literature have focused on the involvement of youth as evaluators or co-evaluators of after-school and other programs intended to serve youth (Checkoway et al. 2003;Checkoway and Richards-Schuster 2004;Holden et al. 2004;London et al. 2003;Suleiman et al. 2006). There has also been growth in efforts to promote youth ''voice'' and engagement in improving schools, particularly with youth whose families and communities are politically marginalized due to racism and poverty (Cammarota and Fine 2007;Cargo et al. 2003;Kirshner 2007;Mediratta et al. 2008;Mitra 2004;Morrell 2007;Nieto 1996;Schensul et al. 2004;Shor 1996). The principles of PAR with youth parallel those for adults, with some variation to reflect the diminished power of youth in adult-supervised settings. Given the heavy skill demands of research and advocacy, a major degree of ''scaffolding'' and alliance-building with adults is expected to occur for youth in the PAR process (Larson et al. 2005;Vygotsky 1978).
Core processes of youth PAR involve the training of young people to identify major concerns in their schools and communities, conduct research to understand the nature of the problems, and take leadership in influencing policies and decisions to enhance the conditions in which they live (London et al. 2003). Key features include an emphasis on promoting youth's sense of ownership and control over the process, and promoting the social and political engagement of youth and their allies to help address problems identified in the research.
Schools are a critical setting for child development, and are intricately related to key developmental milestones in late childhood and adolescence, including academic achievement, peer relationships, pro-social conduct, and involvement in athletics and clubs (Eccles and Roeser 2003;Masten and Coatsworth 1998). Since the early 20th century, theorists and researchers have made cogent critiques regarding the shortcomings of U.S. schools as contexts for the development of children as learners, thinkers, and members of a democratic society (Dewey 1938(Dewey /1997;;Kozol 1991;Sarason 1996;Weinstein 2002a). As Sarason observed, the typical classroom is one in which teachers rather than students ask questions, adults are rendered ''insensitive to what their [children's] interests, concerns and questions are…and children are viewed as incapable of self-regulation'' (Sarason 1996, p. 363). These features of school culture and teacher practices have proven highly resistant to change in public schools, despite the efforts of educational reform movements.
Recent analyses of schools have also emerged from the growing field of youth development. With the goal of informing the practices of school and other youth-serving organizations, reviews of the child development literature have identified the qualities of settings-and the types of transactions between youth and their environments-that are associated with positive development (Eccles and Gootman 2002;Granger 2002). These dimensions are summarized in terms of physical and psychological safety, appropriate structure, supportive relationships, opportunities to belong, positive social norms, support for efficacy and mattering, opportunities for skill building, and the integration of family, school, and community efforts.
As evident from the above discussion, conducting PAR with youth in schools is a process that could potentially target virtually all of the setting-level features that have been linked with healthier development. PAR most explicitly addresses the dimensions of opportunities for developmentally-appropriate and meaningful youth participation, support for efficacy and mattering, the development of relevant skills; and supportive relationships. Examined in light of Sarason's critique of typical classrooms, the PAR process is clearly counter-cultural insofar as it is fundamentally about student-led inquiry, the valuing of students' concerns and expertise, and the opening up of opportunities for students to take on tasks and roles that involve self-regulation as well as participation in governance.
Although PAR presents both developmental opportunities and challenges to school culture in any traditional school environment, we submit that this is particularly the case in middle school settings. There is extensive literature establishing that the transition to middle school in late childhood and early adolescence is a crucial period in the trajectory of intellectual and psychosocial development. This period of development has been associated with reductions in academic motivation and achievement (Eccles et al. 1993;Simmons 1987) as well as increases in depression and other psychological problems (Compas et al. 1997;Galaif 2007). The theoretical and empirical work of Eccles and colleagues provides support for the ''developmental mismatch'' hypothesis-that is, that at least some proportion of the problems in motivation and engagement in school that emerge in early adolescence may be attributable to poor person-environment fit for children at the transition to middle school or junior high (Eccles et al. 1993). Although older children and young adolescents demonstrate growing capacity and desire for autonomy, longitudinal research indicates that youth perceive fewer opportunities to exercise autonomy and participate in making decisions and rules in junior high than they did in elementary schools (Midgley and Feldlaufer 1987). Further, there is evidence that appropriate levels of decision-making and control are particularly salient for motivation and learning in adolescence (Eccles et al. 1993).
The literature on identity formation in childhood and adolescence highlights the importance of a sense of purpose (Damon 2003), and the role of responsibility and service for helping to foster a sense of moral identity (Youniss and Yates 1997). Theoretical frameworks that specifically consider the identity development of youth of color highlight the importance of social position, racism and discrimination, and the influence of their immediate environments (Garcı ´a Coll et al. 1996) as well as youth's appraisals of themselves and of the risk and protective factors that affect their life trajectories (Spencer et al. 1997;Spencer et al. 2003). PAR's emphasis on engaging marginalized groups in research and action as a means of achieving social justice and equity goals enhances the potential developmental relevance of PAR processes for ethnic minority youth. PAR processes that involve youth of color in analyzing and having an impact on the social, economic, and political conditions that shape their schools and communities provide key developmental opportunities for youth to appraise themselves as leaders with a meaningful sense of purpose (Damon 2003;Spencer et al. 2003) as opposed to internalizing the negative stereotypes held by others (Cahill et al. 2008). Further, PAR processes should push youth beyond individual-level explanations of problems faced by communities of color to investigate broader explanatory factors, thereby influencing their appraisal of risk and protective factors.
Thus, in our view, middle schools appear to be settings that could benefit from PAR interventions to help reduce the ''developmental mismatch'' and promote positive identity development for diverse youth. However, in light of extensive evidence for the difficulties in enacting meaningful changes in schools as organizations (Evans 1996;Sarason 1996), we would expect to encounter significant challenges to successfully implementing and sustaining PAR efforts in such settings.
As discussed above, there is extensive theory that broadly informs PAR and other empowerment-oriented interventions (Freire 1994;Zimmerman 1995) and a rich theoretical and empirical literature on the social ecology of schools (Sarason 1996;Trickett and Todd 1972;Weinstein 2002a). There is a small but growing field of research on PAR projects implemented with young people in after-school or summer ''institute'' settings, some of which explicitly involve youth in research and advocacy to improve their schools (Morrell 2007;Wilson et al. 2007) or to address health, social justice, or employment issues that are relevant to their school experiences (Cammarota and Fine 2007; Wallerstein et al. 2004). This existing literature has presented curricula, documented meaningful narratives about the value of YPAR for the youth, adult collaborators, and communities, discussed challenges involved in implementing YPAR given the diverse capacities it demands for youth and adult facilitators, and provided initial evaluation of outcomes for individual participants (Berg et al. 2009;Wallerstein et al. 2004).
The present study builds on and provides a unique contribution to this literature by specifically considering the integration of PAR into the social ecology of classrooms and schools during the ''regular'' school day. Here, we draw on data from our own collaborative implementation of PAR projects in secondary schools to consider two central questions: (1) How do features of middle school settings and the developmental characteristics of the youth promote or inhibit the processes, outcomes, and sustainability of the PAR endeavor? and (2) How can the broad principles and concepts of PAR be effectively translated into key processes to guide intervention activities in schools, both within and outside of the classroom?
In section ''PAR in an Urban Middle School'', as a means of illustrating key opportunities and challenges regarding the implementation of PAR with youth in school settings, we discuss a participatory action research (PAR) project conducted with 6th and 7th graders at an urban middle school (''Tubman''). 1 This project represented the initial year in a multi-year effort involving the collaboration of the school staff, students, staff of a communitybased organization, and a university research team. In the ''Discussion'', building on the existing literature and our own research, we delineate a set of core processes of PAR interventions in schools that we are currently using to assess our work and offer as a potential contribution to this growing field. We then examine the strengths and weaknesses of the Tubman middle school project through the lens of this process framework, and discuss how our implementation and evaluation of our school-based PAR projects has evolved as a result of this early work.
The project discussed here was conducted at ''Tubman,'' a majority-Latino middle school located in a high-SES urban neighborhood in the San Francisco Bay Area. Tubman is in the small to medium range of size for the district, with *500 students in 6th through 8th grades. Two-thirds of the students are economically disadvantaged and one-third is classified by the school district as English language learners. The majority of students who attend the school do not live in the surrounding neighborhood. The intended outcomes of the PAR project, with respect to the school setting, were to establish opportunities for students to participate in school governance and shape school practices, via the sharing of research-based recommendations with administrators aimed at improving the school in areas of concern to the students. Other intended school-level effects included improving alliances between students and adult staff, creating opportunities for students and adults to engage together in inquiry about issues relevant to the school and to students, and enhancing the collective efficacy of students to enact thoughtful and high-quality research and advocacy activities. The intended outcomes for students included the strengthening of knowledge and skills regarding research, communication, strategic thinking, collaborative group work, and advocacy, enhanced sense of positive ethnic identity, sense of purpose, and connection to school, and increased motivation to influence the school setting. These targeted outcomes were part of a conceptual model (see Fig. 1a, b) developed by the first author on the basis of the literature on adolescent development, PAR, and psychological empowerment (Eccles and Gootman 2002;London et al. 2003;Rappaport et al. 1984;Zimmerman 2000) as well as prior school-based pilot work conducted by the first author in which students provided their perspective on the benefits of the YPAR process for them.
The project was part of a larger collaborative research effort to develop and evaluate PAR projects at 6 sites; it was undertaken by a partnership primarily consisting of a local community-based organization (CBO), the university-based team led by the 1st author of this paper, and the school site staff (see Ozer et al. 2008 for a discussion of the partnership and capacity-building efforts). The first author, a European American female from a middle-class background, lives near the school and had conducted research at the school several years prior as part of a prior study on stress and mental health.
We focus here on Year 1 of the collaborative project, a pilot phase in the overall study, in which our universitybased team consulted with a CBO to support the initial implementation of the PAR project in an elective peer counseling class at Tubman and at five high school sites. After 1 year of providing technical assistance and capacity building efforts directly to the teacher and her CBO-based supervisor at Tubman, the plan was for the PAR project at the school to continue-pending feasibility and teacher interest-with direct support from her CBO supervisor but not the university team. The university team continued to engage in an ongoing consultation and technical assistance relationship with the supervisor and the CBO. The CBO's goal was to infuse PAR into its existing youth development programs at 20 middle and high school sites in the school district.
The university team (1 European American male graduate student and 1 Latina female undergraduate, both trained in PAR and in observational research methods) made weekly visits to the Tubman classroom throughout Year 1 to provide technical support and help facilitate the PAR project. The research team also documented the process using observational field notes, which were summarized and excerpted to inform this case discussion. As discussed in detail below, a smaller group of 8 female students in the class were selected to participate in an additional series of small group meetings. Consistent with the policy of the school district and the university IRB, the parents/guardians of these students had provided their signed consent. In recruiting students for the group, the first author worked with the teacher to identify consented students that would reflect a range of academic achievement and whom would likely have some exposure to the problems to be addressed by the group. These meetings took place in the empty school cafeteria during their regular class time. These meetings were audio-taped and transcribed; transcriptions were reviewed to create summaries and quotes for this report.
One year after the end of the Year 1 consultation, the first author conducted follow-up interviews with the principal and teacher to learn about activities that were continued at the site after the consultation ended, the potential impact of the process, and challenges encountered. The first author also conducted an interview with the CBObased supervisor to elicit her insights regarding this project as well as broader lessons learned in supervising teachers conducting PAR projects in 20 middle and high schools. Unfortunately, it was not feasible to conduct follow-up interviews with the students who had participated in the project due to limitations in the initial consenting process that did not allow for re-contacting of the students.
The ethnically-diverse (Latino, Asian American, African American, and European American) class consisted of 32 students, and met daily as an elective class. The classroom space was very small, providing little space for movement or re-grouping of students in the classroom. This was the teacher's first year in a regular teaching position.
The curriculum used by the teacher had been adapted by the CBO, using the structure and activities from two published curricula developed by established youth-led PAR organizations (London 2001;Sydlo et al. 2000). The teacher and university team guided the students in mapping the resources and problems of the school and community, and identifying issues of concern for possible research and action. A range of problems were identified by the students, including drugs, pressure to join gangs, school food, and areas in need of physical improvement at the school (e.g., water fountains did not work, no nets on the basketball court, unclean bathrooms). The teacher and university team then facilitated class discussions to assist them in reaching a group decision about which issues to study more in-depth in their research project. The class ultimately chose to focus their efforts on improving school food and the physical plant of the school with the rationale that these issues affect most students and were potentially ''winnable.'' The class engaged in a photo-voice process (Wang and Burris 1994) to document the problem areas at the school site and develop a storyboard with photographs to share with the principal and other stakeholders. The products of this project were presented to the principal by representatives of the class near the end of the school year. Most of the class members also presented their methodology and findings to a conference sponsored by the CBO in which classes from all 6 schools that were part of the university-CBO-school partnership came together to share their PAR projects.
Several challenges were encountered in the implementation of this project in Year 1. First, although observations by the research team suggested that students' cognitive maturity
Psychological and political empowerment Perceived school connection Intervention Key Processes Youth-level Outcomes Expanded social networks and support Positive ethnic identity, Sense of purpose Skills, efficacy in research, communication, advocacy Positive class climate (e.g. engagement, student perspectives) Group work Opportunities for skill development (e.g. research, advocacy) Networking opportunities Teacher-student power-sharing Classroom PAR Settings Targeted School-Level Outcomes Classroom Alliances between students and adult staff Student -adult inquiry and learning Meaningful student roles in school policies and practices Collective efficacy of students for research, advocacy School PAR was a good fit for the PAR curriculum insofar as they demonstrated language and critical thinking skills that were ''as high or higher than in the high schools,'' the students' social maturity was noted to be uneven:
The challenge here isn't getting the students engaged, but managing their engagement….There is a marked difference between the boys and girls in the class….the boys are…poking each other, talking over each other, seeking attention; the girls are much calmer and more mature seeming. We may need different strategies for each group. (research memo) These and other comments in the research memos reflect the challenge confronted in this project of how to transfer appropriate levels of control to the students. Because the observers were also documenting PAR projects in 9th grade high school classrooms, there were comparisons drawn between the conditions in the middle school versus high school classroom:
The vibe in the class is drastically different than in HS classes. The students are more energetic and unfocused…with constant goofing around…it may be difficult to get them to buy into action research. It may be that we have to construct a series of activities that lead them to investigate, but that they aren't given the level of control of the HS students. (research memo)
The effectiveness of efforts to meaningfully engage less receptive students in this PAR project was undermined by the specific features of limited space and large class size in this urban middle school. The classroom space was inadequate for the class size, which contributed to poor conditions for high-quality discussion, particularly given the teacher's lack of experience in managing classroom dynamics. To increase engagement and student control over the projects, the PAR curricular activities called for dividing the class into small groups for interactive activities at multiple stages of the project. For example, in this research memo excerpt, a graduate student describes a small group exercise in which students drew a map of their school and what affects their health at school and then reported back to the larger class; the students identified as ''difficult'' by the teacher are Latino and African-American: [The teacher] had identified the three boys I was working [with] to all be difficult, and it was hard to get them on task. [Two students] were poking each other with pencils until I got serious and told them I ''didn't want to see it anymore.'' [One student] started out drawing the school and after asking them a lot of questions (…''do you think students here are healthy?'') they started to get into it. They got particularly energetic when talking about the food (''Man, the food is nasty. Why do we have to have the same stuff everyday''), the litter (''people leave their trash all outside and that's why the birds come and then they poop on us'') and the basketball hoops (''The hoop is all bent with no net, no that's not it, give me the pencil, I'll draw it.'')…. Everyone finished for the most part and it was time to pres-ent…The students had a lot of trouble speaking one at a time. It made it difficult to hear everything. [The teacher]…brought out a little sandbag thing -only the student holding it was supposed to be able to speak. It worked more or less. (research memo)
The small group format with students playing a more active role was generally effective in engaging students who appeared not to pay attention to whole-class activities and discussions. Simultaneous discussions, however, created an even higher noise level in the small space.
The memos above indicate that the research team initially attributed the need for more focused attention in engaging several of the Latino and African-American male students to lower maturity levels among boys. This interpretation did not consider the role of teacher capacity and the classroom physical space, nor did it consider how the broader social ecology of the school with respect to differential expectations and treatment for boys of color may have set the stage for their responses to the curriculum and the research team (see p. 23 for a deeper consideration of these issues). As discussed more fully later, the design of the present study does not enable us to disentangle the influence of students' development on the implementation of the YPAR project, versus the influence of the classroom and school conditions in which it was embedded.
Because the classroom size and environment were not conducive to in-depth discussion, the first author and the classroom teacher came up with a strategy to conduct more intensive work with a subgroup of students in the class. In the second semester of Year 1, the first author initiated additional weekly meetings with a group of 8 girls in the class (7 of whom were Latina) to understand more deeply several problems not being addressed by the whole class, and to come up with recommendations for policies or programs to address these issues.
These small group meetings were, not surprisingly, more fruitful than the discussions in the regular classroom context. Excellent rapport was developed between the first author and the students, and among the students, and there appeared to be substantially less concern with self-presentation in the all-female group than in the regular classroom. The teacher expressed to the first author that the students enjoyed having a group for only girls, and they appeared to relish the special role of meeting with the first author and being asked their opinions. A side benefit of this approach was that it temporarily relieved the overcrowding in the regular classroom for that session. Early on in the meetings, the first author shared that she also lived in the neighborhood and was a parent of a child in the public school system. Several of the students responded to this information by expressing surprise that a resident of the surrounding neighborhood would be interested in the students at the school; one said, ''I thought that all of the people in the neighborhood hated us.'' Class, ethnic, and language issues were raised at other times, most notably when the girls were discussing family relationships and several shared their view, learned from TV, that all ''White people were rich'' and that ''White families always sit down and eat dinner together,'' unlike their own families. Almost all of the students were bilingual, as is the first author, and there was some discussion about the value of both languages in the group that was prompted by some exchanges in Spanish by group members.
The small group meetings focused on several mental and physical health issues that had been identified as problems by students in the issue selection phase in the larger class project, but were not studied in the photo-voice project because of their complexity. For each issue, the first author facilitated the students' mapping of ''root'' causes in which they used the metaphor of a tree to name and depict the sources of the problem (the ''leaves'') in terms of factors on the level of the individual, family, school, community, and society. The first author then assisted the students in identifying and reflecting on existing efforts to address it, and in generating ideas for suggested improvement or novel interventions to address it further. These students continued to participate in the full-class photo-voice project on the days that they were not in the small group. The topics considered by the small group included the perceived pressure to join gangs for some Latino students at the school, substance use, and family conflict. Gangs and substance use were the issues that generated the most engagement and interest on the part of the students; we discuss the gang issue below to illustrate the process and the type of insights and recommendations that emerged from the group.
Pressure to join gangs or ''claim colors'' in gang rivalries in their neighborhoods were perceived as highly relevant to the group. Multiple girls reported that relatives or friends had claimed colors or had been pressured to join a gang. Although the school had a strict dress code to prevent the wearing of gang-related colors, students reported that they noticed peers starting to claim colors by ''small things…a hair tie or something, or shoes, or belt. I think the belt shows more, or carrying a bandana out of your pocket. '' Students discussed what they understood to be the root causes of gang involvement, as well as additional conditions that contributed to the problem. They cited a need for protection for recent immigrant youth, lack of awareness of how hard it is to get out of a gang, and lack of connection at home and school as major factors: They don't feel special at home or at school -and the gangs make people feel special and impor-tant…Sometimes parents tell them that they aren't going to do much in their life, that they don't have much to look forward to.
Students also noted that, for a small number of youth, adult family members are already gang-affiliated and membership is a natural thing: If ''part of your family is in a gang -they think it's in their blood.'' One of the lowerachieving students provided her view as to how the Latino honor roll, a school practice designed to encourage the academic engagement and success of Latino students, actually undermined the motivation of Latino students who were below the cut-off, saying that ''you feel like a moron'' if you don't make the list.
As indicated above, students identified a lack of meaningful connection to school and families as enhancing the perceived benefits of gang involvement, although they also acknowledged that even students who did have good ties to home and school might feel the need to claim colors because of pressure and threats. They expressed that hearing messages from adults about not joining gangs might actually make gangs seem more appealing because of ''reverse psychology'' and some youth's desire to do the opposite of what adults tell them to do. They therefore suggested that educational experiences that entailed hearing from young adults in their late teens or early '20's who ''have gone through this kind of thing'' and can talk about how hard it was to get out of gangs would be a good program for the school. The group also suggested ideas for rewarding students for academic improvement, even if they do not make the honor roll, so that they feel that this effort is acknowledged.
Having young adults to talk to at the school was a theme that cut across the group's recommendations for gang prevention, substance abuse, and other areas. They reported that, despite the presence of a school counselor, students feel that ''they have no one to talk to about their problems.'' Two main concerns about talking with teachers or counselors were raised: that the adult will tell their parents, because they think it will be for their own good, ''but this could make it a lot worse.'' The second reason that they expressed is that the counselor and many of the teachers ''wouldn't really understand because they are too old,'' and that they need a younger person whom they feel would really understand what they are going through.
This group, accompanied by the first author, met with the principal near the end of the school year to present their analyses of the issues facing students at Tubman and to share their suggested recommendations for addressing these problems. The principal listened attentively to the students' presentation and expressed appreciation for their feedback; he suggested that there be future meetings in which he could hear the perspectives of students. The students were initially nervous in their presentation but were able to articulate their views and participate in backand-forth exchanges with the principal that were respectful and substantive. Afterwards, they appeared excited and reported being pleased with having had the opportunity to talk with the principal in this way.
Interviews with the teacher, principal, and supervisor (all European American) conducted 1 year after the end of the university consultation in Year 1 of the project suggested several insights regarding PAR at Tubman. The teacher, who had fortunately moved to a much larger classroom space, reported that she had continued with the PAR efforts in the subsequent year by initiating a photo-voice project with her new cohort of students. The new project was not directly based on the work of the prior year's students but instead originated with the new cohort. She enthusiastically discussed plans to engage students in film as well as photography as part of their photo-voice project in her next cohort. She continued to engage youth in peer education and conflict resolution efforts at the school through the regular curriculum, although she did not articulate if or how the photo-voice project informed these educational and service efforts. After the Year 1 project, there was no follow up between the principal and teacher to define or implement specific objectives raised in his prior meeting with the students. Based on our interview and the teacher's planned curriculum, it appears to us that the teacher viewed photo-voice as a valuable tool for promoting the critical thinking and skills of students, but that the potential of PAR for increasing the participation of students in addressing concerns and improving the school was either not understood or was not a priority given limited resources and competing demands.
The principal, in his follow-up interview, reported that several of the problems that the students had identified in the Year 1 PAR project had improved. He attributed these improvements, however, not to suggestions made by the students but rather to broader efforts to improve school safety and climate across the district and at the school: ''Back then it was pretty rough…now…there are no hallways with garbage, the climate has changed. We still do have a lot of trash outside after lunch, and other prob-lems…getting high is a problem.'' In Year 2, the school had initiated a district-wide program to improve school climate that involved the training of students as peer leaders, but, according to the principal, it was difficult to sustain because the 8th graders were trained in the Spring just before graduation and then ''moved on.'' He reported being unaware that the photo-voice project had continued beyond Year 1. Despite the interest he had expressed earlier about having more regular meetings with students like the one in Year 1 with the PAR group, no additional meetings had occurred.
Although he observed that many middle school students are not as comfortable as high school students with ''getting serious'' in talking with adults, he expressed optimism about the PAR process as providing a formal means of student participation beyond the daily conversations that he has with students as a hands-on middle school principal with a regular physical presence at the site: ''I see value in student voice, and that adults are needed to structure the conversation, to express what it is that needs improvement. For students to think and speak like they are growing up. I think it can have an impact here, but hasn't really yet.'' Drug use and ''tagging'' (graffiti using writing and symbols) were two areas in which the principal expressed that he could benefit from PAR-generated data and recommendations:
Substance use is a big area. I would like to know the what and the when. When I talk to students this age about 'why,'' the why isn't too clear. Maybe using it as an escape from a bad home life, or from no home life. But to have them talk truthfully about this… Another area that I really don't understand why is tagging. It has zero redeeming value aestheticallyit's not graffiti -there is no pretense to art. I see it as vandalism. And when I find out who it is sometimes, it's like, ''they did all that?'' It's almost like it's random who did it.
We see several conditions illuminated at Tubman and at other sites that likely influenced the impact and sustainability of PAR, and that may be informative for the conceptualization and implementation of PAR in other school settings. Below, we discuss several features of the school setting that have a bearing on the PAR process for young people and their adult allies in schools, including: constraints on student autonomy, the academic calendar, the size and social network of the adults at the school, and the resources-physical, financial, and time-for nonacademic classes. Although these features would be expected to influence the implementation of any schoolbased intervention, our discussion highlights the ''innovation-specific'' (Wandersman et al. 2008) salience of these conditions for the diffusion of PAR in schools because of its emphasis on research and action intended to make change and disrupt the status quo (Ozer et al. 2008). Following our consideration of PAR within the social ecology of the Tubman case, we draw on the existing literature and our experiences to propose a set of key processes for the implementation of PAR in schools, examine the Tubman experience in light of these key processes, and propose next steps for research and practice.
As noted earlier, middle schools have generally been noted for their emphasis on behavioral control, adult-driven inquiry, and the lack of opportunities for youth to exercise developmentally-appropriate levels of autonomy. Implementing PAR with youth in this context is intended to address these conditions by providing youth-driven inquiry and the opening up of meaningful opportunities for youth to influence school policies and practices. There is a risk, however, that PAR projects within middle school settings will instead reinforce or replicate the ''adultist'' power dynamics that they are trying to address. In the Tubman example, the PAR project was implemented as part of an elective class that students theoretically ''chose'' to take; when they chose the class however, it was not clear to them that they would be conducting PAR. Even when they agreed to participate in the PAR project, it is likely that they did not fully understand what it would entail, and it would be difficult to change their class schedule mid-semester. Thus, although many students expressed enthusiasm for working to improve the school via the PAR project, the Tubman project-like other classroom-based PAR implementations that do not select a small group of motivated youth-faced the challenge of engaging students who were not fully invested in the project. 2Classroom observations suggested that the primary challenges to eliciting the active participation of some students-primarily boys of color-appeared to center on general issues of classroom engagement than objections to the PAR process itself. That is, observations of students who appeared to resist participation in the class activities at various times suggested that they became motivated once they were able to focus on the task at hand. There is extensive evidence for the differential expectations for and treatment of youth in color, particularly males, in U.S. schools (Gregory and Weinstein 2008;Skiba et al. 2002;Weinstein 2002b). Clearly, PAR projects must be differentiated from typical classroom relationships and curricula to avoid ''business as usual'' interactions and role demands from teachers and students alike.
As noted earlier, most of the adults directly involved in the project were from European American backgrounds; nearly all of the students were from ethnic minority backgrounds with a high proportion of Latino youth. It is possible that these differences, which roughly mirror the general demographics of the adults and youth in the school district, may have set the stage for replication of regular classroom dynamics for some students of color who were already experiencing disconnection and disinvestment from classroom activities. Working with students in smaller groups both within and outside of the classroom, however, at times enabled the Tubman project to create interactions that disrupted the typical pattern of engagement and offered opportunities for lower-achieving students to express themselves and contribute meaningfully to the larger project. These small group activities set the stage for peer-to-peer discussion and learning as opposed to questions and answers from the adult teacher.
Although lack of time is likely an issue for nearly every PAR project, whether conducted with adult community members or youth, the academic calendar and competing demands represent formidable challenges for school-based PAR. Unlike adult-led PAR and some youth-led PAR projects enacted in afterschool or other CBO settings, projects implemented in schools must operate within the hard deadlines and ''black-out periods'' of the academic calendar. As noted above, competing interests are particularly strong in schools that are under pressure to improve standardized test scores, where all instructional time must be mapped onto specific state standards. Although the inquiry methods and communication skills integral to PAR curricula likely strengthened learning of basic standards, the PAR process in the elective class at Tubman and other sites was not organized around state learning standards. An additional challenge inherent in the academic structure involves limitations in working with the same cohort over time. Although re-engaging students who have not graduated is possible (unlike an afterschool program, attendance is not optional), students' scheduling limitations may not permit them to re-enroll in the same elective class. This is particularly true for academically underperforming students with limited opportunities for electives.
Conducting PAR projects in after-school youth development programs or other community-based organizations rather than during the school day is an alternative that carries with it potential benefits and disadvantages. On the positive side, PAR projects implemented in afterschool settings likely benefit from greater freedom from the demands of instructional time (if not of homework completion) and from having longer blocks of time to conduct the work. In addition, an organization that continues to work with youth over several years can have greater continuity of effort over time. This structure can potentially minimize the challenges faced by the teacher at Tubman and at other school sites regarding how to promote the sense of ''ownership'' for the PAR process of new cohorts of youth while not abandoning the follow-up activities needed to turn the previous cohort's work into actual change (Ozer et al. 2008).
PAR projects located in after-school or CBO settings are likely to work with a selected group of students that may be less representative of the school. If the focus of the PAR effort is on change within the school, the project will be limited by a more restricted perspective on the school site (although this can potentially be addressed via research conducted by the youth with representative samples of students at the school). PAR programs with self-selected youth-as is the case in the overwhelming majority of programs studied in the existing literature-will also encounter challenges in quantitative evaluations of the impact of PAR on their youth participants because of selection bias.
Size and Social Network of School Tubman is a relatively small school, and the students benefited from access to the principal; he was frequently observed by our team to be interacting with students in the hallways and outside of the school immediately after the school day. In considering the implementation of PAR across 20 district sites, the CBO supervisor further emphasized the size of the school as a critical factor in terms of ''what you can pull off'' in middle school:
The smaller the middle school, the more likely you are to actually do something. They are tighter communities -the teachers know each other -it's easier to get the access you might want. The only way a small school isn't an advantage is if students want structural change. Small schools don't have much capacity for that -for example, if they come up with wanting more electives. At small schools, there is less flexibility on big issues than in big schools. But middle school projects don't tend to focus on structural issues. At a big middle school, there is such a focus on managing the behavior of students, that they aren't that into changes to the system, they are less open to it…and they don't want students in the hallways doing action research.
Another key factor cited by the supervisor beyond the size of the school was the existence of a group of teachers who were already interested in issues like school climate. Because students ''need a lot of adult support,'' this gives the students a committee or group whom they could immediately identify as potential allies. There was no such group at Tubman that was identified by the teacher in Year 1. This would have provided an excellent resource for the students and the teacher to engage other stakeholders, hear their perspectives, strategize about how to bring about desired changes, and provide continuity of effort from 1 year to the next.
The social ties and power of the specific teacher or other adult who facilitates the PAR process can shape it in several ways. First, elective teachers (like the teacher at Tubman) may be more isolated and lack a department of colleagues and potential allies that could help create receptive conditions for the PAR project. This would be expected to be a particular challenge for a new teacher, although prior research on the social organization of schools indicates that even an experienced teacher may have few opportunities to collaborate or exchange ideas with other teachers at the same site (Little and McLaughlin 1993;Lortie 1975). Second, as discussed in detail below, a teacher of an elective subject may be more likely to have a tenuous job at some school sites. Job insecurity can undermine long-range planning from 1 year to the next. It can also create concerns for teachers that they might experience negative repercussions if the students raise politically-sensitive issues in their PAR project.
The size and social network of schools would be expected to be relevant conditions that may influence the implementation of any school-based intervention; however, PAR's focus on increasing the power of youth in shaping school conditions, policies, and practices means that alliances with other stakeholders in the school and community and the capacity to sustain efforts over time are particularly crucial. Students engaging in PAR projects that seek to make changes in schools are operating with limited power in a politically-sensitive environment; forming alliances with more powerful stakeholders such as teachers and administrators and getting them ''bought in'' early on thus improves the likelihood of having a positive impact.
Although conducting PAR via an elective class already focused on youth development was an excellent fit and enabled most of class time to be devoted to the PAR project, it also created challenges in follow-through because of uncertain funding for electives at the case report site and other schools in this under-resourced urban district. At Tubman, there were ongoing negotiations about the role and resources of the elective peer resources program at the site. With the teacher and her supervisor focused on the role and sustainability of the basic program after Year 1, there was less time and focus available to promote the sustainability and impact of the PAR project. According to the supervisor, the issue of program stability is highly salient in determining the scope of work across their 10 middle school sites in the district: Most middle schools barely have electives. Because of test scores, many students are taking double English and double Math. Every principal is super under the gun to deliver test scores. What is possible [in a PAR project] depends on if the school has turned it around or not. Very little is possible with declining enrollment and flat test scores as these schools are shut down for weeks before the test.
How did the developmental level of the youth in our project influence the PAR process, and how was the process adapted to respond to these developmental issues? Uneven social maturity of the students, particularly the boys, was noted in our team's observations of the Tubman classroom in Year 1. This factor represented a challenge in engaging the class in the PAR curriculum, although the inadequate physical space and the inexperience of the teacher also played contributing roles. The combination of these factors was primarily responded to by creating more structure in the regular class and also by creating the additional small group meetings. The latter format enabled in-depth exchange in a quiet place, with sufficient structure to develop relationships and get beyond the ''silliness'' that sometimes ensued before the students settled into a more serious discussion. Prior PAR research with slightly younger students than our sample (5th and 6th graders) also identified group dynamics and social maturity as major challenges in implementation, attributing behaviors such as clowning, put-downs, and ''silly'' responses to questions as part of the early adolescent developmental tasks of identity formation (Wilson et al. 2007). As discussed earlier, some of the challenges in engaging students in PAR activities are probably not solely a function of the developmental level of the students but also of doing PAR in institutional settings in which many students have likely not had access to developmentally-appropriate opportunities to foster students' capacity to express themselves, think critically, and work together in more mature roles.
In reflecting across the 20 middle and high school sites at which the CBO was conducting some type of PAR project, the supervisor of the Tubman project provided additional insights regarding how the process can respond to the developmental needs of the older child and younger adolescent: They [middle school students] need to look for more short-term change. It's hard enough with high school -''that this change might not happen this year'' -but in middle school if they don't see some sort of progress, they are going to lose steam very quickly. Without changing the issue every single time, [we ask]: What are super-concrete actions that all build towards this same thing?… An example from [another middle school] right now is that they did a short survey and found out that racism and stereotypes are a problem. So they are doing three standalone events that were all about racism: A lunchtime activity where students talked to those they wouldn't ordinarily talk to, a multicultural assembly, and an assembly where they put together a game about different people's experiences.
One way to build on this model, consistent with the iterative nature of PAR, would be to integrate formative evaluation research into the action phase that would enable the youth to assess if their actions actually yielded any benefits. As articulated by the supervisor, a clear advantage of more quickly engaging the students in action steps that are relevant to the problem but do not necessarily involve a change in policies or practices is that they can feel that they are making something happen. This can be problematic, however. Presentations and other events are time-consuming; while they raise awareness, they can create the sense that something is being done about the problem without any meaningful mechanisms for change being implemented (Kirshner 2007).
It is also important to recognize that adults' effective implementation of PAR requires substantial experience and support. In contrast to the typical training and skill set of classroom teachers for the instruction of content ''standards,'' facilitation of PAR requires that classroom teachers share power with students and guide them in a flexible process in which the teacher does not have the answers ahead of time and likely needs ongoing technical assistance regarding research and advocacy activities [please see Ozer et al. (2008) for an extensive consideration of technical support and capacity-building for teachers implementing PAR projects in schools and the California Center for Civic Participation and Youth Development (2004) for more general planning for youth participation.] Other research has noted the considerable challenges in effectively training facilitators to lead PAR projects, although in these cases the facilitators were university students or adult volunteers (Helitzer et al. 2000;Wilson et al. 2007).
The existing literature provides an explication of broad principles and several curricula to guide PAR in schools. To our knowledge, however, the field is lacking specific guidelines about how to assess the processes that reflect a high-quality implementation of PAR. We are mindful that PAR projects are inherently flexible and will unfold in differing ways across settings and in response to the particular questions and parameters of each project; however, a common framework may inform practice and guide the assessment of implementation quality necessary for both ''continuous improvement'' and traditional evaluation efforts of PAR interventions. Our collaborative work in supporting and evaluating PAR at multiple school sites has necessitated assessment of processes that reflect a highquality operationalization of PAR with students in classroom and school settings. Building on our initial work with Tubman and other sites, our multi-method intervention research utilizes a design that compares the process and outcomes of classrooms engaged in PAR versus classrooms engaged in a direct-service youth development program. This endeavor has further required us to delineate central PAR processes. Our specification of processes build on prior theory and curriculum development in the PAR field (Cargo et al. 2003;Checkoway et al. 2003;Jennings et al. 2006;London 2001;Schensul et al. 2004;Zimmerman 1995); they further reflect dimensions suggested to us by youth in our own research.
Processes that we view as central to PAR in our assessment system are the iterative integration of research and action, the training and practice of research skills, the teacher's sharing of power with students in the research and action process, and the practice of strategic thinking and strategies for influencing change. Examples of the specific activities we characterize as part of the strategic thinking process include discussion of: root causes to social or health problems, information about how rules or policies are made, how to develop recommendations based in research, and how to develop alliances with various stakeholders.
Processes that promote a high-quality implementation of PAR but are not unique to it include: expansion of the social network of the youth, opportunities and guidance for working in groups to achieve goals, and the development of skills to communicate with other youth and adult stakeholders. All of these processes were reported by the youth participants in our PAR classes as constituting very different experiences from their regular classroom activities, and meaningful aspects of their learning in the PAR process. Class climate has been studied extensively as a key factor in prior educational research; in our evaluation work, we are assessing classroom climate dimensions that we believe facilitate effective implementation of PAR, including the teacher's emphasis on student perspectives, the teacher's flexibility regarding classroom projects or structure, and the engagement of the students in the classroom activities (Pianta et al. 2006). Efforts to measure PAR processes through the development of a reliable and valid observational tool are currently in progress.
The Tubman project experienced mixed success. In light of the challenges faced by this 1st year teacher and her students in Year 1, having this class effectively engage in several phases of the PAR process and make meaningful presentations to their principal and to a conference of high school students were legitimately viewed as major achievements by the students and the adults involved in the project. Further, the teacher continued to implement and value the photo-voice process beyond the initial consultation and the principal remained optimistic about the potential for youth ''voice'' in the school. No specific improvements in the problems identified by the students or the school site, however, could be attributed to the project.
The Tubman project represents an initial effort that provided multiple lessons useful for the evolution of our work in school-based PAR. With respect to research, students were trained in and practiced photo-voice as a research method and learned about the strengths and limitations of other methods such as surveys and interviews as part of their conference with other students engaged in PAR in the district. With respect to practicing strategies for strategic thinking and influencing change, all students in the Tubman project analyzed the roots of issues of concern to them although the girls' group received more intensive practice in both analysis of problems and generation of possible solutions. The recommendations of the girls' group, however, would ideally have been grounded by further research. In terms of students' sharing of power, the students exercised power in their choice of topics and research methods used; they chose what to photograph and organized their own presentation. As discussed earlier, however, the teacher provided much direction and structure over the classroom activities.
Regarding general processes that support but are not unique to PAR implementation, the Tubman project provided guidance for group work and opportunities for participants to develop their skills in communicating with youth and adult stakeholders. Beyond the meeting with the principal, the first year of the project fell short in its expansion of the participants' social network and in the building of alliances with school faculty, administrators, and the student body to address the problems identified and promote the implementation of recommended changes in policies or practices. Not surprisingly, ending Year 1 without achieving clear agreements among the principal, teacher, students, and supervisor about specific action steps, a timeline, and accountability for follow-up undermined the potential impact. Finally, the importance of class climate dimensions such as student engagement for effective implementation was highlighted previously.
In retrospect, it is clear that more attention should have been paid to long-range planning and building alliances with teachers and others during the Year 1 initial implementation of the project at Tubman. Long-range planning was challenging to enact, however, for several reasons. First, with the university team in the role of consultant, the issue of planning for next year can be raised (as it was at Tubman), but the power to establish these agreements and commitments remains with the CBO and school sites, who are often in reactive rather than proactive stances when dealing with pressing funding and staffing uncertainties. Second, as evident at Tubman and at other sites where we have worked, the students' successfully engaging in the PAR process to the point of developing research-driven recommendations and events are legitimately experienced as ''wins'' for the youth and adults. Given all of the challenges at the sites, it is sometimes difficult for all of those involved in the project to see beyond the specifics of the project to think deeply about sustainable change.
Third, although it would be ideal to engage potential adult allies within and outside of the school for students' research-driven change efforts earlier in the project, to help pave the way for receptiveness in the action phases, it is our experience that it takes time before students build the skills and confidence to engage with adults in a collegial manner. That is, once students have engaged in the research and have findings to share, they tend to realize the expertise that they are bringing to the table to discuss with adults; these data presentations also help to legitimize their expertise in the eyes of the adults. Thus, preparing the ground with potential adult allies early in the process, before research recommendations are ready, is a key role for the adult facilitators and for youth who feel ready for this; we have also seen effective approaches in our projects in which data-gathering from potential adult allies is used to build relationships and buy-in. Reflections on lessons learned with Tubman and other sites have spurred the CBO and university team to emphasize alliance-building and planning for continuity in the training of teachers, curriculum, and supervision throughout the PAR process.
The kinds of potential allies that youth might want to engage will partially depend on the issue they choose to address; for example, the issue of safety and violence in specific neighborhoods might include meeting with neighborhood and ethnic associations, elected officials, local CBO's, health care providers, and the police and transportation departments. Issues focused on school equity and conditions would likely engage the site council, school board, PTA, and CBO's focused on the educational system. Progress on some issues may be aided by these external alliances; in some cases, however, the involvement of outsiders may not be helpful (e.g., involving non-school members in students' inquiry and action regarding teaching practices or hiring might increase defensiveness and limit constructive collaboration).
This discussion emphasizes the promise of PAR in middle schools, providing a case example that illustrates multiple opportunities and challenges for PAR practice in middle school. In this case, PAR appeared to be developmentallyappropriate and meaningful for students when basic classroom dynamics were addressed, and was particularly fruitful in a small-group context that facilitated the strengthening of relationships among the youth and with the adult facilitator. Infusing this program into an elective program supervised by a local CBO committed to youth development practice, and providing ongoing technical assistance to the CBO, utilized an existing niche and strengthened existing resources in the school and district. This approach helped the sustainability of PAR efforts because of the CBO's long-term relationships with the school sites. It also provided an elective ''space'' for the implementation of PAR in low-performing schools, where implementation in regular academic classes would not have been feasible due to the great demands on instructional time. Elective ''spaces,'' however, were more tenuous than expected, and uncertain funding for electives undermined longer-term planning.
In our ongoing research, we are studying in-school PAR classes over multiple years at diverse urban sites to help understand the conditions that support effective implementation of this complex intervention (Biglan et al. 2000). Of particular interest is how to strengthen the continuity of effort across semesters and years while allowing for each new cohort to ''own'' the project, and how youth can impact the climate and governance of their schools despite multiple challenges and competing demands. Our current research further emphasizes the perspectives of the youth participants (as well as adults) on PAR implementation and impact; a limitation of the Tubman project described here was its lack of student perspectives regarding impact. As alluded to earlier, we are working to develop reliable and valid measures of processes and outcomes of school-based PAR with the goal of contributing to research and practice in this growing field. PAR holds tremendous potential for providing the means by which students can initiate inquiry, develop skills, and provide recommendations to improve the developmental quality and fit of middle schools for their development as thinkers and citizens.
Place and participant names are masked in this article.
The authors acknowledge the contributions of an anonymous reviewer regarding this issue of student autonomy and how the existing expectations for students' behavior may create challenges for PAR in middle school settings.
Acknowledgments This research was supported by a
Open Access This article is distributed under the terms of the Creative Commons Attribution Noncommercial License which permits any noncommercial use, distribution, and reproduction in any medium, provided the original author(s) and source are credited.
The individualized reference for defining small for gestational age (SGA) at birth has gained popularity in recent years. However, its utility on fetal assessment has not been evaluated. The authors compare an individualized with an ultrasound reference in predicting poor perinatal outcomes. Data from a large clinical trial in predominantly white US women (1987)(1988)(1989)(1990)(1991) with singleton pregnancies (n ¼ 9,526) were used. The individualized reference classified fewer SGA fetuses than the ultrasound reference, but the risks of adverse outcomes were similar between fetuses classified by both references. The risk increased substantially only when the percentiles fell below the 5th percentile (likelihood ratio positive at birth ¼ 2.68 (95% confidence interval (CI): 2.00, 3.58) and 3.13 (95% CI: 2.34, 4.18) for ultrasound and individualized references, respectively). SGA fetuses defined by either the individualized or ultrasound reference alone had risk ratios of adverse outcomes of 1.91 (95% CI: 0.77, 4.77) and 1.18 (95% CI: 0.37, 3.77), respectively, compared with normal fetuses (the difference between these 2 risk ratios, P ¼ 0.71). The authors conclude that neither the ultrasound-based nor the individualized reference does well in predicting adverse perinatal outcomes. The 5th percentile may be a better cutpoint than the 10th percentile in defining SGA.
Defining fetal growth restriction has been a longstanding challenge. Currently, most clinicians and researchers use small for gestational age (SGA) (i.e., the smallest 10% of fetuses or newborns at a given gestational week) as a surrogate for fetal growth restriction (1). One of the primary deficiencies of the current definition of fetal growth restriction is that it is based on an absolute fetal size, irrespective of maternal and fetal genetic and physiologic factors. For example, fetuses from a large white mother and a small Asian mother are both assessed by the same criterion, which could deem one fetus to be SGA and the other to be within the normal range, even though what is normal can be expected to differ according to maternal characteristics. As a result, both under-and overdiagnosis of fetal growth restriction may occur, which has significant clinical implications.
Clinicians and researchers have increasingly recognized that fetal growth should be evaluated according to the extent to which a particular fetus has fulfilled its growth potential for a given maternal and fetal profile (2). Brenner et al. (3) were the first to advance this notion of an individualized reference, proposing an adjustment for race, sex, and parity when assessing birth weight by gestational age. Gardosi et al. (4) later developed a detailed methodology to automate the procedure that used a fetal weight curve and adjusted for maternal race, parity, pre-or early pregnancy weight and height, and infant's sex. Their methodology continues to gain greater acceptance (5).
Currently, 3 types of SGA references are commonly usedreferences based on birth weight (e.g., by Alexander et al.) (6), estimated fetal weight (e.g., by Hadlock et al.) (7), and individualized reference (4). The birth weight reference is flawed in early gestation weeks because babies born preterm are often growth restricted. That segment in the birth weight reference has artificially lower percentile limits (8); that is, SGA is likely to be underdiagnosed in early gestation. Several past studies found that infants classified by the individualized reference as SGA had a significantly higher overall mortality and morbidity than SGA infants classified by the birth weight reference (9-13). The higher perinatal mortality among infants classified as SGA by the individualized reference is largely due to the inclusion of more preterm births in this group (8,14). Given the important deficiency with the birth weight reference in preterm births, researchers are now seeking better references.
The individualized reference adopts the fetal weight reference as the base but adjusts for other factors (4). Some studies of the risk of stillbirth (15) and neonatal death (16) have suggested that the advantages of the individualized reference, relative to a simple ultrasound-based fetal weight reference (i.e., without adjustment for maternal or fetal characteristics), are rather limited. It remains unclear whether the individualized reference is superior if perinatal morbidity (instead of mortality) is used as the outcome.
Although assessment of newborn size at birth is important for pediatric care and research for long-term outcomes, recognizing SGA in utero is as crucial for obstetric care and research on fetal programming. Thus, existing research leaves open the questions about how best to identify SGA fetuses either in utero or at birth that have a higher risk of exhibiting adverse outcomes. No study to our knowledge has evaluated the prenatal utility of individualized reference. Therefore, this paper focuses on the ultrasound estimated fetal weight and compares the fetal weight and individualized references with regard to their ability to predict perinatal outcomes.
The Routine Antenatal Diagnostic Imaging with Ultrasound (RADIUS) trial was a multicenter study of pregnant women at low risk for adverse outcomes. The trial was designed to test the hypothesis that routine screening with standardized ultrasonography on 2 occasions would reduce perinatal morbidity and mortality. A detailed description of the trial is provided elsewhere (17). Briefly, from November 1987 to May 1991, pregnant women who were aged 18 years or older, who spoke English, whose last menstrual period (LMP) was known within 1 week, and whose gestational age was less than 18 weeks at recruitment were potentially eligible for the study. The study excluded women who had a previous stillbirth, prior SGA infant, irregular menstrual cycles, discrepancy between uterine size and dates of more than 3 weeks, diabetes, chronic hypertension, or chronic renal disease. In total, the RADIUS trial recruited 15,151 women, including 15,018 singleton pregnancies; 90% of the subjects were white.
Eligible women were randomly assigned to either the ultrasound screening group or the control group. Women in the former group underwent sonographic examinations at both 15-22 and 31-35 weeks of gestation. All of the study participants, regardless of group, could undergo ultrasonography at any time for medical or obstetric indications, as identified by their physicians. The standardized fetal biometry included biparietal diameter, head circumference, abdominal circumference, and femur length. Women were interviewed at recruitment, and information on demographic characteristics and reproductive history was recorded. Ante-partum, intrapartum, and neonatal information was abstracted from the antenatal medical records and from inpatient hospital records.
At enrollment, each woman reported the first day of her LMP. We used ultrasound measurements of the first fetal biometry to validate the gestational age, relying on Hadlock's formula (18) to calculate an ultrasound-based LMP. If the self-reported LMP and ultrasound-based LMP differed by more than 7 days for women prior to 21 weeks of gestation, or by more than 10 days for gestational ages between 21 and 26 weeks, the ultrasound-based LMP was substituted for the self-reported LMP. Gestational age at delivery was calculated according to the corrected LMP. Ultrasound measurements at 30 weeks or later were used to calculate the estimated fetal weight on the basis of head and abdominal circumferences and femur length (19).
We selected women who had an ultrasound examination at 30 weeks of gestation or later because that is when most ultrasound examinations in late gestation were done in this study. A total of 6,787 and 2,027 women had at least 1 ultrasound scan between 30-33 weeks of gestation and at 34 weeks or later, respectively. The sample used in our analysis included a total of 9,526 births delivered at 30 weeks or later (some of them had birth weight, but not estimated fetal weight at 30 weeks or after). All of the subjects had complete information on birth weight, gestational age, maternal race, parity, infant's sex, maternal prepregnancy weight and height, and perinatal morbidity and mortality.
To define SGA, we applied 2 widely used references to our study population. The first was an ultrasound-based fetal weight reference (7), which was created on the basis of a cross-sectional cohort of predominantly low-risk white women in the United Sttates with a normal perinatal outcome. The second was an individualized reference, which adjusted the fetal weight reference for maternal race/ ethnicity, parity, prepregnancy or early pregnancy height and weight, and infant's sex in the US population (www. gestation.net) (20). We used the 10th and 5th percentiles in each reference and compared the proportion of infants classified by these references as SGA.
In the RADIUS trial, the list of adverse perinatal outcomes encompassed the following: fetal death or neonatal death up to 28 days of age, grade IV retinopathy of prematurity, bronchopulmonary dysplasia, mechanical ventilation required for more than 48 hours, necrotizing enterocolitis, intraventricular hemorrhage, subdural or cerebral hemorrhage, neonatal seizure, placement of chest tube, neonatal sepsis, oxygen required for more than 48 hours, birth trauma, or a stay of more than 5 days in a special care nursery. Because of the small number of subjects with severe adverse outcomes, we created a composite outcome that includes any of the above conditions. We compared the incidence of the composite outcome among infants classified by both references and, for each reference, we calculated the likelihood ratio (21) in predicting the adverse outcome. Finally, because this is a secondary data analysis using deidentified data, an institutional review board exemption was obtained.
Table 1 presents the proportions of SGA infants that were based on different reference types. The ultrasound reference classified more SGA infants (<10th percentile): 8% at 30-33 weeks, 13% at 34 weeks, and 11% at birth. The numbers of SGA infants that the individualized reference identified are as follows: 6%, 10%, and 8%, respectively. Table 1 also shows that the vast majority of adverse perinatal outcomes occurred in fetuses/infants weighing above the 10th percentile of either reference. The ultrasound and individualized references yielded a very similar incidence of adverse outcomes in SGA infants (<10th percentile overall) (all P > 0.05). The results also suggest that the incidence of adverse outcome is substantially higher only when the weight falls below the 5th percentile. The likelihood ratios indicate that only weight below the 5th percentile may have predictive power for adverse outcomes. These findings are robust even when we separate preterm and term births (results not shown).
We then examined the association between weight below the 5th percentile, based on the ultrasound and individualized references, and the incidence of adverse outcomes (Table 2). The SGA cases classified by these 2 references overlap considerably (70%-77% of SGA cases). The cases identified by both references had a much higher risk of adverse outcomes than those identified by just 1 of the references and those not classified as SGA cases by either reference. When estimated fetal weight was considered, subjects classified as SGA only by the individualized reference seem to have a higher risk ratio of adverse outcomes (risk ratio ¼ 1.91, 95% confidence interval (CI): 0.77, 4.77) than those classified as SGA by the ultrasound reference alone (risk ratio ¼ 1.18, 95% CI: 0.37, 3.77). A similar pattern of results was observed for birth weight. However, the differences in corresponding risk ratios were not statistically significant (P ¼ 0.71 and P ¼ 0.23, respectively). Furthermore, because of substantial overlap, the number of extra cases that can be identified by the individualized reference is limited.
Our study, using data from the RADIUS trial, shows that the 5th percentile is a better cutoff point for defining SGA than the 10th percentile with regard to its ability to predict adverse perinatal outcomes. However, even with the 5th percentile, neither the ultrasound nor the individualized reference for SGA has high predictive power. With advanced perinatal care, the majority of infants whose weight is below the 10th percentile survive well without any significant morbidity and mortality. Our study shows that the risk of adverse perinatal outcomes increased meaningfully only after the weight falls below the 5th percentile. This finding raises the question of whether in a contemporary population SGA should be redefined as having a weight below the 5th percentile rather than the traditional 10th percentile.
It is well established that race, parity, sex, and maternal prepregnancy height and weight influence fetal weight. For example, male fetuses are, on average, 100 g heavier than female fetuses at birth (22). The mean birth weight of white infants is approximately 200 g greater than that of black infants (22). In principle, therefore, the individualized reference should be able to assign a fetus/infant to a more appropriate weight percentile than a simple ultrasoundbased reference. Fine tuning the weight percentile, however, does not seem to yield substantial gains as far as predicting adverse perinatal outcomes.
Several reasons may explain this phenomenon. First, the likelihood of being classified by both references as being below the 5th percentiles is high; that is, there is a large overlap. The fine tuning affects mostly the borderline SGA cases. These cases tend to have lower mortality and morbidity than the more severe cases.
Second, errors in the ultrasound-based estimated fetal weight may have further reduced the potential improvement. These findings are consistent with those of previous studies (15,16).
Third, the correlation between weight percentiles and risk of adverse perinatal outcomes is disappointingly low. Similar to previous studies, the current study found that the vast majority of adverse perinatal outcomes occurred in fetuses/ infants with weights above the 10th percentile (Table 1) (23). Conversely, the risk of adverse outcomes increases substantially only when the estimated fetal weight or birth weight is below the 5th percentile in both preterm and term births. Thus, any improvement in percentile assignment of-fered by the individualized reference is further offset by the low correlation between the percentile and adverse outcomes. These deficiencies may explain why the individualized reference does not provide a substantial overall advantage over the simple ultrasound reference in our study.
One could argue that the conditions included in the perinatal composite outcome are not specific to disorders of fetal growth and that the improved assignment of weight percentiles by the individualized reference may still be important for more subtle and long-term effects, such as child neurodevelopment and adult diseases in later life. Fetal growth restriction is a consequence of many causes and an indicator of compromised fetal status. Thus, fetal growth restriction itself does not necessarily cause adverse perinatal outcomes but is associated with a wide variety of perinatal mortality and morbidity. We selected a number of perinatal outcomes that are clinically important and priority concerns of both obstetricians and neonatologists (24,25). Consequently, these outcomes should be the ''gold standard'' in assessing the prenatal utility of these references, even though they may not be good indicators for long-term effects.
As in previous studies (15,16), our study population also has the limitation of being relatively homogeneous-a predominantly white population. It is reasonable to question whether the advantage of the individualized reference mainly reflects in a racially diverse population. Further studies with a diverse population are needed to address this issue. Nonetheless, findings from previous studies and ours suggest that adjusting for other characteristics (namely, maternal height and weight, parity, and sex of the infant) in the individualized reference may not improve the classification as much as previously reported when the individualized reference was compared with a birth weight reference (9)(10)(11)(12)(13).
At the same time, our study has several advantages over existing published research. For one, we used data from a large, carefully conducted prospective trial. More importantly, our study provides the first evidence using ultrasoundbased estimated fetal weight, which is more relevant to obstetric practice than the birth weight data examined in past studies.
In conclusion, neither the ultrasound nor the individualized references for SGA do well in predicting adverse perinatal outcomes in pregnancies of predominantly white women. Research in a racially diverse population is needed to demonstrate whether the individualized reference has substantial advantages over the simple ultrasound reference. Finally, the 5th percentile may be a better cutpoint than the 10th percentile in defining SGA.
Abbreviations: CI, confidence interval; LR, likelihood ratio.
Abbreviations: CI, confidence interval; RR, risk ratio.
This study was supported by the
Author affiliations: Epidemiology Branch, the Eunice Kennedy Shriver National Institute of Child Health and Human Development, National Institutes of Health, Bethesda, Maryland (Jun Zhang, Jagteshwar Grewal, Gila Neta, Mark Klebanoff); and Department of Clinical Epidemiology, Bremen Institute for Prevention Research and Social Medicine, Bremen University, Bremen, Germany (Rafael Mikolajczyk).
By the repeated use of the doubly labeled water method (DLW), this study aimed to investigate (1) the extent of changes in energy expenditure and physical activity level (PAL) in response to increased agricultural work demands, and (2) whether the seasonal work demands induce the changes in the fairly equitable division of work and similarity of energy needs between men and women observed in our previous study (Phase 1 study; Kashiwazaki et al., [1995]: Am J Clin Nutr 62: 901-910). In a rural small agropastoral community of the Bolivian Andes, we made the follow-up study (Phase 2, 14 adults; a time of high agricultural activity) of the Phase 1 study (12 adults; a time of low agricultural activity). In the Phase 2 study, both men and women showed very high PAL (mean6SD), but there was no significant difference by sex (men; 2.18 6 0.23 (age; 64 6 11 years, n 5 7), women; 2.26 6 0.25 (63 6 10 years, n 5 7)). The increase of PAL by 11% (P 5 0.023) in the Phase 2 was equally occurred in both men and women. The factorial approach underestimated PAL significantly by 15% (P < 0.05). High PAL throughout the year ranging on average 2.0 and 2.2 was attributable to everyday tasks for subsistence and domestic works undertaking over 9-11 h (men spent 2.7 h on agricultural work and 4.7 h on animal herding, whereas women spent 7.3 h almost exclusively on animal herding). The seasonal increase in PAL was statistically significant, but it was smaller than those anticipated from published reports. A flexible division of labor played an important role in the equitable energetic increase in both men and women. Am. J.
Since the doubly labeled water (DLW) method has been established as an important tool for the measurement of total energy expenditure (TEE) in man, there have been many reports detailing TEE from a wide range of population groups under free-living conditions. As reviewed by Black et al. (1996), the method has been applied in diverse circumstances ranging from premature infants to the extreme elderly, and from bed-bound patients to athletes performing at the limits of human endurance. However, these studies have been made predominantly in affluent groups, typically urban populations of industrialized societies.
In contrast, very few DLW studies have been performed in rural populations of developing countries. Currently reports are limited to those of Gambian farmers (women: Heini et al., 1991;Singh et al., 1989;men: Heini et al., 1996), Andean male and female agropastoralists in Bolivia (Kashiwazaki et al., 1995), male farmers in rural south India (Borgonha et al., 2000), Bangladeshi lactating women working for tea estates (Rosetta et al., 2005), and lactating women of 3-6 month postpartum from a Mexican subsistence community (Butte et al., 1997). Except for the latter study, a high physical activity level (PAL) ranging from 1.90 to 2.4 is found in these populations.
In general, these studies are cross-sectional and performed during a particular time of the year. Several earlier reports (Adams, 1995;Bleiberg et al., 1980;Brun et al., 1979Brun et al., , 1981;;Ferro-Luzzi et al., 1990;Lawrence et al., 1987;Panter-Brick, 1993, 1996;Schultink et al., 1993) deduce that rural populations (particularly in non-industrialized societies) are exposed to seasonal energy imbalance, evidenced by body weight changes due to naturally occurring seasonal variations in food availability and to the changes in work demands for agricultural activities during the year. Although the reported PAL values obtained from DLW on nonindustrialized rural societies are more reliable than those derived from other methods, they do not provide a comprehensive picture by themselves. When seasonal changes in physical activity have to be considered, the estimate of energy requirements becomes even more difficult. Unfortunately there is little information available on extent of change in TEE and PAL between the peak and slack seasons of agricultural activity caused by varying obligatory tasks for subsistence.
The purpose of this study was to address this issue by comparing energy expenditure and PAL at two contrasting time and work intensity periods of the year in a rural agropastoral population. In 1990, we obtained data from an Andean Bolivian community during January and February, the preharvest season when labor demands for agriculture are at a minimum (Phase 1 : Kashiwazaki et al., 1995). This work reports a return visit to the same community during the season of more intensive agricultural activity (August-September), when field preparation and potato planting increases the total number of daily tasks being undertaken, which is achieved in part by a redistribution of the work load between family members (Phase 2).
The repeated use of the DLW method in this rural community of the less developed world has allowed us to investigate firstly, whether energy expenditure intensifies markedly in response to increased agricultural activity, and secondly, whether the seasonal work demands induce the changes in the fairly equitable division of work and similarity of energy needs between men and women observed in the Phase 1 study. A timed-record of activities collected for 3-4 days during the DLW study were used as complementary information to examine these questions.
The return visit to the same community as the previous study was made in August 1997 (Phase 2), a season contrasting to that of the previous study, as it is when the ground is leveled for potato planting. This season requires greater labor demands than the preharvest season of our previous Phase 1 study (Kashiwazaki et al., 1995). We anticipate a change in several aspects of labor demands between the two periods. The additional activities of field preparation is likely to promote a change or redistribution of work schedule for each of the family members, and an understanding of the impact of this on energy expenditure is of great interest in the realm of nutritional adaptation as well as clinical nutrition in rural populations of nonindustrialized societies.
This study was made in Vilacollo (Kashiwazaki et al., 1995), an estancia consisting nine households (eleven households in the year of 1990-1991) at high altitude (about 4,000 m above sea level) near the Bolivian border with Chile. The subjects are Aymara-speaking people living in a small community with outlying pastures and small fields for potato growing. The climate in this region is most severe from June through August, the temperature at night falling to between 0 and 2108C, with daytime temperatures of 108C and humidity less than 30%. In summer, December through February, nighttime temperatures still fall to between 0 and 38C, with daytime temperatures reaching between 12 and 208C.
Sixteen adult subjects (age greater than 18) and four children and adolescents were recruited to the 1997 study, representing eight of the possible nine households. Due to out-migration of two households and some young family members, five of the adult subjects were overlapped with the previous study (Kashiwazaki et al., 1995), and other adults were the subjects newly participated in this Phase 2 study. All procedures of this study were reviewed and approved by the Ethical Committee at the University of Occupational and Environmental Health, Kitakyushu (former affiliation of HK).
The procedures of the study were the same as those of our previous study (Kashiwazaki et al., 1995). Before the doubly labeled water was administered, subjects underwent a general health check, anthropometric measurements and were asked to provide a predose urine sample for baseline isotopic measurements. The body weight of subjects in minimal clothing was measured with a digital balance after subjects voided their bladders to the nearest 0.05 kg. Typical clothing was also weighed for correction to nude weight. Measured subject's weight was used to estimate the isotope dosing requirements. At the end of the study, no significant change in their body weight was detected.
Each subject received a mixed oral dose of 2 H 2 18 O containing 0.062 g 2 H 2 O/kg and 0.149 g H 2 18 O/kg. The subject was then given a cup of water or a light refreshment drink to ensure that the entire dose was consumed. Subjects were instructed to avoid large fluid and food intakes for the next 4 h, and the first postdose urine sample was collected after a 6-h equilibration period. For 14 days thereafter, urine samples were collected daily in early or mid morning after subjects had voided overnight urine. To avoid evaporation and isotopic exchange with atmospheric air, urine sample was immediately transferred in duplicate to a small air-tight vial (10 ml) closed by a screw cap, tightly sealed with silicon tape, and kept in a dark and cool room.
After the entire field research was completed, all urine samples were sent to Cambridge (UK) for stable isotope analysis.
The postabsorptive resting metabolic rate was determined in duplicate and on two successive days in the early morning, according to standard procedures for measurement of basal metabolic rate (BMR). For simplicity and comparative purpose, we expediently defined and used these data as BMR in this article. The day before the BMR measurement, each subject assigned to the DLW study was asked to accept a visit by one investigator to his or her house in the early morning, at about 06:00, usually before the subject woke up. While subjects were remaining in their own bed at complete rest and after an overnight fast, a facemask was attached to them to collect expired air. The procedure of measurement was the same as that described previously (Kashiwazaki et al., 1986(Kashiwazaki et al., , 1990)). No control was made for room temperature, which ranged from 4 to 88C (outdoor temperature ranged from 25 to 28C). During the measurement, subjects were laid on their bed and covered with sufficient clothing and blankets to provide a comfortable temperature of 25-288C over most of their body. Only the subject's face was exposed to room temperature, which may have affected the measurement of BMR, producing a slightly higher value than that measured in controlled laboratory conditions (Kashiwazaki et al., 1986(Kashiwazaki et al., , 1990)). These conditions were the same as those with the Phase 1 study (Kashiwazaki et al., 1995), providing the comparable BMR data. After 5 min of stabilization, expired air was collected with a Douglas bag, twice, for 5 min. Tests for leaks were made at each measurement. The volume of expired air was measured with the certified dry gas meter (Shinagawa Seisakusho, Tokyo), and the sample of expired air was analyzed for O 2 and CO 2 concentrations in duplicate with a Roken-shiki portable gas analyzer (Shimada, Tokyo; a modified Haldane gas analysis apparatus that can be used without a supply of electric power). BMR was calculated by using Weir's equation (Weir, 1949). The accuracy and precision of the apparatus were carefully checked in Tokyo before its application in the field. By simultaneous measurements of expired air with a paramagnetic oxygen analyzer and infrared carbon dioxide analyzer (Expired Gas Monitor, model 1H21; San-Ei Sokki, Tokyo), accuracy and precision was proved to be acceptable; differences in the measurements between the two apparatuses were negligible: the mean difference and 95% confidence limits of agreement for nine measurements of standard air were 0.08 6 0.14% in oxygen and 20.12 6 0.22% in CO 2 . Precision as CV was 0.4% for oxygen and 1.8% for CO 2 on repeated measurements of expired gas. The measured BMR was highly correlated (r 5 0.89, P < 0.001) with the BMR estimated from body weight by using the equations of the FAO/WHO/UNU (1985), and the mean differences on average of about 1% between the measured BMR and the estimated BMR was not statistically significant by paired t-test. In the Phase 1 study, measured BMR was 5.6% lower than estimated BMR (Kashiwazaki et al., 1995). This means that the calculated PAL (TEE/BMR) by using the measured BMR would include a potential underestimation by 4%-5% in the subjects of Phase 2 study.
Measurement of isotopes and calculation of carbon dioxide production 2 H and 18 O composition of the urine samples, and samples of dilute doses were measured using a Sira 10 dual inlet mass spectrometer (Micromass, Wythenshawe, UK). For 2 H, a 0.4 ml of urine was aliquoted into a disposable 3.5 ml glass vial containing a reusable platinum catalyst (Finnigan MAT, Bremen, Germany) and fitted with a rubber septum. Isotopic equilibration with hydrogen gas at 2 bars and 258C and was complete after 3 h. The hydrogen was then admitted to the mass spectrometer for isotopic determination. Measurements were made against a sample of H 2 gas similarly equilibrated with water of natural abundance and were corrected for interference of H 31 . Laboratory standards calibrated with values of 251.45% and 763.24% relative to standard mean ocean water (SMOW)/standard light antarctic precipitation (SLAP) (147.75 and 274.64 ppm) were run as unknowns and the true enrichment of the analyzed samples calculated from these. Precision of the measurements was 1.58% (0.24 ppm). After use, the catalyst rods were washed with deionized natural abundance water and stored in air at 408C for at least a day before reuse.
For 18 O enrichments, 3 ml aliquots of urine were equilibrated with 13 ml CO 2 at 400 mbars and 25 6 0.18C for 6 h on a shaker bench (Isoprep system Micromass, Wythenshawe, UK). After admission to the mass spectrometer, the CO 2 was measured relative to a cylinder of gas traceable to international standards, and the isotopic composition expressed relative to SMOW. The precision of these measurements was 0.15% (0.3 ppm).
Dilution spaces for 2 H (Nd) and 18 O (No) and disappearance rate constants for 2 H (Kd) and 18 O (Ko) were calculated by the multipoint method (Coward, 1988). From these, calculation of CO 2 production was made as previously (Kashiwazaki et al., 1995).
The respiratory quotient (RQ) required for the calculation of energy expenditure from CO 2 production was assumed to be equivalent to the food quotient (FQ), which was taken as 0.93 as the mean estimate from our previous study (Kashiwazaki et al., 1995). Averaged 24 h total energy expenditure (TEE) was calculated from the mean daily CO 2 production by using Weir's formula (Weir, 1949): kJ=LCO 2 ¼ 4:63 þ 16:49=RQ Propagation of error analysis was performed on each measurement to obtain an estimate of individual errors, which is the product of the standard errors of the flux rates (Ko and Kd), and the pool sizes (No and Nd) resulting in a standard error for the CO 2 production uncorrected for fractionation (Cole et al., 1990). The errors computed by this procedure, contain analytical error and physiological variation from day-to-day changes in flux rates of 18 O and 2 H.
Total body water (TBW) determined by DLW was calculated from 18 O and 2 H dilution spaces combined. Fat-free mass (FFM) was estimated by dividing TBW by 0.732. Body fat was calculated by subtracting FFM from body weight.
Time-allocation study was conducted contemporaneously with the DLW study. A subset of six households composing of 15 subjects (12 adult subjects) was selected (four subjects from two households were not observed because of the remoteness of their houses from the majority of dwellings). Two well-trained assistants (both of them native speakers of Aymara, and one native to the area) observed the assigned subjects over 3-4 days. One observer was able to record usually two-three subjects of one household. The activities of each individual were recorded, and the times when changes in major activity occurred were noted. Records were kept from 6:30 am to 7:00-8:30 pm (about 12-14 h per day). Where there were missing periods or more detail required, the observations were augmented by ad hoc interviewing of the subject.
Recorded activities were classified into the following main categories and subcategories, I: subsistence (IA: daily herding and animal care, IB: farming work), II: household chores (IIA: cooking and food processing, IIB: washing and sweeping), III: eating and discretionary activities (IIIA: eating, IIIB: rest, chatting, and reading), IV: social relation (mainly time spent for attending the meeting with neighboring communities), and V: overnight sleep (when bed-time and wake-up time were missing, we assumed the rest of time was spent in overnight sleep). For each of the activity subcategories, the time weighted energy cost (EC tw ) was calculated by using the energy costs (PAR: physical activity ratio), adopted from the published list of FAO (2004).
All data are expressed as mean 6 standard deviation (SD), unless otherwise noted. Statistical analyses were performed by using SPSS 11J statistical software for Windows (SPSS Japan, Tokyo, Japan). Correlation analysis was used to assess the association between energy expenditure and other variables. Differences between measurements were analyzed by unpaired Student's t-test (pooled-or separate-variance estimates after homogeneity of variance test). To test the effect of factors (sex, phase; year, and interactions), least-squares means were also computed and adjusted for weight as a covariate by using the general linear models (GLM) procedure in group comparisons of the measurements on adult subjects. When the probabilities were <0.05, the statistical tests were regarded as significant.
Table 1 shows individual data on physical characteristics, DLW variables, total energy expenditure (TEE), BMR, and PAL. Data on children and adolescents are also presented for reference. The following analyses are limited to those for adults. Although only 7 years elapsed between the two studies, the mean age of the adult subjects (male; 64 6 11 years, female; 63 6 10 years) resulted in the increase by about 25 years. This is partly due to the aging of overlapped subjects with the previous study, but mostly due to new subjects selected from the remained elderly people resulted from out-migration of the younger members. Medical checks made by one of the authors (JOR) before giving stable isotope found no severe health problems in any of the subjects, however, many of them complained of poor vision, and shoulder-or head-ache usually associated with fatigue.
An overview of the individual data for the adult subjects is shown in Figure 1; in which physical activity level (PAL: TEE/BMR) values are plotted in comparison with those from the Phase 1 study. In neither season was there a sex-difference in the PAL of adult subjects. In both male and female subjects, medians (indicated by gray bars in the box of the figure) shifted toward higher PAL in the season of Phase 2 (1997). As shown, two female subjects had PAL 1.5 extremely lower than other subjects. Sub-ject F12 in Phase 1 (1990) was pregnant, and F4 in Phase 2 (1997) had been recently hospitalized and undergone surgery. She belonged to a household, which did not hold herds of animals, and her major physical activity was household chores. Her husband (M2) was working as a teacher at the elementary school of this area, and their lifestyle and physical activity pattern differed from most of the other adults in this rural estancia. Similar low PALs were found among Yakut men and women transitioning away from a subsistence herding lifestyle (Snodgrass et al., 2006). In the following analyses, subjects of M2 and F4 in Phase 2 (1997), and F12 in Phase 1 (1990) were excluded.
Body mass index (BMI), estimated FFM, and fat% suggest that the subjects have a normal body composition (Table 2). They appeared to have no tendency to an increased body fat with age as has been often observed in many societies of the developed countries (no significant correlations were observed between age and body composition; data not shown). Despite the advanced and older age of the subjects, when compared with the Phase 1 study, no statistical differences were detected in BMI, FFM, or fat% between Phase 1 and Phase 2.
Energy expenditure and physical activity indicators of the adult subjects were examined to test for the differences by both gender and phases (Table 3). Two procedures were applied. One was by comparing the least-squares means of TEE and activity energy expenditure (AEE, derived as TEE2BMR) with body weight as the covariate, as most activities are presumed to be weight bearing. The second was comparing TEE and AEE normalized by weight. Neither method indicated any gender differences, only differences between the two phases, where the extent Subjects participated in the previous study. e After a few days of giving DLW, F1 had repeated attacks of diarrhea and was inactive during the DLW period. This was reflected in the relatively high water turnover rate as seen in Kd and Ko.
of increase in the labor intensive season was almost identical for male and female adults, e.g., 11% increase in PAL and about 25% in AEE adjusted for weight.
Timed-activity data on 12 adults (six males and six females) allowed us to examine the physical activity patterns in detail, expressed as time spent and estimated time-weighted energy cost (EC tw ) (Table 4). Both male and female subjects spent most of their time undertaking subsistence activities (male; 437 6 70 min, female; 444 6 92 min), which was about 30% (male; 30% vs. female; 31%) of time during 24 h. Of a total time spent on subsistence activity, men spent 2.7 h/24 h on agricultural work and 4.7 h/24 h on animal herding, whereas women spent 7.3 h/24 h almost exclusively on animal herding. In five major categories, although time spent in subsistence and sleep showed no gender differences, three activity categories (household chores, eating and other discretionary activities, and activity for social relations) had gender differences in allocated time. Gender differences in activity patterns appeared more obvious in subcategories. Women spent most of their time in daily herding and animal care, and rarely in other farming work. Compared with men, women spent more than twice the time undertaking household chores, less than half the time for light leisure time activities (chatting, rest, and reading), and little time in activities involving in the community meeting. When compared on the basis of EC tw , the results with respect to gender difference were the same as those times allocated. More than 50% (male; 56% vs. female; 52%) of average daily energy expenditure was spent on subsistence activities. For other everyday activities combined (activities II, III, and IV), EC tw were 20% of TEE in male and 24% in female (sex difference was not detected). Sum of EC tw did not differ between males and females. When EC tw was transformed into estimated PAL est (PAL est 5 EC tw /24 h) derived from the timed-activity record and compared with the PAL derived from DLW, the values of PAL est were significantly lower by 15% in both males (1.92 6 0.12 vs. 2.21 6 0.23, n 5 6, P 5 0.03 by paired t-test) and females (1.90 6 0.09 vs. 2.24 6 0.27, n 5 6, P 5 0.02 by paired t-test).
High PAL of rural Andean Aymara has two notable features. Firstly, there was no decreasing trend in PAL with aging and senior adults over 60 years had PAL > 2. Secondly, the increase of work demands for agricultural activities resulted in an 11% increase in PAL of both men and women compared with the Phase 1 study (Kashiwazaki et al., 1995). This increase in PAL, or the extent of seasonal fluctuation, was smaller than those anticipated from published reports on agricultural communities of less developed world (Bleiberg et al., 1980;Brun et al., 1981;Heini et al., 1996;Lawrence and Whitehead, 1988;Singh et al., 1989). Particularly, important issues from this study are the high level of energy expenditure throughout the year, and the contrast between rural and urban lifestyle relevant to; (1) PAL with aging, (2) seasonality of physical activity and obligatory work demands, (3) outmigration of younger generation to cities, and (4) the importance of qualitative information on activity patterns, which the DLW study alone cannot provide (Durnin, 1996;Irwin et al., 2001).
There are a limited number of reports on TEE using DLW for healthy senior adults over 60 years old. For example, the DLW studies of 574 measurements compiled and summarized by Black et al. (1996) from the subjects of industrialized urban settings of affluent societies suggested a range of PAL between 1.2 and 2.5 for sustainable lifestyles, in which only six reports were for senior adults. Four reports targeted for healthy senior subjects living independently; had no participant aged over 80 years, and the PAL values ranged from 1.4 to 1.8. Higher PAL values were reported in subjects with a high level of leisure activity compared with other subjects of similar age, but a PAL of greater than 2.0 was rare even in the very active subjects. Physical activity patterns also differ from that of senior age groups in the urban industrialized world. In studies of European elderly subjects, most of time was spent laying and seated (70-80%), about 20% undertaking standing activities of light to moderate intensity, and less than 20% of time spent walking or performing recreational activities (Morio et al., 1997;Visser et al., 1995). Generally, the distribution of the PAL for the elderly is lower when compared with those younger than 65 years of age. The high levels of PAL in the rural agropastoral Aymara were mainly attributable to the every day tasks and subsistence activities necessary to maintain their domestic works and animal care etc, whereas high levels of PAL in senior citizens of the Western European society is limited to those who maintain a high level of leisure time activity.
Seasonal fluctuations of work demands and food availability have been major issues when related to health and nutrition in rural communities of the less developed world. The variations in body weight, energy balance, and energy expenditure can result in long term weight cycling or the risk of shortage in energy intake, and this has been reported as having an adverse effect on health (Adams, 1995;Bleiberg et al., 1980;Brun et al., 1979Brun et al., , 1981;;Ferro-Luzzi et al., 1990;Singh et al., 1989). These studies are in groups of subjects who exclusively rely on agriculture. In subjects of this study, no apparent nutrition-related health problems were observed. As indicated in Table 2, there was no significant difference in body weight, BMI, and fat% between the subjects of two phases. High PAL values similar to our subjects have been reported only from DLW studies in Gambian women (mean PAL 5 1.97; Singh et al., 1989) or Gambian males (mean PAL 5 2.4; Heini et al., 1996). These reports are based on research during the peak agricultural season. No report exists illustrating the extent of variation in PAL of rural subjects evaluated by DLW during the course of the year. In this respect, energy expenditure data measured on the Andean agropastoral Aymara subjects are unique in their high PAL throughout the year ranging on average between 2.0 and 2.2. With respect to the constancy of high PAL not only in middle-aged but also in the elderly men and women, the work demands for their subsistence activities and out-migration of younger generation to cities could have been the interrelated compounding factors.
Relative importance of animal herding and agriculture for their subsistence Animal care requires one or two of the family members to lead them from the corral to the appropriate grassplots at the extensive foothills of local mountains or pampas. At least one of the family members should watch and move the herds to other grassplots during the daytime, and then lead them back to the corral before sunset. Herding activity involves time-consuming tasks to watch grazing animals, and it is unwise to curtail time for this activity for two reasons. By watching animals and leading them to the appropriate grassplots, the herders prevent their animals from (1) overgrazing the patchy grassplots and (2) underfeeding and malnutrition of animals. Alpacas, llamas, and sheep are their primary wealth and source for their subsistence, providing them wools for textiles, exchanging outside foods, cash income, and dung fertilizer as well as supplying them food as meat. A sheep is gener- ally slaughtered every 1-2 month for family consumption, or one llama/alpaca is consumed every 3-6 months. Shearing for wool from one animal happens every 2 years. A caravan of llamas is sometimes used to transport goods to an open-air market about 150-200 km away from this region. All family members know the necessity of providing animals with some level of constant care throughout the year, which is crucial for their survival under the harsh environmental conditions. Seasonal fluctuation in herding activity is less obvious than that of agricultural activity. A similar observation has been reported from the study in an agropastoral community in the foothills of Himalayas by Panter-Brick (1996), who concluded that seasonality as evaluated by time consumed for activity was highly significant for agriculture, forest work, and travel, but it was not detected for animal husbandry. In contrast to animal herding, the relative importance of agriculture for their subsistence is much less than other rural community. Small plots, where soil and climatic conditions permit, are cultivated in expectation of a good harvest. The preparation of fields for planting potatoes and other feasible crops in this area is one of the few tasks that truly require the strength of adult males. However, heavy reliance on agriculture is perilous to people of this area because of the risk of damage to plants by unexpected frost and drought during the growing season. The harvest from the small plots of potato field is variable and generally not sufficient enough to support the annual food consumption in a household. It would cover only about 30-60% of needs. Kim et al. (1991) reported, based on a seven-day food consumption survey in September 1988 in this area, that locally produced foodstuffs such as potato, meat, and eggs provided 37% of total energy intake and that the foodstuffs produced outside the community were of critical importance for their food and nutrition. They relied heavily on animal herding for sources of wool and meat to bring cash income or would exchange them for potatoes, cereals, and other foods in the local market. No apparent changes in their lifestyle and the importance of outside foods were observed since 1988 or 1990.
Out-migration of the younger generation of family members to attractive big cities, such as La Paz and El Alto, is another important factor leading both senior adult men and women to very high PAL. In a period of 10 years, since the household census survey of 1988, the population size of this community has decreased to 34% of 1988's total population. This decrease was exclusively due to the outmigration of the younger generation between 10 and 40 years at the time in 1988. The decrease of household members was substantial from 4.8/household to 2.9/household. Their subsistence and household activities, which had previously been shared among three-four household members, has become solely the tasks of the remaining household members; typically a husband and wife aged 50-80. If the additional work demands were only shortterm and urgent, the logical solution would have been to recruit manpower from outside the household. However, all households, being virtually in the same situation with a shortage of manpower, have resulted in the only solution of existing members spending more time on subsistence and household chores. Thus, out-migration of the younger generation combined has been the compounding factors that underlie the senior adults having high physical activity levels. This may not be an aspect unique to this area, but also may be occurring in other rural communities in developing societies. The long-term consequences of outmigration and community aging in rural areas of developing societies, deserves much more attention from social and ecological perspectives of health, physical activity, and nutritional status.
In many rural third-world communities, women contribute substantial time and energy to subsistence farming, even though they are childbearing (Panter-Brick, 1989). A study based on the observational timed-activity records, from rural areas of the Ivory Coast, reported that women consistently work more hours than men and spend less time on leisure activities (Levine et al., 2001). Identifying the use of time during 24-h, similar gender differences in time spent on work were observed in our Aymara subjects. Women spent 2 h more doing household chores and about 1 h less time on light leisure activities than men spent. The total time spent working on subsistence activities and household tasks was longer among women than men by 2.2 h (women; 657 min, men; 524 min). Observed differences in time allocated by men and women are one of the cross-sectional scenes during the year. As time needed for agricultural activities increases, men have to curtail their time spent on animal care and herding or other activities. Both husband and wife equally adjust and share their time to meet the extra work demands and because activities such as animal herding and care are crucial for their short-term survival, these tasks are unlikely to be curtailed to compensate. A flexible division of labor between husband and wife enabled the husband to engage in agricultural activities and to curtail his time for animal care. This results in an increased workload for both men and women; while men engaged in agricultural activities, women spent longer time in animal care and herding, hence compensating the time curtailed by the men. This brought almost equal increase of PAL (11%) and AEE (25%) derived from DLW data both in men and women. A similar pattern was observed in EC tw estimated from the timed-activity records, in which the activities undertaking over many hours showed no significant energetic differences by sex. In rural areas of no electricity like this community, the total time available for subsistence working is governed by daylight hours, usually about 11 h in this season (sunrise at about 7:00 am and sunset at about 18:30). The time spent on subsistence activities and household tasks was 11 h in women and 8.7 h in men, suggesting very little margin of time available for additional work, particularly in women. The observed PAL during this period appears to represent their peak for the year.
Timed-activity records provide information not only on life-style and activity pattern, which DLW study alone cannot, but it is also used as the essential data to estimate energy expenditure in the factorial method. Several calculations, using the factorial approach, reported there was a systematic underestimation of PAL, which was 15%, both in men and women when compared with the DLW method. Similar or much greater levels of underestimation was also reported in studies validated by the DLW method (Haggarty et al., 1994;Roberts et al., 1991) and studies with heart rate monitoring (Leonard et al., 1995(Leonard et al., , 1997;;Spurr et al., 1996). However, other recent DLW studies have reported that the factorial method provides a close estimate of TEE in groups of free-living adolescents (Bratteby et al., 1997) and elderly subjects (Morio et al., 1997), or in some cases, an overestimation of about 8% in males of 27-65 years (Conway et al., 2002;Irwin et al., 2001). Leonard et al. (1997) have pointed out, by reviewing studies on the factorial approach validated by DLW and HR methods, that underestimation with the factorial approach is more substantial at moderate to high activity levels than at very low activity levels. These and our report indicate that the factorial approach is prone to underestimate TEE. The reasons for systematic error remain to be examined further, but it must be emphasized that the observed underestimation in the factorial approach does not diminish the overall importance of timed-activity data.
a b (weight2FFM) 3 100/weight.
c
F4 and M2 were husband and wife not engaging in the agropastoral activities, without the domestic animals. M2 was working as a teacher at the elementary school of this area. The major physical activity of F4 was household chores. Their physical activity pattern differed from most of the other adults in this rural estancia. Eight months before, F4 was hospitalized for a surgical operation and treatment of injury suffered from an unfortunate accident of stone hitting on the head. d
BMI, body mass index expressed as weight/(height (m)) 2 ; FFM, fat free mass derived from dilution spaces of 18 O and 2 H. (for details, see Text and Table 1); Fat (%) 5 (Weight2FFM) 3 100 / Weight. a Statistical significance of the factor effects tested by GLM procedure. b BMR adjusted for FFM as covariate and the data are adjusted mean 6 SE. ns, not significant; P, statistically significant with the probability of error.
c PAR, Physical activity ratio (energy expenditure of the activity/BMR), adopted from Tables5-1FAO 2004. d
Denotes t-test by separate-variance estimate, after the F value used to test homogeneity of variance and its probability.
The authors are grateful to
Background-Left ventricular hypertrophy (LVH) is common in patients with end-stage renal disease (ESRD) and an independent risk factor for premature cardiovascular death. Left atrial volume (LAV), measured using echocardiography, predicts death in patients with ESRD. Cardiovascular magnetic resonance (CMR) imaging is a volume-independent method of accurately assessing cardiac structure and function in patients with ESRD.
Study Design-Single-center prospective observational study to assess the determinants of allcause mortality, particularly LAV, in a cohort of ESRD patients with LVH, defined using CMR imaging.
Setting & Participants-201 consecutive ESRD patients with LVH (72.1% men; mean age, 51.6 ± 11.7 years) who had undergone pretransplant cardiovascular assessment were identified using CMR imaging between 2002-2008. LVH was defined as left ventricular mass index >84.1 g/m 2 (men) or >74.6 g/m 2 (women) based on published normal left ventricle dimensions for CMR imaging. Maximal LAV was calculated using the biplane area-length method at the end of left ventricle systole and corrected for body surface area.
Results-54 patients died (11 after transplant) during a median follow-up of 3.62 years. Median LAV was 30.4 mL/m 2 (interquartile range, 26.2-58.1). Patients were grouped into high (median or higher) or low (less than median) LAV. There were no significant differences in heart rate and mitral valve Doppler early to late atrial peak velocity ratio. Increased LAV was associated with higher © 2010 Elsevier Inc. This document may be redistributed and reused, subject to certain conditions.
mortality. Kaplan-Meier survival analysis showed poorer survival in patients with higher LAV (log rank P = 0.01). High LAV and left ventricular systolic dysfunction conferred similar risk and were independent predictors of death using multivariate analysis.
Limitations-Only patients undergoing pretransplant cardiac assessment are included. Limited assessment of left ventricular diastolic function.
Conclusions-Higher LAV and left ventricular systolic dysfunction are independent predictors of death in ESRD patients with LVH.
Patients with end-stage renal disease (ESRD) have an increased risk of premature cardiovascular disease. Echocardiography has identified abnormalities in left ventricular structure and function-left ventricular hypertrophy (LVH), left ventricular systolic dysfunction (LVSD), and left ventricular dilation-that independently confer a poorer prognosis. These changes are common and have been termed "uremic cardiomyopathy." LVH is present in approximately 67% of patients with ESRD and is the most common manifestation of uremic cardiomyopathy. Moreover, it is an independent risk factor for sudden cardiac death, heart failure, and cardiac arrhythmias in both the general population and patients receiving hemodialysis. Although common, the presence of LVH alone has a variable prognosis. Furthermore, reversal of LVH in patients with ESRD has proven difficult, and attempts have been made to identify additional abnormalities that predict death and are amenable to intervention.
Previous studies measuring left ventricular mass index (LVMi: defined as left ventricular mass corrected for body surface area [BSA]) in patients with ESRD have used echocardiography. However, estimation of LVMi is inaccurate because of changes in intravascular volumes during the inter-and intra-dialytic period and during dialysis and geometric assumptions that rely on intraventricular diameter to calculate LVMi using conventional echocardiography. Cardiovascular magnetic resonance (CMR) imaging provides detailed volume-independent measurement of cardiac structure and is considered the most accurate method for assessing ventricular dimensions in patients, including those with stage 5 chronic kidney disease.
Left atrial dilation (corrected for BSA or height) measured using echocardiography is an independent predictor of mortality in the general population, and in patients with hypertension and ESRD. Causes of increased left atrial volume (LAV) in patients with ESRD include mitral valve disease, fluid overload, and impaired left ventricular diastolic relaxation and filling. LAV can be reliably and reproducibly measured using echocardiography and CMR imaging using the biplane area-length method. To this end, we postulated that increased LAV conferred poorer prognosis in ESRD patients with LVH. The aim of this study was to identify the prognostic effect of cardiac abnormalities, particularly increased LAV, in a cohort of ESRD patients with LVH identified using CMR imaging.
This was a single-center prospective observational study to assess the determinants of all-cause mortality, particularly LAV, in a cohort of patients with ESRD with LVH defined using CMR.
The Renal Transplant Unit at the Western Infirmary, Glasgow, provides transplant services to a population of 2.8 million people in the West of Scotland. The transplant waiting list has 300 patients at any time; approximately 120 new patients are waitlisted and approximately 80 adult transplants are performed annually.
Since 2002, we have used CMR imaging as part of the standard assessment of patients who are referred for cardiovascular assessment before their inclusion on the kidney transplant waiting list. These patients were referred for pretransplant assessment because of advancing age or past/current history of cardiovascular disease.
All patients receiving maintenance hemodialysis therapy from our unit were studied on a nondialysis day with the aim to perform all investigations at the individuals' "dry weight." Only patients with evidence of LVH on CMR imaging were entered into the study. To ensure that only nonvalvular causes of left atrial dilation were assessed, patients with mild to severe mitral valve disease on echocardiography, based on American Society of Echocardiography guidelines, were excluded from the study. In addition, all patients were in sinus cardiac rhythm at the time of scanning.
All-cause mortality was the principle outcome in this study. To characterize factors associated with death, demographic information, past clinical history (at time of CMR scanning), and CMR measurements were recorded.
Non-gadolinium-enhanced CMR imaging was performed using a 1.5-Tesla magnetic resonance imaging scanner (Sonata; Siemens Medical, www.medical.siemens.com) to assess LVM and function. Scans were performed on the day after their hemodialysis session. A fast imaging with steady-state precession (true FISP) sequence was used to acquire cine images in long-axis planes (vertical long axis, horizontal long axis, and left ventricular outflow tract) followed by sequential short-axis left ventricular cine loops (8-mm slice thickness, 2-mm gap between slices) from the atrioventricular ring to the apex. Imaging parameters, which were standardized for all participants, included the following values: repetition time, 3.14 ms; echo time, 1.6 ms; flip angle, 60°; voxel size, 2.2 × 1.3 × 8.0 mm; and field of view, 340 mm.
LVM and LAV were analyzed by 2 observers (R.K.P. and A.G.M.J.) blinded to patient clinical characteristics. LVM was measured from short-axis cine loops using manual tracing of epicardial and endocardial end-systolic and end-diastolic contours, calculated using analysis software (Argus; Siemens Medical), and indexed for BSA (thus providing LVMi). According to established normal values, LVH was defined as LVMi >84.1 g/m 2 (male) or >76.4 g/m 2 (female). LVSD was defined as left ventricular ejection fraction <55%, and left ventricular dilation was defined as end-diastolic volume (EDV)/BSA >111.7 mL/m 2 (male) or >99.3 mL/ m 2 (female) or end-systolic volume (ESV)/BSA >92.8 mL (male) or >70.3 mL (female).
The biplane area-length method for ellipsoid bodies was used to measure LAV. Horizontal and vertical long-axis cine images were used to obtain images of the left atrium at maximal filling. Atrial lengths and areas were measured from both views, and LAV was calculated. LAV was corrected for BSA (LAV/BSA). Left atrial appendages were included in these measurements.
Echocardiography was performed by an experienced echocardiographer (A.F.C.) using an Acuson Sequoia C512 machine (Siemens Medical). Diastolic function was assessed using pulsed-wave Doppler from apical 4-chamber views to measure the ratio of early (E) to late (A) mitral inflow peak flow velocity (E:A ratio).
Data are described as mean ± standard deviation for normally distributed data) or median and interquartile range (IQR) for non-normal data. Survival data including survival time (mean ± SD) are shown as Kaplan-Meier graphs (with statistical comparison using log-rank test). These data were also analyzed using Cox multivariate survival analysis to assess the influence of multiple clinical and cardiac variables on outcome. Variables identified as significantly influential on outcome by univariate analysis were entered into a backward stepwise regression model. All analyses were performed using SPSS, version 15.0 (SPSS Inc, www.spss.com).
From 312 patients with ESRD assessed for kidney transplantation between 2002 and 2008, we identified 201 patients with LVH. Median follow-up was 3.62 years (IQR, 1.2-5.2) and transplant-censored follow-up was 1.69 years (IQR, 1.0-3.9). Mean age of patients was 51.6 ± 11.8 years and 72.1% were men. The first column in Table 1 lists renal replacement therapy mode, medical history, and cardiac drug history of the cohort.
Seventy-one patients received a kidney transplant during the study period. There were 54 (26.9%) deaths during a 6.65-year follow-up period. Eleven deaths occurred after kidney transplantation. Overall survival at 12, 24, 36, 48, and 60 months was 92%, 87%, 79%, 77%, and 74%, respectively. Median LAV/BSA was 30.4 mL/m 2 . The distribution of LAV/BSA for this cohort is shown in Fig 1. There was no significant correlation between LVMi and LAV/ BSA (r = 0.03; P = 0.7).
To identify determinants and consequences of increased LAV, we divided patients into high LAV (LAV/BSA equal to or higher than median; n = 100) or low LAV (LAV/BSA less than median value; n = 101; Table 1). High LAV was significantly associated with treatment with statins. Low LAV was significantly associated with male sex. There was no significant difference in age, number of patients who underwent kidney transplantation, BSA, and renal replacement type or duration between the high-and low-LAV groups. Furthermore, there was no significant difference in number of patients with diabetes mellitus, smoking history, and cardiovascular medical history (namely ischemic heart disease, chronic heart failure, cerebrovascular and peripheral vascular diseases). On comparison of cardiac medications, there were no other significant differences between the low-and high-LAV groups.
Examining the entire cohort (Table 2), mean heart rate during CMR imaging was 77 ± 26 beats/ min. Mean ejection fraction was 63.1% ± 14.4%, LVMi was 117.3 ± 31.1 g/m 2 , EDV/BSA was 86.3 ± 31.4 mL/m 2 , and ESV/BSA was 34.1 ± 25.3 mL/m 2 . Fifty (24.9%) patients had LVSD and 49 (24.4%) had left ventricular dilation. Doppler mitral valve inflow velocity measurement showed a mean peak E wave velocity of 0.74 ± 0.2 cm/s, mean peak A wave of 0.75 ± 0.2 cm/s, and E:A ratio of 1.04 ± 0.5.
We also compared structural CMR imaging and mitral valve echocardiography Doppler results between the low-and high-LAV groups. There were no significant differences between heart rate during CMR imaging, ejection fraction, left ventricular myocardial mass, EDV/BSA, ESV/ BSA, and number of patients with LVSD or left ventricular dilation between groups. As a basic marker of diastolic function, Doppler mitral valve inflow velocity measurement showed no difference in E:A ratios between the low-and high-LAV groups.
We initially examined the effect of left atrial and ventricular abnormalities on patient survival. Divided into quartiles, increasing LAV was significantly associated with higher mortality: quartile (Q)1, 5 (10.0%) deaths; Q2, 13 (26.0%) deaths; Q3, 17 (34.0%) deaths; and Q4, 19 (37.3%) deaths; P = 0.01. Furthermore, increasing LAV was associated significantly with poorer prognosis (Fig 2A: P = 0.01).
Similarly, LVSD was associated with a significant reduction in mean survival time (Fig 2B: P = 0.02). Left ventricular dilation was associated with a non-significant decrease in patient survival (no left ventricular dilation, 5.2 ± 1.9 years vs left ventricular dilation, 4.7 ± 3.7 years; P = 0.2).
Table 3 lists univariate and multivariate Cox survival analyses for patient clinical and cardiac characteristics. Univariate analyses showed that LVSD, LAV/BSA, and history of ischemic heart disease were significantly associated with death. Kidney transplantation was associated with a significant survival benefit. Advancing age increased the risk of death, but this did not reach statistical significance. Multivariate analysis (Table 3) was performed and showed LVSD, LAV/BSA, and history of ischemic heart disease as independent predictors of mortality. Kidney transplantation was independently associated with decreased mortality.
Patients with ESRD have an increased risk of premature cardiovascular mortality. In contrast to the general population, sudden, presumed arrhythmic, cardiac death, rather than myocardial infarction or heart failure, is the most common cause of death. Modification of traditional risk factors (eg, dyslipidemia) does not alter prognosis significantly, and reversal of LVH (the most common abnormality of uremic cardiomyopathy) is difficult. Alternative, potentially reversible, myocardial abnormalities have been sought to provide a target for intervention that may decrease cardiovascular death in this patient population.
Against this background, we prospectively assessed the effect of additional myocardial abnormalities and clinical history on survival in a cohort of ESRD patients with LVH. In particular, we investigated the prognostic effect of LAV, which has been shown previously using echocardiography as an independent predictor of death in dialysis patients. LAV can be calculated reliably from 2-dimensional echocardiography and CMR measurements. We restricted this study to patients with LVH to identify other potentially modifiable characteristics that predict death in patients with ESRD and established uremic cardiomyopathy.
Elevated LAV (higher than the median) was less common in male patients. This differs from previous studies that have shown removal of sex-related differences in LAV when corrected for body size. Neither sex nor BSA had an effect on survival in our analyses. There was no significant difference in cardiovascular disease history between groups (greater or less than median LAV). Furthermore, heart rate, LVEF, LVMi, and left ventricle chamber size (at enddiastole and end-systole) were similar in both groups. This suggests that left atrial size was not a marker of reduced diastolic filling time or impaired left ventricular systolic emptying.
In the survival analysis, higher LAV was significantly associated with poorer survival (Fig 2A). Multivariate analysis also showed that increased LAV/BSA and presence of LVSD independently predicted death in ESRD patients with LVH. As previously demonstrated, a clinical history of ischemic heart disease was a significant independent predictor of death, and kidney transplantation was independently associated with significantly improved survival. These data confirm previous studies investigating LAV and survival. However, in previous studies no analyses were performed in patients with pre-existing myocardial abnormalities. We believe that the strength of this study lies with accurate assessment of LVMi using CMR imaging, which provides an accurate, volume-independent, and reproducible method of measuring LVM.
In the present study, LAV was not significantly correlated with LVM, suggesting that increased LAV is not caused solely by impaired atrial emptying into a large, poorly compliant, left ventricle. In other patient populations, increased LAVs are considered to reflect the long-term effects of increased ventricular filling pressures. Thus, when filling pressures are increased, the atria (like the ventricles) will enlarge in response to pressure and chronic volume overload. LAV was not associated with diastolic dysfunction in our population as measured using E:A ratio. We did not obtain tissue Doppler or pulmonary venous blood flow velocity data to fully evaluate diastolic function, and it is likely that diastolic dysfunction is common in ESRD patients with LVH. Thus, we postulate that in ESRD patients with LVH, left atrial enlargement is caused largely by diastolic dysfunction and also chronic fluid overload due to expansion in intravascular volume.
A decrease in LAV has been achieved in patients with mitral valve disease and atrial fibrillation; however, its effect on overall prognosis is unknown. Whether tight control of fluid volume status in patients with ESRD similarly decreases LAV and mortality will require a controlled clinical trial.
We have previously shown that LVSD is associated with (often asymptomatic) ischemic heart disease and, in turn, poorer survival. This presumably is caused by occlusive large-vessel disease and inadequate growth of penetrating epicardial vessels in response to cardiac myocyte hypertrophy. Ventricular action potential propagation and recovery are impaired in the presence of LVH and LVSD, increasing the risk of ventricular re-entrant tachyarrhythmias. The prognostic benefits of improving systolic function of ESRD patients, using pharmacological approaches or modification of dialysis remain to be assessed.
In contrast to previous studies using echocardiography to assess left ventricular dimensions, higher LVMi was not a significant predictor of mortality in this cohort of patients. We have previously shown inaccuracies of echocardiography measurements in patients with ESRD, and it is likely that the use of CMR imaging to accurately assess myocardial mass will provide more reliable prognostic information in the future.
We accept that this study has some limitations. Patients recruited to this study were being assessed for kidney transplantation and may not be representative of all patients with ESRD. However, since these patients were considered healthy enough to be considered for a kidney transplant, we believe these results would be relevant to other patients with more significant comorbid conditions. In addition, we obtained limited information regarding ventricular diastolic function in our cohort and hopefully, as more detailed methods (eg, tissue Doppler) are utilized, these data will become available.
In conclusion, in ESRD patients with LVH, increased LAV and the presence of LVSD are independent predictors of death and may provide novel factors that may be amenable to modification to improve cardiovascular prognosis. Patel et al. Page 9
Published as: Am J Kidney Dis. 2010 June ; 55(6): 1088-1096. Patel et al. Page 10 Published as: Am J Kidney Dis. 2010 June ; 55(6): 1088-1096. Sponsored Document Sponsored Document Sponsored Document Sponsored Document Sponsored Document Sponsored Document Patel et al. Page 11 Published as: Am J Kidney Dis. 2010 June ; 55(6): 1088-1096.
Published as: Am J Kidney Dis. 2010 June ; 55(6): 1088-1096.
a Assessed using cardiovascular magnetic resonance imaging (ejection fraction <55%).
Published as: Am J Kidney Dis. 2010 June ; 55(6): 1088-1096.
a Never smoking is the reference category.
Published as: Am J Kidney Dis. 2010 June ; 55(6): 1088-1096.Sponsored DocumentSponsored DocumentSponsored Document
Published as: Am J Kidney Dis. 2010 June ; 55(6): 1088-1096.
Support: The study was supported by the
The authors declare that they have no relevant financial interests.
Intelligence is a highly heritable trait for which it has proven difficult to identify the actual genes. In the past decade, five whole-genome linkage scans have suggested genomic regions important to human intelligence; however, so far none of the responsible genes or variants in those regions have been identified. Apart from these regions, a handful of candidate genes have been identified, although most of these are in need of replication. The recent growth in publicly available data sets that contain both whole genome association data and a wealth of phenotypic data, serves as an excellent resource for fine mapping and candidate gene replication. We used the publicly available data of 947 families participating in the International Multi-Centre ADHD Genetics (IMAGE) study to conduct an in silico fine mapping study of previously associated genomic locations, and to attempt replication of previously reported candidate genes for intelligence. Although this sample was ascertained for attention deficit/hyperactivity disorder (ADHD), intelligence quotient (IQ) scores were distributed normally. We tested 667 single nucleotide polymorphisms (SNPs) within 15 previously reported candidate genes for intelligence and 29451 SNPs in five genomic loci previously identified through whole genome linkage and association analyses. Significant SNPs were tested in four independent samples (4,357 subjects), one ascertained for ADHD, and three population-based samples. Associations between intelligence and SNPs in the ATXN1 and TRIM31 genes and in three genomic locations showed replicated association, but only in the samples ascertained for ADHD, suggesting that these genetic variants become particularly relevant to IQ on the background of a psychiatric disorder.
Intelligence is a highly heritable complex trait, for which it is hypothesized that many genes of small effect size contribute to its variability [McClearn et al., 1997;Plomin, 1999]. Almost a decade after the completion of a rough draft of the human genome sequence, major efforts have been undertaken to identify common variations related to inter-individual differences in intelligence. Plomin and coworkers [Plomin, 1999;Plomin et al., 2001Plomin et al., , 2004;;Butcher et al., 2005Butcher et al., , 2008] ] conducted several genome wide association (GWA) studies and showed significant association of a functional polymorphism in ALDH5A1 (aldehyde dehydrogenase 5 family) (MIM: 610045) on chromosome 6p with intelligence. Whole genome linkage scans for intelligence [Posthuma et al., 2005;Buyske et al., 2006;Dick et al., 2006;Luciano et al., 2006] reported two areas of genome-wide significant linkage for general intelligence on the long arm of chromosome 2 (2q24.1-31.1) and the short arm of chromosome 6 (6p25-21.2), and several areas of suggestive linkage (4p, 7q, 14q, 20p, 21p), following Lander and Kruglyak guidelines [1995]. The region on chromosome 6 (6p25-21.2) overlaps with the locus (6p24.1) identified in the genomewide association study performed by Butcher et al. [2008]. Converging evidence from these whole genome studies provides support for the involvement of six different chromosomal regions, 2q24.1-31.1, 2q31.3, 6p25-21.2, 7q32.1, 14q11.2-12, and 16p13.3, in human intelligence (see Table I).
Apart from whole genome searches, several candidate genebased association analyses have also reported significant associations with human intelligence [for a review see Posthuma and de Geus, 2006]. Based on a literature search, we identified 16 genes that have been associated with intelligence, as measured with an intelligence quotient test (IQ) at least once (P-value 0.05); DTNBP1 (dystrobrevin-binding protein 1) (MIM: 607145), ALDH5A1 (aldehyde dehydrogenase 5 family, member A1) (MIM: 610045), IGF2R (insulin-like growth factor 2 receptor) (MIM: 147280), CHRM2 (cholinergic muscarinic receptor 2) (MIM: 118493), BDNF (brain-derived neurotrophic factor) (MIM: 113505), CTSD (cathepsin D) (MIM: 116840), DRD2 (dopamine receptor D2) (MIM: 126450), KL (klotho) (MIM: 604824), APOE (apolipoprotein E) (MIM: 107741), SNAP25 (synaptosomal-associated protein, 25 kDa) (MIM: 600322), PRNP (prion protein (p27-30)) (MIM: 176640), CBS (cystathionine-beta-synthase) (MIM: 236200), COMT (catechol-O-methyltransferase) (MIM: 116790), DNAJC13 (DnaJ (Hsp40)) (GeneID: 23317), FADS3 (fatty acid desaturase 3) (MIM: 606150), and TBC1D7 (TBC1 domain family, member 7) (GeneID: 51256) (see Table II).
One of the major hurdles in identifying genes for complex traits is the need for replication to distinguish false positives from genuine associations. Of all reported genetic association studies in the literature, only 4% have shown replicable association according to a 2002 search [Hirschhorn et al., 2002]. At present, searching for ''genetic'' and ''association'' in PubMed gives 69950 hits (June 2010), while adding the keywords ''replicated'' or ''validated'' results in 1,318 studies. In other words, in this rough scan around 2.0% of the total reported genetic associations are reports of validated genetic association. The field of intelligence shows no exception. Of the 16 genes mentioned above, only three (CHRM2 [Comings et al., 2003;Gosso et al., 2006bGosso et al., , 2007;;Dick et al., 2007], SNAP25 [Gosso et al., 2006a[Gosso et al., , 2008b]], and BDNF [Tsai et al., 2004;Harris et al., 2006]) have shown replicated association with intelligence across independent samples. Several other genes (e.g., COMT, DTNBP1) have repeatedly shown association to a range of cognitive traits, but have not been replicated for association with intelligence as measured with an IQ test [Small et al., 2004;Savitz et al., 2006]. The reasons for lack of replication are many and include different ethnicity, insufficient sample size, different phenotype, opposite effect direction, or the fact that no replication was attempted at all.
The recent growth in publicly available data sets that contain whole genome association data as well as a wealth of phenotypic data serves as an excellent resource for rapid replication efforts. In the public domain, the Genetic Association Information Network (GAIN)-International Multi-Centre ADHD Genetics (IMAGE) sample is the sole GWA sample with information on IQ scores. In the current article, we use data from the IMAGE project, to (a) attempt replication of previous association findings for the 16 genes associated with normal intelligence at least once, and (b) explore the six chromosome regions previously implicated in human intelligence. Associations found in the IMAGE sample (discovery sample) are subsequently attempted for replication in four independent samples. Of these four samples one is ascertained for attention deficit/hyperactivity disorder (ADHD)-as is the IMAGE sample-and three are population-based samples. This allows to investigate whether associated single nucleotide polymorphisms (SNPs) found with the IMAGE sample are discovered due to an association with intelligence in an ADHD population, or are more generally associated with intelligence.
Subjects of the IMAGE project have been described in detail elsewhere [Brookes et al., 2006;Kuntsi et al., 2006;Neale et al.,
TABLE I. Summary of Genomic Loci Previously Associated With Intelligence Locus Refs. Previous population 2q24.1-31.1 Posthuma et al. [2005] Study ¼ 1, population ¼ 1 and 2, N ¼ 950 Luciano et al. [2006] Study ¼ 1, population ¼ 1 and 2, N ¼ 836 2q31.3 Butcher et al. [2008] Study ¼ 1, population ¼ 3, N ¼ 3,195 6p25-21.2 Posthuma et al. [2005] Study ¼ 1, population ¼ 1 and 2, N ¼ 950 Luciano et al. [2006] Study ¼ 1, population ¼ 1 and 2, N ¼ 836 7q32.1 Butcher et al. [2008] Study ¼ 1, population ¼ 3, N ¼ 3,195 14q11.2-12 Buyske et al. [2006] Study ¼ 2, population ¼ 1, N ¼ 1,115 16p13.3 Butcher et al. [2008] Study ¼ 1, population ¼ 3, N ¼ 3,195 Study 1 is a family study, 2 is the COGA (Collaborative Studies on Genetics of Alcoholism) family study. Population 1 is from the Netherlands, 2 is from Australia, 3 is from the United Kingdom. N indicates sample size.
TABLE II. Overview of Genes Previously Associated With Intelligence at Least Once Gene Chr Gene size SNP Position Type Previous P-value Refs. Previous population DNAJC13 3 121371 rs1378810 133736780 Intron 0.0007 Butcher et al. [2008] Study ¼ 3, population ¼ 5, N ¼ 3,195 TBC1D7 6 35001 rs2496143 13419830 Intron 0.037 Butcher et al. [2008] Study ¼ 3, population ¼ 5, N ¼ 3,195 DTNBP1 6 140233 rs1018381 15765048 Intron 0.008 Burdick et al. [2006a, b] Study ¼ 1, population ¼ 1, N ¼ 339 ALDH5A1 6 42238 rs2760118 24611568 Coding-non-synonymous 0.001 Plomin et al. [2004] Study ¼ 4, population ¼ 1, N ¼ 594 IGF2R 6 137452 rs3832385 160446894 mrna-utr 0.02 Chorney et al. [1998] Study ¼ 4, population ¼ 1, N ¼ 102 CHRM2 7 148372 rs8191992 136351847 mrna-utr <0.017 Comings et al. [2003] Study ¼ 4, population ¼ 1, N ¼ 828 rs8191992 136351847 mrna-utr 0.036 Dick et al. [2007] Study ¼ 2, population ¼ 1, N ¼ 1,113 rs1378650 136355690 -0.028 Dick et al. [2007] Study ¼ 2, population ¼ 1, N ¼ 1113 rs1424548 136360299 -0.037 Dick et al. [2007] Study ¼ 2, population ¼ 1, N ¼ 1113 rs2350780 136243508 Intron 0.016 Dick et al. [2007] Study ¼ 2, population ¼ 1, N ¼ 1113 rs2350786 136327109 Intron 0.016 Dick et al. [2007] Study ¼ 2, population ¼ 1, N ¼ 1113 rs6948054 136331340 Intron 0.04 Dick et al. [2007] Study ¼ 2, population ¼ 1, N ¼ 1113 rs7799047 136322097 Intron 0.02 Dick et al. [2007] Study ¼ 2, population ¼ 1, N ¼ 1113 rs324640 136339535 Intron <0.001 Gosso et al. [2006b] Study ¼ 3, population ¼ 2, N ¼ 667 rs324650 136344200 Intron <0.01 Gosso et al. [2006b] Study ¼ 3, population ¼ 2, N ¼ 667 rs2061174 136311939 Intron <0.01 Gosso et al. [2007] Study ¼ 3, population ¼ 2, N ¼ 762 rs2061174 136311939 Intron 0.016 Dick et al. [2007] Study ¼ 2, population ¼ 1, N ¼ 1113 BDNF 11 66856 rs6265 27636491 Coding-non-synonymous 0.046 Tsai et al. [2004] Study ¼ 4, population ¼ 3, N ¼ 114 rs6265 27636491 Coding-non-synonymous 0.001 Harris et al. [2006] Study ¼ 4, population ¼ 4, N ¼ 904 CTSD 11 11237 rs17571 1739169 Coding-non-synonymous 0.01 Payton et al. [2006] Study ¼ 4, population ¼ 1, N ¼ 767 FADS3 11 91903 rs174455 61412693 Intron 0.013 Butcher et al. [2008] Study ¼ 3, population ¼ 5, N ¼ 3195 DRD2 11 65564 rs2075654 112794275 Intron 0.05 Gosso et al. [2008a] Study ¼ 3, population ¼ 2, N ¼ 762 KL 13 49708 rs9536314 32526137 Coding-non-synonymous 0.011 Deary et al. [2005] Study ¼ 4, population ¼ 4, N ¼ 915 APOE 19 3611 rs28931577 50103741 Coding-non-synonymous 0.009 Deary et al. [2002] Study ¼ 4, population ¼ 4, N ¼ 466 rs769455 50103879 Coding-non-synonymous 0.009 Deary et al. [2002] Study ¼ 4, population ¼ 4, N ¼ 466 SNAP25 20 88588 rs362602 10241527 -0.005 Gosso et al. [2006a] Study ¼ 3, population ¼ 2, N ¼ 762 rs363039 10168495 Intron 0.001 Gosso et al. [2006a] Study ¼ 3, population ¼ 2, N ¼ 762 rs363050 10182256 Intron 0.0002 Gosso et al. [2006a] Study ¼ 3, population ¼ 2, N ¼ 762 PRNP 20 15437 rs1799990 4628250 Coding-non-synonymous 0.006 Kachiwala et al. [2005] Study ¼ 4, population ¼ 4, N ¼ 915 CBS 21 23120 rs5742905 43356252 Coding-non-synonymous 0.02 Barbaux et al. [2000] Study ¼ 4, Population ¼ 1, N ¼ 202 COMT 22 27221 rs4680 18331270 Coding-non-synonymous 0.05 Gosso et al. [2008a]
Study 1 is a control and case schizophrenia study, 2 is the COGA (Collaborative Studies on Genetics of Alcoholism) family study, 3 is a family study, 4 is a general population study. Population 1 is from the USA, 2 is from the Netherlands, 3 is from China, 4 is from Scotland, 5 is from the UK. N indicates sample size.
2008]. Briefly, 947 European Caucasian nuclear families (2,844 individuals) from eight countries (Belgium, England, Germany, Holland, Ireland, Israel, Spain, and Switzerland) were included in the analysis. Families had been recruited based on having one child with ADHD and another who would provide DNA and quantitative trait data. In addition, both parents had to be available for DNA sampling. IQ scores were available for 606 unrelated probands (for which we also had genotyping data, see below), of which 554 were males, with a mean age of 10.99 (SD 2.74). IQ was measured with the WISC-III-R (Wechsler Intelligence Scales for children) [Wechsler, 1991] or the WAIS-III-R (Wechsler Adult Intelligence Scale) [Wechsler, 1997] when appropriate (for children aged 17 and older).
The Verbal subtests Vocabulary and Similarities, and the Performance subtests Picture Completion and Block Design from the WISC were used to obtain an estimate of a child's IQ (prorated following procedures described by Sattler [1992]). Age-appropriate national population norms were available for each participating site included in the IMAGE sample and these were used to derive standardized estimates of intelligence [Sonuga-Barke et al., 2008]. Standardized Full-Scale IQ (FSIQ) scores had a median of 101.6 and a mean of 100.7 (SD 15.7). Skewness of the distribution of IQ scores was 0.063 while kurtosis was À0.075. The Shapiro-Wilk test was non-significant (P ¼ 0.517) suggesting that the distribution of IQ in the IMAGE sample did not deviate form a normal distribution (see Fig. 1).
The parents of the probands filled out the Conner's questionnaire, which provides a quantitative measure of ADHD symptoms. Correlations between the symptom scores on the Conner's Questionnaire and IQ were À0.066 (P ¼ 0.074.) for the total score, À0.029 (P ¼ 0.442), for the inattention score, and À0.084 (P ¼ 0.024) for the hyperactivity/impulsivity score. Although this sample was originally ascertained for ADHD, and ADHD and IQ have been reported to be associated [Frazier et al., 2004], these findings suggest that in this sample IQ scores are normally distributed (as would be expected in a population-based sample) and are at most very weakly related to ADHD symptom scores. As there were mean fluctuations across collection sites, we calculated Z-scores within each site/country. The use of Z-scores ensures that there are no mean IQ differences left across subpopulations in the IMAGE sample and therefore rules out spurious associations due to the known subpopulation structure.
The IMAGE study was genotyped as part of the GAIN initiative, a public-private partnership of the FNIH (Foundation for the National Institutes of Health, Inc.) that currently involves NIH, Pfizer, Affymetrix, Perlegen Sciences, Abbott, and the Eli and the Edythe Broad Institute of MIT and Harvard University (http:// www.fnih.org). Genotyping was conducted at Perlegen Sciences using their genotyping platform, which comprises approximately 600,000 tagging SNPs designed to be in high linkage disequilibrium with untyped SNPs for the HapMap populations. Genotype data were cleaned by NCBI (The National Center for Biotechnology Information). Quality control analyses were processed using the GAIN QA/QC Software Package (version 0.7.4) developed by Gon‚ calo Abecasis and Shyam Gopalakrishnan at the University of Michigan. Details of the genotyping and data cleaning process for the IMAGE study (study accession phs000016.v1.p1) have been reported elsewhere [Neale et al., 2008].
Briefly, we selected only SNPs with a minor allele frequency (MAF) !0.05 and Hardy-Weinberg equilibrium (HWE) (P ! 1 Â 10 À6 ). Genotypes causing Mendelian inconsistencies were identified by PLINK and removed (http://pngu.mgh.harvard.edu/ purcell/plink/) [Purcell et al., 2007]. We additionally removed SNPs that failed the quality control metrics for the other two GAIN Perlegen studies (i.e., Major Depression Disorder [dbGAP study accession, phs000020.v1.p1) and Psoriasis (dbGAP study accession, phs000019.v1.p1), see Neale et al., 2008]. With this filtering, 384,401 SNPs were retained in the final data set. One genomic intelligence locus (7q32.1) could not be included in the analysis because all 10 SNPs inside this relatively small area failed the quality control. The APOE was also not included as no SNPs were genotyped in or near this gene. Fifteen genes (ALDH5A1, BDNF, CBS, CHRM2, COMT, CTSD, DNAJC13, DRD2, DTNBP1, FADS3, IGF2R, KLOTHO, PRNP, SNAP25, and TBC1D7) and five genomic areas (2q24.1-31.1, 2q31.3, 6p25-21.2, 14q11.2-12, and 6p13.3) were thus included in the association analysis. From the cleaned data set, we selected all genotyped SNPs that lie in these candidate genes and genomic loci including 10 kb both upstream and downstream of each gene or genomic locus.
To increase coverage in the targeted genomic areas, we used the imputation approach implemented in MACH [Li et al., 2006], which imputes genotypes of SNPs that are not directly genotyped in the data set, but that are present on a reference panel. MACH is a cohort, consisting of 1,670 Australians (793 male, 877 female) from 741 families with mean age of 16.4 (SD ¼ 4). FSIQ was assessed with the Multidimensional Aptitude Battery [MAB; Jackson, 1984]. Five subtests were administered (three Verbal: Information, Arithmetic, Vocabulary; two Performance: Spatial, Object Assembly) and from these a standardized FSIQ measure was obtained. FSIQ had a mean of 112.6 (SD ¼ 12.8). Genotyping was done using the Illumina 610K SNP platform and Illumina BeadStudio software, with 529,721 SNPs passing QC. Data were imputed to $2.3 million SNPs with the use of the phased data from the HapMap samples (CEU; build 36, release 22) and MACH.9, described in detail in Medland et al. [2009], (see Project 5: ADOL deCODE). Individual SNPs were tested for association with the family-based score test implemented in Merlin. This study was approved by the QIMR human research ethics committee and informed written consent was obtained from all participants.
Lothian Birth Cohort 1936 (LBC1936) sample. The LBC1936 consisted of 1,091 individuals who, at the age of $11 years, participated in the Scottish Mental Survey of 1947, when they took a validated mental ability test, the Moray House Test No. 12 (MHT). Briefly, at a mean age of 69.6 years (SD ¼ 0.8) participants of LBC1936 were recruited to a study to investigate the causes of cognitive ageing. They underwent a series of cognitive, physical, and biochemical tests at the Wellcome Trust Clinical Research Facility (WTCRF) at the Western General Hospital, Edinburgh. For this study, a general cognitive ability factor was derived from principal components analysis of six Wechsler Adult Intelligence Scale-III UK (WAIS-III) subtests (matrix reasoning, letter number sequencing, block design, symbol search, digit span backwards, and digit symbol), as described previously [Luciano et al., 2009]. The general cognitive ability factor scores were corrected for age in days and sex, and converted to IQ scores (mean ¼ 100; SD ¼ 15). DNA was isolated by standard procedure at the WTCRF Genetics Core, Western General Hospital, Edinburgh from 1,071 individuals. Twenty-nine samples failed quality control preceding the genotyping procedure. The remaining 1,042 samples (all blood-extracted) were genotyped at the WTCRF Genetics Core using the Illumina610 -Quadv1 chip. These samples were then subjected to the following quality control procedures after which 1,005 samples remained. All individuals were checked for disagreement between genetic and reported gender (n ¼ 12). Relatedness between subjects was investigated and, for any related pair of individuals, one was removed (n ¼ 8). Samples with a call rate 0.95 (n ¼ 16), and those showing evidence of non-Caucasian ascent by multidimensional scaling, were also removed (n ¼ 1). SNPs were included in the analyses if they met the following conditions: call rate !0.98, MAF !0.01, and HWE test with P ! 0.001. The final number of SNPs included in the genome-wide association study was 549,091. IQ scores and genotype were available for 976 individuals. Genomic coverage was extended to $2.5 million common SNPs by imputation using the HapMap phase II CEU data (NCBI build 36 (UCSC hg18)) as the reference sample and MACH software. SNPs with low imputation (r 2 < 0.30), low MAF (<0.01), and divergence from HWE (P < 0.001) were excluded so that respective SNP and sample call rates were 0.98 and 0.95.
Statistical power. The primary (IMAGE) sample of 606 subjects had sufficient (80%) statistical power to detect SNPs that explained at least 1.3% of the variance for direct replication (significance level 0.05) (Genetic Power Calculator) [Purcell et al., 2003], which is in the order of effect sizes of SNPs reported previously. The sample size of the meta-analysis including the two ADHD samples (606 þ 216 ¼ 822) was sufficient to detect genetic effects explaining 2% of the variance, given a Bonferroni corrected significance level of 0.001. The sample size including all samples (N ¼ 4,963) was sufficient to detect SNPs explaining 0.35% (i.e., <1%) of the variance (significance level of 0.001).
Replication analysis. All populations were imputed using MACH and imputed SNPs were included in our analysis if quality score > 0.9 and r 2 > 0.3 and MAF > 0.05. IQ scores were all corrected for effects of age and sex and transformed to Z-scores and standardized such that the mean was 100 and SD ¼ 15, within each sample, for comparison of effect sizes across samples.
Although replication across different samples provides information on the genuineness of an initial association, meta-analysis appropriately weighs the effect and sample sizes across different replications samples. We thus conducted a metaanalysis, in which the primary sample was included to increase statistical power [Skol et al., 2006]. We used a stepwise approach, in which we first ran a combined analysis based on the two samples ascertained for ADHD, and then conducted a meta-analysis on all 4,963 subjects. The meta-analysis was conducted using the METAL program (http://www.sph.umich.edu/csg/abecasis/metal/). MET-AL creates a single summary P-value for each SNP from all samples together. For each marker, an arbitrary reference allele is selected and a Z-statistic, characterizing the evidence for association, is used as input. The Z-statistic summarizes the magnitude and the direction of an effect relative to the reference allele. An overall Z-statistic and P-value are then calculated from the weighted average of the individual statistics. Weights are proportional to the square root of the number of individuals examined in each sample, and selected such that the squared weights sum to 1.0. Outcomes of the metaanalyses were tested against a Bonferroni corrected threshold of significance (P < 0.001).
Most previously reported associations of genes with intelligence included intronic SNPs with no clear function. This suggests that they might be controlling RNA signaling networks or that other SNPs in LD might be the actual causal variant. We used imputation to increase coverage. We do note; however, that even after imputation, not all of the originally reported SNPs were available in the current sample. Of the 15 candidate genes, six genes showed at least one SNP with a P-value <0.05 (see Table III).
Of the five previously reported genomic loci (2q24.1-31.1, 2q31.3, 6p25-21.2, 14q11.2-12, and 16p13) investigated here, we observed P-values <0.0025 in three regions (6p25-21.2, 2q24.1-31.1, and 14q11.2-12) (see Table IV). Genomic areas 2q31.3 and 16p13.3 showed no association with IQ (all P-values >0.15). On a SNP level, there were three independent SNPs in intergenic and non-coding regions with P-values 2.0 Â 10 À4 inside the 2q24.1-31.1 and 14q11.2-12 areas (Table IV). The lowest P-values were
This study aimed to replicate association of previously reported candidate genes for IQ as well as to fine-map previously linked genomic areas. As available samples differed in ascertainment method (i.e., ascertained for ADHD or population based) we tested for SNP associations with IQ in an ADHD background and in a non -ADHD, general population, background.
In the primary analysis, we found weak evidence for the association of some of the previously reported genes with IQ: IGF2R (five SNPs with P-value <0.05), DTNBP1 (five SNPs with P-value <0.05), ALDHA5A1 (one SNP with P-value <0.05), BDNF (two SNPs with P-value <0.05), DRD2 (two SNPs with P-value <0.05), and CHRM2 (two SNPs with P-value 0.03). None of SNPs previously associated with IQ showed association in the current study (P-value >0.05). The lack of replication can either indicate a false positive finding in previous studies, or might be explained by the ascertainment for ADHD in our primary sample. Although association between IQ and ADHD in the current sample was not significant, and IQ was distributed normally in the IMAGE sample, previous reports [e.g., Kuntsi et al., 2004]
Results from the primary association analysis in the genomic loci implicated three intergenic regions (2q24.1-31, 6p25-21.2, and 14q11.2-12). The nominally significant SNPs from the candidate genes, and the top SNPs from the genomic regions, were included in a stepwise combined analysis. When we combined the two samples ascertained for ADHD (totaling 822 subjects), we found that five SNPs were associated with IQ. None of these SNPs were inside candidate genes previously implicated, but instead were located in two genomic areas: 6p25-21.2 and 14q11.2-12. Two of these SNPs were inside two genes: rs17606174 was in the second intron of the ATXN1 gene, and rs2023472 in exon 5 on TRIM31. However, when we combined all samples, none of these SNPs showed a significant association with intelligence. However, we cannot exclude the possibility of type I error given the total number of tests performed within the discovery sample only. These results provide suggestive evidence that the ATXN1 and TRIM31 genes, and several other SNPs in areas 6p25-21.2 and 14q11.2-12, are related to IQ, but only on the background of ADHD.
In the primary IMAGE association results, ATXN1 has 25 SNPs with P-value <0.05, and most of them are located in the second intron of ATXN1, nearby an alternative splicing region. ATXN1 is present in the nucleus of the neurons of the basal ganglia, pons and cortex, and in both cytoplasm and nucleus of Purkinje cells of the cerebellum [Servadio et al., 1995]. Expansion of a (CAG)n repeat in ATXN1 (previous called SCA1 gene) causes spinocerebellar ataxia-1 (SCA1) in humans (MIM: 164400) [Orr et al., 1993;
TABLE IV. Replication Results in the Genomic Loci Previously Associated With Intelligence in the IMAGE Cohort SNP (G/I) a Minor/major allele MAF Rank P-value Chr Position Type Closest gene Distance to gene Genomic location 2q24.1-31.1 (from 154475832 to 177730691 bp) total SNPs tested ¼ 7,819 in 182 genes rs4972741 (I) G/A 0.12 1 0.00017 2 172823906 Intergenic AC104088.1 À64,355 rs6721348 (I) C/T 0.12 2 0.00018 2 172826755 Intergenic AC104088.1 À61,506 rs10172929 (G) G/T 0.13 3 0.00031 2 164756952 Intergenic AC092684.1 0 rs16844374 (G) C/T 0.15 7 0.00127 2 160394348 Intronic LY75 0 rs10201330 (I) T/C 0.09 4 0.00132 2 177056271 Intergenic AC017048.3 22,948 rs4289149 (G) A/G 0.18 8 0.00150 2 172834736 Intergenic ITGA6 165,264 rs995711 (G) G/T 0.12 5 0.00174 2 164123635 Intergenic FIGN À34,517 rs11896469 (G) C/T 0.44 6 0.00230 2 176388492 Intergenic EXTL2P1 27,369 Genomic location 6p25-21.2 (from 5945435 to 41007859 bp) total SNPs tested ¼ 18,651 in 809 genes rs12204969 (I) C/T 0.12 1 0.00018 6 16802156 Intronic ATXN1 0 rs17606216 (G) C/T 0.12 2 0.00018 6 16796594 Intronic ATXN1 0 rs993600 (G) G/A 0.16 3 0.00027 6 22153623 Within non-coding gene RP1-67M12.1 0 rs2023472 (G) A/G 0.42 4 0.00028 6 30183843 Intergenic TRIM31 5,241 rs6929819 (G) G/A 0.43 5 0.00033 6 33670832 Intergenic C6orf227 1,739 rs195371 (G) G/A 0.23 6 0.00034 6 37412364 Intergenic TBC1 3,464-rs6929774 (I) T/C 0.42 7 0.00039 6 33670698 Intergenic C6orf227 1,605 rs17606174 (G) T/C 0.13 8 0.00050 6 16795524 Intronic ATXN1 0 Genomic location 14q11.2-12 (from 21269202 to 28322992 bp) total SNPs tested ¼ 2,964 in 233 genes rs2807822 (I) T/C 0.47 1 0.00010 14 27554764 Intergenic AL445384.1 25,600 rs3811222 (I) A/G 0.10 2 0.00066 14 22020854 Intronic TRAC 0 rs762578 (I) T/G 0.11 3 0.00069 14 22020088 Intronic TRAC 0 rs1872159 (G) T/C 0.09 4 0.00080 14 22017743 Intronic TRAC 0 rs7149201 (I) C/T 0.20 5 0.00178 14 23034259 Intergenic NGDN 17,017 rs877726 (G) T/A 0.23 6 0.00230 14 27557719 Intergenic AL445384.1 À28,555 Only SNPs with a P < 0.0025 are shown. Genome build 36. a G and I indicate genotyped and imputed SNPs, respectively.
TABLE V. Meta-Analysis of the Top SNPs From the Primary Analysis Gene Genomic area SNP A IMAGE (N ¼ 606) DUKE (N ¼ 216) ALSPAC (N ¼ 1,495) QIMR (N ¼ 1,670) LBC1936 (N ¼ 976) IMAGE and DUKE (N ¼ 822) ALL (N ¼ 4,963) B P B P B P B P B P Z P Z P Intergenic 2q24.1-31.1 rs10172929 T 4.64 0.0003 0.23 0.84 À0.03 0.96 À0.94 0.35 0.95 0.34 Intergenic 2q24.1-31.1 rs10201330 T 5.43 0.0013 1.28 0.27 À0.76 0.38 0.19 0.88 1.32 0.19 CHRM2 7q33 rs10271552 T À3.18 0.0398 À0.42 0.65 À0.23 0.76 À1.38 0.23 À1.72 0.09 Intergenic 2q24.1-31.1 rs11896469 T 2.55 0.0023 À0.01 0.71 0.47 0.32 À0.38 0.58 1.18 0.24 ATXN1 6p25-21.2 rs12204969 T 4.98 0.0002 À2.12 0.01 À0.99 0.14 0.49 0.61 0.65 0.51 BDNF 11p14 rs12273363 T À2.61 0.0131 0.88 0.68 À0.42 0.54 0.13 0.82 0.70 0.43 À2.00 0.0456 0.95 0.34 BDNF 11p14 rs12288512 A 2.67 0.0114 À0.88 0.68 0.42 0.55 À0.13 0.82 À0.70 0.43 2.04 0.0410 1.32 0.19 LY75 2q24.1-31.1 rs16844374 T 3.77 0.0013 ATXN1 6p25-21.2 rs17606174 T À4.47 0.0005 À3.95 0.16 0.12 0.02 0.40 0.53 À0.76 0.41 À3.80 0.0001 À0.22 0.83 ATXN1 6p25-21.2 rs17606216 T 4.91 0.0002 À2.15 0.01 À0.86 0.20 0.43 0.65 À0.67 0.51 IGF2R 6q26 rs1805075 A À4.47 0.0189 À0.30 0.82 0.76 0.50 À0.62 0.71 À0.73 0.47 Intergenic 14q11.2-12 rs1872159 T 4.77 0.0008 2.61 0.46 À0.30 0.76 0.58 0.43 0.75 0.47 3.32 0.0009 1.92 0.06 Intergenic 6p25-21.2 rs195371 A À3.65 0.0003 À0.14 0.00 0.62 0.28 À0.23 0.78 À2.41 0.02 TRIM31 6p25-21.2 rs2023472 A À3.09 0.0003 À1.58 0.35 0.19 0.75 0.95 0.05 0.23 0.73 À3.65 0.0003 À0.25 0.80 DTNBP1 6p23 rs2619545 T À2.34 0.0362 2.07 0.34 0.25 0.73 À0.34 0.56 0.87 0.30 À1.41 0.1589 À1.98 0.05 ALDH5A1 6p23 rs2760138 A À3.05 0.0478 Intergenic 14q11.2-12 rs2807822 T À3.63 0.0000 0.39 0.82 0.09 0.89 0.77 0.10 À0.19 0.78 À3.34 0.0009 0.54 0.59 Intergenic 14q11.2-12 rs3811222 A 5.15 0.0007 0.91 0.79 À0.47 0.68 0.59 0.41 0.98 0.33 3.16 0.0016 À1.87 0.06 Intergenic 2q24.1-31.1 rs4289149 A 3.36 0.0015 À3.52 0.14 0.00 0.99 1.18 0.05 À0.62 0.48 2.08 0.0380 1.67 0.10 DRD2 11q23 rs4630328 A À1.79 0.0475 0.95 0.14 0.18 0.71 0.33 0.66 À0.74 0.46 Intergenic 2q24.1-31.1 rs4972741 A À5.51 0.0002 À1.15 0.24 À0.23 0.78 0.80 0.51 À1.87 0.06 CHRM2 7q33 rs6467694 T À3.86 0.0097 À1.47 0.59 0.92 0.32 1.11 0.14 À0.94 0.41 À2.54 0.0112 À0.73 0.47 DRD2 11q23 rs6589377 A 1.71 0.0504 À0.29 0.87 À0.92 0.13 À0.18 0.71 À0.53 0.46 1.65 0.0986 1.92 0.06 Intergenic 2q24.1-31.1 rs6721348 A À5.51 0.0002 À1.09 0.26 À0.28 0.73 0.79 0.51 0.04 0.97 Intergenic 6p25-21.2 rs6929774 T À3.09 0.0004 1.98 0.23 0.05 0.18 À0.63 0.19 0.43 0.53 À2.62 0.0089 À0.77 0.44 C6orf227 6p25-21.2 rs6929819 A 3.08 0.0003 À1.09 0.26 À0.05 0.20 0.65 0.18 À0.41 0.55 2.65 0.0081 0.86 0.39 Intergenic 14q11.2-12 rs7149201 T 3.72 0.0018 À2.30 0.42 0.36 0.62 À0.39 0.49 0.49 0.54 2.45 0.0144 1.10 0.27 DTNBP1 6p23 rs760666 A À2.34 0.0223 0.87 0.69 À0.33 0.64 0.13 0.81 À0.11 0.90 À1.85 0.0644 À0.91 0.36 Intergenic 14q11.2-12 rs762578 T 4.62 0.0007 2.32 0.40 À0.31 0.73 0.73 0.29 0.27 0.79 3.40 0.0007 1.89 0.06 DTNBP1 6p23 rs7758659 T À2.34 0.0225 0.88 0.69 À0.34 0.63 0.13 0.81 À0.11 0.89 À1.85 0.0648 À0.91 0.36 IGF2R 6q26 rs8191818 T À4.49 0.0200 À2.79 0.35 À0.36 0.80 0.65 0.56 À0.47 0.78 À2.49 0.0126 À0.92 0.36 IGF2R 6q26 rs8191821 T 4.52 0.0196 2.79 0.35 0.34 0.80 À0.65 0.56 0.47 0.78 2.50 0.0124 0.92 0.36 IGF2R 6q26 rs8191898 T 4.47 0.0185 2.78 0.35 0.31 0.82 À0.76 0.50 0.47 0.78 2.52 0.0117 0.86 0.39 DTNBP1 6p23 rs875462 T 2.21 0.0318 À1.20 0.57 0.39 0.57 À0.17 0.76 0.59 0.47 1.64 0.1012 1.11 0.27 Intergenic 14q11.2-12 rs877726 A À3.09 0.0023 À0.55 0.79 0.00 0.92 0.65 0.23 À0.41 0.59 À2.75 0.0059 À0.67 0.50 DTNBP1 6p23 rs9296983 A À2.33 0.0229 0.87 0.69 À0.33 0.64 0.16 0.77 À0.11 0.90 À1.84 0.0657 1.10 0.27 IGF2R 6q26 rs9457827 T 4.47 0.0187 2.78 0.35 0.30 0.82 À0.76 0.50 0.62 0.71 2.52 0.0119 0.89 0.37 Non-coding gene 6p25-21.2 rs993600 A 4.11 0.0003 1.41 0.49 0.33 0.67 0.00 0.99 0.13 0.88 3.55 0.0004 1.71 0.09 Intergenic 2q24.1-31.1 rs995711 T 4.02 0.0018 2.16 0.44 1.19 0.25 À0.67 0.40 0.27 0.84 3.15 0.0017 1.47 0.14 A is the reference allele; B is the beta effect;
P is the P-value, Z is the Z-statistic as provided in the meta-analysis.
Bold indicates below the Bonferroni corrected threshold of significance (<0.001) and same direction of effect in the meta-analysis. Banfi et al., 1994]. It was also reported that mice lacking ATXN1 are characterized by decreased exploratory behavior, pronounced deficits in the spatial version of the Morris water maze test, and impaired performance on the rotating rod apparatus [Matilla et al., 1998], pointing to the possible role of ATXN1 in learning and memory.
In the primary IMAGE association results, TRIM31 has 23 SNPs with P-value <0.05 and most of them are located in the 5 0 region and in intron 1 of TRIM31. The protein encoded by this gene is a member of the tripartite motif (TRIM) family. The TRIM motif includes three zinc-binding domains, a RING, a B-box type 1 and a B-box type 2, and a coiled-coil region [Meroni and Diez-Roux, 2005]. Other members of the TRIM family (TRIM3, MIM: 605493) were reported to modulate NGF-induced neurite outgrowth in PC12 cells [El-Husseini and Vincent, 1999].
In summary, we found very little support for genetic variants in genes that have previously been associated with intelligence. In addition, this study did provide tentative support for a role of the ATXN1 and TRIM31 genes in previously associated linkage areas for intelligence in the context of a psychiatric disorder, that is, ADHD. This suggests that genetic variants important for IQ in a non-psychiatric population may not necessary overlap with genetic variants important for IQ in a psychiatric population.
ACKNOWLEDGMENTS We thank all the persons who kindly participated in this research. The
Grant sponsor:
Markov Chain-based haplotyper, which obtains an imputation of each unknown genotype using short stretches of DNA that are shared among unrelated individuals. The reference panel used was HapMap III phased data in MACH input format, which is publicly available for download from the MACH website (http:// www.sph.umich.edu/csg/abecasis/MaCH/download/).
Genomic coverage of the candidate regions was extended to $1.5 Mb common SNPs by imputation using the HapMap phase III CEU data (NCBI build 36 (UCSC hg18)) as the reference sample. Imputed SNPs were selected if r 2 was above 0.3 with the reference allele. Additionally, a quality threshold of 0.90 for imputation was set to be included in further association testing.
Gene coverage was determined by the sum of the typed and imputed SNPs as well as the tagged SNPs (based on HapMap information) divided by the total known common SNPs (again based on HapMap information) within a gene, using WGAviewer [Ge et al., 2008]. On average, after imputation, gene coverage was 85% in the candidate genes, with 100% coverage for DNAJC13, TBC1D7, DTNBP1, ALDH5A1, BDNF, and CTSD. In total, we analyzed 672 SNPs in the candidate genes and 29451 SNPs in the genomic loci.
We carried out association testing using an additive linear regression model implemented in PLINK for genotyped markers, and in MACH2QTL [Li et al., 2009], for imputed SNPs, taking into account dosage information. All IQ scores were precorrected for sex and age and no other covariates were included in the model. As mentioned above, Z-scores were calculated within each of the different sites included in IMAGE, such that there were no mean differences in IQ between sites. Analyses included only SNPs with a minimum 80% genotyping rate and individuals with <20% of missing genotype data. SNPs in candidate genes that had a nominal P-value <0.05, and the top five SNPs from the genomic regions, were selected for testing in the four replication samples.
Four replication samples totaling 4,357 independent subjects were available for replication of top findings of the IMAGE sample. One sample was ascertained for ADHD, and three samples were population-based samples.
DUKE cohort. The DUKE cohort consisted of 216 Americans from 108 families with a DSM-IV diagnosed ADHD-affected proband [Kollins et al., 2008]. Families were enrolled from two collection sites: Duke University Medical Center, Durham, NC, and University of North Carolina, Greensboro, NC. All participating family members provided written informed consent that had been approved by the institutional review board at the ascertaining institution. The WAIS-III was administered to individuals 17 years of age or older, and the WISC-IV was given to children ages 6-16. The Wechsler Preschool and Primary Scale of Intelligence-3rd edition (WPPSI-III) was used for children under the age of 6 [Wechsler, 2002]. FSIQ was estimated for both adults and children from the vocabulary and block design subtests (M ¼ 109.5 109.5 and SD ¼ 12.9). Parents and children were genotyped using the Illumina Infinium HumanHap300 duo chip (Illumina, Inc., San Diego, CA). Quality of the Illumina data was assessed using PLINK (http://pngu.mgh.harvard.edu/purcell/plink/) [Purcell et al. 2007]. SNPs (315,980) were submitted for quality checks. Call rates exceeded 98% for all individuals, one individual was excluded due to a gender discrepancy, and two individuals were excluded due to per-family Mendelian errors in excess of 1%. Out of the 315,980 SNPs submitted, 6,109 SNPs were excluded based on a MAF <0.05, 13 SNPs were excluded due to Mendelian errors in >4 families, and 629 SNPs were excluded due to deviations from HWE (P < 0.000001). In total, 3 individuals and 6,751 SNPs did not pass our quality control checks. Two Centre d'Etude du Polymorphism Humain (CEPH) controls and blinded duplicates were used for every 94 samples and required to match 100%. Data were genomewide imputed with the use of the phased data from the HapMap samples (CEU; build 36, release 22) and MACH. Association analysis was carried out using QTDT (http://www.sph.umich. edu/csg/abecasis/QTDT/). QTDT adopts the between/within model as used by Fulker et al. [1999] and Purcell et al. [2007] as implemented in the QFAM package. We tested for population stratification by comparing the between and within family components of association, using a variant of the orthogonal model [Abecasis et al., 2000]. None of the tested SNPs showed sign of stratification in this population.
The Avon Longitudinal Study of Parents and Children (ALSPAC) is a large population-based, prospective birth cohort consisting initially of over 13,000 women and their children recruited from the Bristol area, UK in the early 1990s [Golding et al., 2001]. ALSPAC has extensive data collections on health and development of children and their parents from the 8th gestational week onwards. Ethical approval for the study was obtained from the ALSPAC Law and Ethics Committee and the local research ethics committees. FSIQ within ALSPAC was measured at the age of 8 with the WISC-III [Wechsler et al., 1992]. A short version of the test consisting of alternate items only (except the coding task) was applied by trained psychologists [Joinson et al., 2007]. Verbal (Information, Similarities, Arithmetic, Vocabulary, and Comprehension) and Performance (Picture Completion, Coding, Picture arrangement, Block Design, and Object assembly) subtests were administered; the subtests were scaled and scores for FSIQ derived. ALSPAC (1,543) children were initially genotyped at 317,504 SNPs on the Illumina HumanHap317K SNP chip. Individuals exhibiting cryptic relatedness, non-European ancestry, high genome-wide heterozygosity, and/or missing rates were excluded as described in Timpson et al. [2009], leaving 1,518 individuals in the analysis of whom 1,495 had information on FSIQ within a range of AE4 SD (M ¼ 106.8, SD ¼ 15.6). Markers with MAF <1%, SNPs with >5% missing genotypes and markers that failed an exact test of HWE (P < 5 Â 10 À6 ) were excluded from further analyses leaving 310,505 SNPs that passed quality control. GWAS analysis was performed on sex and population stratification-adjusted (first five principal components from Eigenstrat analysis) [Price et al., 2006] Z-standardized IQ scores. Genome-wide imputation was done using the HapMap phase I-II CEU data (release 22, NCBI build 36) as the reference sample and MACH software.
The QIMR adolescent cohort is a population-based observed for rs2807822, P ¼ 1 Â 10 À4 ; rs4972741, P ¼ 1.7 Â 10 À4 ; and rs6721348 P ¼ 1.8 Â 10 À4 .
To confirm whether the nominally significant SNPs (P < 0.05) from the candidate genes and the top SNPs (P < 0.0025) in each of the genomic regions with IQ were simply due to chance, we tested these SNPs in the replication samples.
We attempted replication in four independent cohorts. We first performed an association analysis of the 17 nominally associated SNPs (P-value <0.05) in the candidate genes, and the 22 most strongly associated SNPs in the genomic areas in each population (total of 39 SNPs), using the same reference allele for each SNP across different populations. The MAF of the tested SNPs across the five samples were comparable (see Supplementary Table S1).
We first conducted a combined analysis on only the two samples ascertained for ADHD. We then combined all five samples to test whether the significant SNPs were associated with intelligence in a general context, or merely in an ADHD background. Although IQ was normally distributed in both samples ascertained for ADHD, association of a SNP with IQ in an ADHD background may differ from association of that SNP with intelligence in a non-ADHD background.
When combining the two samples ascertained for ADHD we found that of all tested SNPs, 12 had a P-value <0.05 (same direction of effect) of which 6 showed evidence for associated after Bonferroni correction (P < 0.001) for multiple testing. For one of these SNPs (rs2807822, intergenic, 14q11.2-12), however, the effect was in opposite direction in the two samples ascertained for ADHD, also indicated by a significant heterogeneity effect (P ¼ 0.04; see Supplementary Table S2). Three other SNPs were in intergenic areas 6p25-21.2 (one SNP) and 14q11.2-12 (two SNPs), while two SNPs were in genic areas: rs17606174 (P ¼ 0.00018), located in the second intron of ATXN1 (ataxin 1) (MIM: 601556), and rs2023472 (P ¼ 0.0003), located in exon 5 on TRIM31 (tripartite motifcontaining 31) (MIM: 609316). Allelic effect sizes were in the order of 3-4 IQ points in the combined DUKE and IMAGE samples. When we combined all five samples, none of these associations were significant, even though some of the SNPs showed similar direction of effects in some of the replication samples. We provide results in Table V.
TABLE III. Results of 15 Candidate Genes for Intelligence in the IMAGE Cohort GENE Previous associated SNP (G/I) a P-value with previous associated SNP nSNPs tested Coverage SNP density, kb/SNP nSNPs, P < 0.05 Most significant SNP (G/I) a Position Type Best P-value DNAJC13 rs1378810 (I) 0.642 45 1 31.41 0 rs12637073 (I) 133666251 Intronic 0.096 TBC1D7 rs2496143 (I) 0.8568 47 1 9.20 0 rs480122 (G) 13425063 Intronic 0.588 DTNBP1 rs1018381 -65 1 24.55 5 rs760666 (G) 15589121 Intronic 0.020 ALDH5A1 rs2760118 (I) 0.8328 46 1 13.52 1 rs2760138 (I) 24620816 Intronic 0.047 IGF2R rs3832385 -88 0.967 18.62 5 rs8191898 (I) 160418955 Intronic 0.018 CHRM2 rs8191992 -81 0.942 21.10 2 rs6467694 (G) 136197456 Upstream 0.010 rs1378650 (G) 0.8284 rs1424548 (I) 0.3888 rs2350780 (I) 0.9788 rs2350786 (G) 0.6316 rs6948054 (I) 0.322 rs7799047 -rs324640 -rs324650 (I) 0.2514 rs2061174 -BDNF rs6265 (G) 0.1018 29 1 29.86 2 rs12288512 (I) 27704247 Upstream 0.011 CTSD rs17571 (G) 0.2932 6 1 59.01 0 rs3740621 (I) 1728373 Upstream 0.081 FADS3 rs174455 (I) 0.6665 9 0.875 41.54 0 rs174626 (G) 61393633 Downstream 0.050 DRD2 rs2075654 -57 0.982 15.02 2 rs4630328 (I) 112839419 Intronic 0.047 KL rs9536314 (G) 0.6873 52 0.933 13.39 0 rs17763040 (G) 32543384 Intergenic 0.142 SNAP25 rs362602 (G) 0.2254 71 0.913 0 rs362990 (G) 10224221 Intronic 0.063 rs363039 (I) 0.6062 rs363050 (G) 0.3723 PRNP rs1799990 (G) 0.944 21 0.778 20.63 0 rs6084833 (I) 4620759 Intronic 0.135 CBS rs5742905 -22 0.675 19.83 0 rs1788490 (I) 43340620 Intergenic 0.189 COMT rs4680 (G) 0.6209 33 0.633 14.62 0 rs9332377 (I) 18335692 Intronic 0.08 Genome build 36. a G and I indicate genotyped and imputed SNPs, respectively.
Objective-The purpose of this study was to identify risk factors associated with striae gravidarum (SG).
Study design-A cross-sectional study of 112 primiparous women delivering at a private teaching hospital was conducted. Participants were assessed during the immediate postpartum period for evidence of SG. Presence and severity of SG were compared to characteristics of women using t tests and Chi-square tests.
Results-Sixty percent of the study participants had developed SG. Women who developed SG were significantly younger (26.5 ± 4.5 vs 30.5 ± 4.6; P < .001) and had gained significantly more weight during pregnancy (15.6 ± 3.9 vs 38.4 kg ± 2.7; P < .001). Birthweight (BW), gestational age at delivery, and family history of SG were associated with moderate/severe SG.
Conclusion-Maternal age and weight gain during pregnancy are associated with SG. BW, family history of SG, and gestational age at delivery are associated with moderate/severe SG.
Striae distensae or "stretch marks," referred to as striae gravidarum (SG) when they occur in pregnancy, are a common skin problem of considerable cosmetic concern to many patients. They are characterized clinically by linear bands that are initially erythematous to violaceous and gradually fade to become skin colored or hypopigmented atrophic lines that may be thin or wide. SG occur on the abdomen, breasts, buttocks, hips, and thighs and usually develop after the 24th week of gestation.
The cause of SG remains unknown but clearly relates to changes in the structures that provide the skin with its tensile strength and elasticity. Mechanical stretching of the skin in association with hormonal factors has been implicated in the pathogenesis. It has been postulated that some hormones, like estrogen, relaxin, and adrenocortical hormones, decrease the adhesiveness between collagen fibers and increase ground substance, which results in the formation of striae in areas of stretching. Striae may form due to structural connective tissue changes that include realignment and reduced elastin and fibrillin in the dermis. However, some studies have shown that although SG tend to occur in areas of maximum skin stretching, there is no correlation between the degree of striae formation and the extent of body size enlargement during pregnancy. A recent study has found a correlation between the presence of striae and pelvic relaxation, a condition associated with decreased collagen content. This study was limited in that it used self-reported data, and physical exams were not performed.
The data on prevalence of SG and the risk factors associated with their development are scant and often contradictory. It is estimated that up to 90% of pregnant women develop SG; however, some authors report the prevalence to be as low as 50%. Proposed risk factors for the development of SG include family history, race, skin type, birthweight (BW), baseline body mass index (BMI), age, weight gain, and poor nutrition; however, most of these have not been substantiated.
In preparation for a clinical trial on prevention of SG, we conducted a study to determine the incidence of SG in primiparous women in our population and identify the risk factors associated with their development. In addition to previously studied risk factors, we looked at other factors, not studied in the past, that may theoretically affect the risk of developing SG such has smoking history and fetal gender.
A cross-sectional study was conducted at a large private teaching hospital in Beirut, Lebanon, after obtaining institutional review board approval. All primiparas with singleton gestations delivering during a 6-month period (February-July 2005) were invited to participate in the study irrespective of gestational age at delivery. Women were identified through the Delivery Suite logbook on a daily basis.
After obtaining written consent to participate in the study, all eligible participants were assessed during the postpartum period before their discharge from the hospital using a 22-item data collection tool. Information was collected from the medical charts about socioeconomic status, gestational age at delivery, total weight gain during pregnancy, current weight, fetal gender, and birthweight. Socioeconomic status was determined based on the third party coverage with patients admitted on the expense of the Ministry of Health labelled as low socioeconomic status and those who had private insurance as high socioeconomic status. Patients were also asked about the use of creams for prevention of SG during pregnancy, smoking history, and family history of stretch marks. Family history of SG was considered positive if the woman's mother and/or sister had developed SG during her pregnancy. Skin type was determined by interview questions based on the Fitzpatrick classification, which is based on how often a person burns and how well they tan when exposed to the sun.
Presence of SG on the abdomen, thighs, and breasts was assessed by 1 of 3 researchers based on a scale developed and validated by the research team. The scale is based on the total surface area of the affected body part that is covered by SG: < 25% was rated as mild, 25-50% as moderate, and > 50% as severe. The scale provided a useful way to incorporate the number of SG as well as the width of SG covering the affected area. In return for their participation, women were given a packet of brochures about the postpartum period, breast feeding, and newborn care.
Assuming a 50% prevalence of SG, a total of 113 patients would be required to achieve a clinical significance of 15% at a power of 90% and a significance level of .05. The data were entered and analyzed using SPSS 13.0 (Chicago, IL). Two different outcomes were considered: (1) women with any SG (mild, moderate, or severe) on either the abdomen, thighs, or breasts versus those with no SG in any of those sites; and (2) women with moderate and/or severe SG in any of the 3 sites versus women who had either mild SG or none. Comparing the 2 outcomes across women's characteristics was done by either performing chi-square tests if the variables were categorical or t tests if the variables were continuous. Statistical significance was set at α = 0.05.
During the study period, 532 women delivered at the hospital. Of these, 163 were eligible for participation in the study. Forty-one women were discharged before it was possible to approach them, and 9 were not interested in participating in the study. One woman who was eligible was not approached because her infant was stillborn. Of the 112 women that were assessed, 1 had missing information for SG on the abdomen, thigh, and breast and another had no SG on abdomen and thigh but had missing information on the breast. Thus, both patients were excluded from the final analysis.
All eligible patients who were not formally assessed were compared to the women (n = 110) included in the study. No significant differences were found in maternal age, socioeconomic status, gestational age at delivery, fetal gender, or birthweight.
Of the 110 women enrolled in the study, 67 (61%) developed SG in at least 1 of the assessed sites. Fifty-three women (48%) developed SG on the abdomen, 27 (25%) developed SG on the breasts, and 27 (25%) developed SG on the thighs during their pregnancy. Figure 1 shows the percentage of women who developed SG at 1, 2, or all 3 sites. Of the abdominal striae, 17 (32%) were mild, 18 (34%) moderate, and another 18 (34%) severe. The severity of the striae on the breasts and thighs of the women were similar to each other with 19 (70%) reported as mild, 7 (26%) as moderate, and only 1 (4%) as severe. Figure 2 depicts the risk of developing moderate/severe SG by number of sites involved.
Most of the women in the study (93%) delivered at term. Seven percent delivered before 37 weeks of gestation with 1 woman delivering as early as 27 weeks. Eleven of the women (10%) delivered past 40 weeks of gestation. Maternal age ranged between 19 and 46 years. The majority of the women were between 24 and 34 years of age (77%), and the mean maternal age was 28 years. The weight gained during the pregnancy ranged from 3 to 33 kg with a mean of 14.4 kg. The BW ranged between 677 and 4115 g with a mean BW of 3143 g. Table 1 summarizes some antenatal and fetal characteristics in both groups. Women who developed SG were significantly younger and had gained significantly more weight during pregnancy compared to those who did not. BW and gestational age at delivery were strongly associated with risk of developing moderate/severe SG.
The predominant skin type in our population was Fitzpatrick III (41%) and IV (32%). Twenty percent of the women had skin types I/II, and only 8% had skin types V/VI. Most of the women (88%) were nonsmokers, and 45% were of a low socioeconomic status. Only 6% were current smokers and continued to smoke during their pregnancy, and 7% had smoked in the past but were no longer smoking at diagnosis of pregnancy. Sixty-seven women (61%) had used a cream or lotion during their pregnancy in an attempt to avoid the development of SG, and 19 (17%) had used more than 1 cream or lotion. There was a large variation in the types of creams used. The most commonly used products were cocoa butter (11%), baby oil (10%), and almond oil (5%). Sixty-five (59%) of the infants delivered were males and 47 (43%) of them were females. No relationship was noted between skin type, socioeconomic status, smoking, cream use, fetal gender, or family history and the risk of developing SG. However, women with family history of SG were more likely to develop moderate/severe SG than were those with no family history of SG. These findings are summarized in Table 2.
This study provides a clinical assessment of the prevalence of SG and associated risk factors in a cohort of racially homogeneous women at a single tertiary-care referral center. The evaluation was based on a new scoring system that was developed by the researchers. This is 1 of the few studies in which SG were quantified by clinical assessment rather than relying on the woman's own evaluation of her SG. To the best of our knowledge, our study is the only study that evaluated SG on the breasts and thighs and not merely the abdomen as in previously published studies in the literature. This was confirmed by a MEDLINE search from 1966 to March 2006, using the keywords "striae gravidarum," "stretch mark," and "pregnancy." Our finding that 24% of women developed SG on their thighs or breasts shows that these areas are also significantly affected by SG.
We found that the prevalence of SG is 60%, consistent with previous reported figures. SG have a predilection to the abdomen, the site of involvement in 47% of the women; 24% had SG on the thighs and/or breasts. The correlation we identified between weight gain during pregnancy and BW and the development of SG are consistent with the findings by Davey. Although gestational age at delivery and BW were similar in those who developed SG and those who did not, both BW and gestational age at delivery were significantly larger for those who developed more severe SG. These factors are all probably interrelated and are related to some extent to the degree of stretching of the skin. Thomas et al noted that women with higher BMI and larger babies have more stretch marks. Similar to Thomas et al, we found that younger women were more likely to develop SG, though this finding is not consistent with other studies.
Although we did not find that family history of SG was significantly correlated with the development of SG, we did find that women with a positive family history of SG were more likely to develop moderate/severe SG, suggesting that genetic factors do play a role in the development of SG. The other study that assessed the role of family history found a strong relationship between family history of SG and the risk of their development. However, that study was a voluntary and self-administered questionnaire that did not take into account factors like maternal age, number of births, and time elapsed since the delivery date.
The literature on the association between race and risk of the development of SG is conflicting with some researchers finding that nonwhites were more likely to develop SG than white women, while others found that women with lighter skin were more likely to develop SG. Unfortunately, our study population was too racially homogeneous to determine any differences with regard to SG risk related to skin type.
We felt it would be worthwhile to investigate the effects of smoking and fetal gender, mostly due to the theoretical effect of tobacco and fetal sex hormones on SG development since both are known to affect connective tissue properties. However, the proportion of smokers in our cohort was too small to show any difference with regard to SG development. We found that fetal gender did not correlate with SG development.
A large proportion of our population was using 1 or more creams/lotions in an attempt to prevent the development of SG; however, we found no correlation between cream use and SG development. In 1972, Davey found that the use of oil and massage significantly reduced the likelihood of SG development. The nature of our study makes it difficult to draw any conclusions regarding a possible advantage or role of some of the creams utilized. It is very likely that users of lotions/creams started to apply these topical treatments after they noted that they were developing SG in the hope of minimizing their appearance.
The appearance of SG may be influenced by population genetics. If that is the case, the significance level (α) should probably be set to lower levels which may affect our results. Genetic studies to establish whether such a linkage exists have already been proposed to clarify this issue.
Pregnant women often request information regarding their risks of developing SG and means to prevent their appearance during their prenatal visits. Our findings can help physicians answer some of these questions when counseling patients about their risk of developing SG. Although some of the factors associated with SG are not modifiable (ie, age and family history), other factors such as weight gain during pregnancy are. Future research should focus on preventive methods that may reduce the likelihood of SG development. More specifically the prophylactic use of creams and lotions should be further investigated to determine once and for all if these treatments have any benefit. Sponsored Document Sponsored Document Sponsored Document Sponsored Document Sponsored Document Sponsored Document Osman et al. Page 8
* Statistically significant.
†
Current Smoking was labeled 'yes' if the patient smoked during pregnancy, 'no' if she has never smoked, and 'stopped' if smoking was stopped before pregnancy.
Published as: Am J Obstet Gynecol. 2007 January ; 196(1): 62.e1-62.e5.Sponsored DocumentSponsored DocumentSponsored Document
Published as: Am J Obstet Gynecol. 2007 January ; 196(1): 62.e1-62.e5.
The authors would like to acknowledge
The endometrium has a remarkable capacity for efficient repair; however, factors involved remain undefined. Premenstrual progesterone withdrawal leads to increased prostaglandin (PG) production and local hypoxia. Here we determined human endometrial expression of interleukin-8 (IL-8) and the roles of PGE 2 and hypoxia in its regulation. Endometrial biopsy specimens (n = 51) were collected. Endometrial cells and explants were exposed to 100 nmol/L of PGE 2 or 0.5% O 2 . The endometrial IL-8 concentration peaked during menstruation (P < 0.001) and had a significant proangiogenic effect. IL-8 was increased by PGE 2 and hypoxia in secretory but not proliferative explants, which suggests that exposure to progesterone is essential. In vitro progesterone withdrawal induced significant IL-8 up-regulation in proliferative explants primed with progestins, but only in the presence of hypoxia. Epithelial cells treated simultaneously with PGE 2 and hypoxia demonstrated synergistic increases in IL-8. Inhibition of HIF-1 by short hairpin RNA abolished hypoxic IL-8 induction, and inhibition of NF-κB by an adenoviral dominant negative inhibitor decreased PGE 2 -induced IL-8 expression (P > 0.05). Increased menstrual IL-8 is consistent with a role in repair. Progesterone withdrawal, hypoxia, and PGE 2 regulate endometrial IL-8 by acting via HIF-1 and NF-κB. Hence, progesterone withdrawal may activate two distinct pathways to initiate endometrial repair.
Menstruation exhibits many of the classic hallmarks of inflammation. The withdrawal of progesterone in the late secretory phase of the cycle triggers a cascade of inflammatory mediators, leading to a dramatic influx of leukocytes into the premenstrual endometrium. After shedding, the human endometrium exhibits a remarkable and immediate regenerative capacity. This cyclical injury and repair is tightly controlled and, unlike resolution of inflammation at other sites in the body, does not involve loss of function or scarring. However, the precise local mechanisms involved in this efficient repair have not yet been fully elucidated. Aberrations may lead to menstrual disorders including heavy menstrual
bleeding and dysmenorrhea. Delineation of the physiologic processes of the endometrium could result in new therapeutic targets for these common debilitating conditions. In addition, the efficient endometrial model may provide an informative comparator for other tissue sites associated with problematic scarring or persistent inflammation.
Withdrawal of progesterone occurs in the late secretory endometrium as the corpus luteum regresses. Progesterone withdrawal leads to up-regulation of endometrial cyclooxygenase-2 (COX-2) and subsequent increased levels of prostaglandins (PGs), namely, PGE 2 and PGF 2α . PGF 2α induces myometrial contractions and vasoconstriction of the endometrial spiral arterioles. Consequently, it is believed that there is an episode of transient hypoxia in the uppermost endometrial zones. The existence of hypoxia was confirmed in a murine model of menstruation using pimonidazole, a marker of pO 2 less than 10 mm Hg. The luminal portion of the endometrial functional layer was demonstrated to be intensely hypoxic during simulated menstruation, with negligible detection of pimonidazole by day 5. It was hypothesized that PGF 2α along with other endometrial vasoconstrictors induces hypoxic conditions in the human perimenstrual endometrium to increase repair gene expression. The role of the other major prostaglandin present during the premenstrual phase, PGE 2, is not fully understood. It was proposed, therefore, that PGE 2 may also independently increase expression of genes responsible for endometrial repair.
Interleukin-8 (IL-8, CXCL8) is a CXC chemokine, best known for its role as a potent chemoattractant for neutrophils and T cells. In addition, it has mitogenic properties and a key role in angiogenesis in vivo. These processes are fundamental for endometrial shedding and repair. The present study demonstrated significant changes in IL-8 mRNA and protein expression during the menstrual cycle, with maximal expression at menstruation. Concentrations of IL-8 secreted by menstrual endometrium exhibited significantly greater angiogenic potential in vitro than did concentrations secreted by mid-secretory endometrium. IL-8 expression is up-regulated in endometrial epithelial cells by hypoxic conditions and by PGE 2 , with a synergistic increase observed in the presence of both factors. An in vitro model of progesterone withdrawal also increased IL-8 expression in human endometrial tissue, but only with the addition of hypoxic conditions. The presence of indomethacin, a COX enzyme inhibitor, attenuated the increase in IL-8 expression in this model. These observations suggest a role for progesterone withdrawal in the initiation of endometrial repair and indicate that subsequent hypoxia and PGE 2 are necessary for increased expression of IL-8, an angiogenic factor with a putative role in the repair process.
Human endometrial biopsy specimens were collected from women undergoing hysterectomy or investigation in the gynecologic outpatient setting (n = 51). Ethical approval was obtained from the Lothian Research Ethics Committee, and written informed consent was obtained from all participants before tissue collection. Participants were aged 31 to 52 years (median, 41 years; mean, 41 years). All women reported regular menstrual cycles (duration, 21 to 35 days) and had not taken exogenous hormones or used an intrauterine device during the 3 months before endometrial biopsy. Women with known uterine disease such as large myomas (>3 cm) and endometriosis were excluded. Endometrial biopsy specimens were collected using an endometrial suction curette (Pipelle; Laboratoire CCD, Paris, France). Immediately after collection, tissue was divided and i) placed in RNA stabilization solution (RNA Later; Ambion (Europe) Ltd., Warrington, UK), ii) stored at -70°C for RNA extraction, iii) fixed in neutral buffered formalin for wax embedding or iv) placed in PBS for in vitro culture. The specimens were dated according to the criteria of Noyes et al based on histologic appearance, which was consistent with the participants' reported last menstrual period. In addition, serum samples were collected from each woman at biopsy to determine circulating serum progesterone and estradiol concentrations, and were consistent for both last menstrual period and histologic assessment. For analysis, biopsy specimens were classified as proliferative, early secretory, mid secretory, late secretory, or menstrual (Table 1). Seven women consented to undergo a second endometrial biopsy, and returned for this procedure three to six months after insertion of the levonorgestrel-releasing intrauterine system (LNG-IUS) for treatment of subjective report of heavy menstrual bleeding.
Endometrial biopsy specimens (secretory phase, n = 7; proliferative phase, n = 3) were divided into three equal explants and incubated for at least 16 hours on raised platforms in 24-well plates just covered with serum-free RPMI 1640 medium plus 50 μg/ml of penicillin, 50 μg/ml of streptomycin, and 5 μg/ml of gentamicin (all from Sigma Aldrich, St. Louis, MO), and 8.4 μmol/L of indomethacin. The next day, two explants were treated with vehicle under normoxic conditions, 1 with 21% O 2 , 5% CO 2, and 37°C, and one with 100 nmol/L of PGE 2 . The last explant was placed in a sealed hypoxic chamber (Coy Laboratory Products Inc., Grass Lake, MI) set at 0.5%O 2, 5% CO 2, and 37°C for 24 hours.
Five endometrial biopsy specimens from the proliferative phase were divided into 8 equalsized explants and placed on raised platforms in four wells of 2 × 24-well plates. All explants were treated with 1 μmol/L of medroxyprogesterone acetate (MPA) for 24 hours. Explants were then treated with either 1 μmol/L of MPA plus vehicle, 1 μmol/L of MPA plus 8.4 μmol/L of indomethacin (a COX enzyme inhibitor), 1 μmol/L of MPA and 1 μmol/ L of RU486 (a progesterone-receptor antagonist) plus vehicle, or 1 μmol/L of MPA and RU486 plus 8.4 μmol/L of indomethacin. One plate was placed in normoxic conditions, and the other in hypoxic conditions, for 48 hours.
Human Ishikawa endometrial adenocarcinoma cells (European Collection of Cell Cultures, Centre for Applied Microbiology, Wiltshire, UK) stably expressing the EP2 receptor (EP2S) were maintained in Dulbecco modified Eagle medium nutrient mixture F-12 with glutamax-1 and pyridoxine, supplemented with 10% fetal calf serum, 1% antibiotic (stock 500 IU/ml of penicillin and 500 μg/ml of streptomycin), and 200 μg/ml of G418 at 37°C. Primary human endometrial stromal cells were isolated from mid-secretory endometrial tissue (n = 3) via enzymatic digestion as previously described, and were maintained in RPMI 1640 medium plus 50 μg/ml of penicillin, 50 μg/ml of streptomycin, and 5 μg/ml of gentamicin (all from Sigma Aldrich).
Approximately 4 × 10 5 EP2S or 3 × 10 5 human endometrial stromal cells were seeded in 6well plates. The following day, cells were incubated for at least 16 hours in serum-free culture medium containing antibiotics and 8.4 μmol/L of indomethacin. Cells were then treated with either vehicle or 100 nmol/L of PGE 2 and placed at 37°C, 21% O 2 , and 5% CO 2 for 2, 4, 8, 24, and 48 hours or placed in hypoxic conditions (0.5%O 2 and 5% CO 2 ) in a sealed chamber (Coy Laboratory Products Inc.) for the same amount of time. Alternatively, EP2S cells were pretreated with vehicle or 5 nmol/L of echinomycin (a specific inhibitor of HIF-1 DNA binding activity). After 1 hour, cells were stimulated for 6 hours with vehicle, 100 nmol/L of PGE 2 with or without 5 nmol/L of echinomycin,or hypoxia with or without 5 nmol/L of echinomycin. A short-hairpin RNA (shRNA) sequence against human HIF-1α and scrambled control oligonucleotide (TIB MOLBIOL) were donated by Prof. T. Cramer (Charité-Universitätsmedizin Berlin, Berlin, Germany). A 19-nucleotide sequence derived from human HIF-1α mRNA (U22431; bp 1470 to 1489) was used and was termed HIF-1α/ shRNA. Cells were transiently transfected with lentivirus at a multiplicity of infection of 10 for 24 hours. Cells were incubated in serum-free medium overnight before treatment with 100 nmol/L of PGE 2 or placed in the hypoxic chamber for 8 hours. Cells were washed with PBS and harvested, and RNA or protein was extracted for PCR or Western blot analysis. To determine the role of NF-κB in IL-8 up-regulation, EP2S cells were seeded at a density of 1 × 10 5 . The following day, cells were infected with an adenovirus containing a dominantnegative Iκ-Bα mutant, which maintains NF-κB in a cytoplasmic location, or control adenovirus (Ad-d1703) at a total multiplicity of infection of 50 for 8 hours. Ad-d1703 and Ad-Iκ-Bα have been described previously. Cells were serum-starved with 8.4 μmol/L of indomethacin for at least 16 hours before treatment with 100 nmol/L of PGE 2 or hypoxic conditions for 6 hours.
Protein was extracted from endometrial cells with a cytoplasmic protein lysis buffer (10 mmol/L of HEPES, pH 7.8), 10 nmol/L of KCl, 2 mmol/L of MgCl 2, 1 mmol/L of dithiothreitol, 0.1 mmol/L of EDTA, and 10% Nonident P-40) containing protease inhibitors (Complete Mini Protease Inhibitor Cocktail; Roche Diagnostics, Ltd., Lewes, UK). After centrifugation at 13,000 rpm for 1 minute at 4°C, the cytoplasmic fraction supernatant was removed and stored at -80°C. The nuclear fraction was extracted using a nuclear protein lysis buffer (50 mmol/L of HEPES [pH 7.8], 50 nmol/L of KCl, 300 mmol/L of NaCl, 0.1 mmol/L of EDTA, 1 mmol/L of dithiothreitol, and 10% glycerol) containing protease inhibitors (Roche Diagnostics, Ltd), followed by agitation for 20 minutes at 4°C and centrifugation at 13,000 rpm for 5 minutes at 4°C. The nuclear fraction supernatant was removed and stored at -80°C. Protein content was determined using protein assay kits (Bio-Rad; Hemel Hempstead, UK).
For detection of HIF-1α and β-actin, 10 μg of nuclear protein was resuspended in a 2:1 ratio with Laemmli buffer (125 mmol/L Tris-HCl [pH 6.8], 4% SDS, 5% 2-mercaptoethanol, 20% glycerol, and 0.05% bromophenol blue) and denatured for 5 minutes at 90°C. Proteins were separated on 4% to 12% Bis-Tris gels (NuPAGE Novex; Invitrogen Corp., Carlsbad, CA) and transferred onto polyvinylidene difluoride membrane (Millipore Corp., Billerica, MA). Membranes were blocked overnight in 5% milk solution in Tris-buffered saline solution and Tween 20 (50 mmol/L of Tris HCl, 150 mmol/L of NaCl, and 0.05% v/v of Tween 20). After washing with Tris-buffered saline solution and Tween 20, the membranes were incubated with mouse monoclonal anti-HIF-1α antibody (BD Biosciences, Oxford, UK) (1:250) and rabbit polyclonal anti-β-actin (Abcam, Cambridge, UK) (1:5000). After washing, the membrane was incubated with horseradish peroxidase-conjugated goat antimouse IgG (DAKO Corp, Carpinteria, CA) or horseradish peroxidase-conjugated mouse anti-rabbit IgG (Sigma Aldrich) at 1:20,000. The chemiluminescent horseradish perioxidase substrate (Immobilon; Millipore Corp.) was used for immunoreactive protein detection according to the manufacturer's instructions.
Expression of IL-8 mRNA in endometrial tissue and Ishikawa cells was determined using quantitative RT-PCR (Taqman) analysis. Total RNA from cells and endometrial biopsy specimens was extracted using a kit (RNeasy Mini Kit; Qiagen Ltd, Sussex, UK) according to the manufacturer's instructions. Samples were treated for DNA contamination via DNA digestion during RNA purification. After extraction, RNA was quantified using a spectrophotometer (NanoDrop 1000, version 3.7; ThermoScientific, Wilmington, DE) and stored at -80°C. Quality of the RNA was assessed using a bioanalyzer (Agilent 2100 Bioanalyser System) in combination with RNA 6000 nano chips (Agilent Technologies, Palo Alto, CA).
RNA samples were reverse transcribed using 5.5 mmol/L of MgCl 2 , 0.5 mmol/L each of deoxynucleotide triphosphates, 2.5 μmol/L of random hexamers, 0.4 U/μlL of RNA inhibitor, and 1.25 U/μL of multiscribe reverse transcriptase (all from PE Biosystems, Warrington, UK). The mix was aliquoted into individual tubes, and 200 to 400 ng of RNA was added. A tube with no reverse transcriptase and a further tube with water were included to control for DNA contamination. After mixing, samples were incubated for 20 minutes at 25°C, 60 minutes at 42°C, and 5 minutes at 95°C. cDNA samples were subsequently stored at -20°C.
To measure cDNA expression, a reaction mix was prepared containing Taqman buffer (5.5 mmol/L of MgCl 2 , 200 μmol/L of deoxyadenosine triphosphate, 200 μmol/L of deoxycytidine, 200 μmol/L of deoxyguanosine, and 400 μmol/L of deoxyuridine triphosphate), ribosomal 18S primers and probe (Applied Biosystems, Warrington, UK), and specific forward and reverse primers and probe for IL-8 and EP2: IL-8 forward primer, 5′-CTGGCCGTGGCTCTCTTG-3′; reverse primer, 5′-TTAGCACTCCTTGGAAAACTG-3′; and probe, 5′-CCTTCCTGATTTCTGCAGCTCTGTGTGAA-3′; and EP2 forward primer, 5′-TGAAGTTGCAGGCGAGCA-3′; reverse primer, 5′-GACCGCTTACCTGCAGCT-3′; and probe, 5′-CCACCCTGCTGCTGCTGCTTCT-3′. After mixing, 36-μL aliquots were placed in separate tubes, and 1.5 μL of cDNA was added. Into one aliquot, 1.5 μL of water was added as a no template control. Triplicate 12-μL samples were placed in a PCR plate. PCR was performed using ABI Prism 7900 (Applied Biosystems). Data were analyzed and processed using Sequence Detector version 2.3 (PE Biosystems). Expression of target mRNA was normalized to RNA loading for each sample using the 18S ribosomal RNA as an internal standard.
Endometrial tissue from women at each stage of the menstrual cycle was collected in PBS (n = 20), weighed, and incubated for 24 hours on raised platforms in 1 ml of serum-free RPMI 1640 medium with 50 μg/ml of penicillin, 50 μg/ml of streptomycin, and 5 μg/ml of gentamicin (all from Sigma Aldrich). IL-8 protein secretion into the culture medium by EP2S cells and endometrial biopsy specimens after 24 hours was quantified using an in-house enzyme-linked immunosorbent assay as described previously. A mouse monoclonal anti-human IL-8 capture antibody and a biotinylated polyclonal goat anti-human IL-8 detection antibody were used (R&D Systems, Oxford, UK). Protein concentrations in the conditioned medium were normalized to tissue weight.
IL-8 was immunolocalized in endometrial tissue sections as previously described. In brief, slides were dewaxed and rehydrated before antigen retrieval in 0.01mmol/L of sodium citrate on high power in a pressure cooker for 5 minutes. Primary antibody (rabbit polyclonal, 1:100) was added overnight at 4°C. After incubation with secondary antibody (goat anti-rabbit, 1:200) and avidin biotin peroxidise complex (ABC Elite; Vector Laboratories, Peterborough, UK), staining was detected with liquid biaminobenzidine (DAB kit; Zymed Laboratories, Inc., South San Francisco, CA). Localization and intensity of immunostaining were evaluated blindly by two independent observers using a previously validated semiquantitative scoring system (J.A.M.). Intensity was graded using a three-point scale (0 = no staining, 1 = mild staining, and 2 = strong staining). The percentage of cells stained at each of these intensities was assessed in each cellular compartment. A value was derived for each compartment using the sum of these percentages after multiplication by the intensity of staining.
Matrigel, 100 μL (BD Biosciences, Bedford, MA), was added in each well of a 48-well plate and allowed to polymerize for 1 hour at 37°C. Human umbilical vascular endothelial cells were seeded at a density of 2 × 10 4 in 200 μL of EBM-2 medium (Lonza, Walkersville, MD) supplemented with GA1000 and ascorbic acid SingleQuots (Lonza). Cells were then treated with 250 μL of culture supernatant from menstrual and mid-secretory tissue explants incubated in vitro for 24 hours (40 mg of tissue per milliliter of RPMI medium) or 0.5 or 20 ng of recombinant human IL-8 (R&D Systems) in 250 μL of medium. Each dose of IL-8 was assessed in triplicate in three separate experiments. Capillary tube formations were visualized after 8 hours. Images were captured in the same position in each well using an inverted microscope at ×5 magnification. Branch points of the formed tubes were counted by an observer (J.A.M.) blinded to the sample origin, and an average of the replicates was determined after unblinding.
For mRNA expression in explants and cell culture, results are given as fold increase where relative expression of mRNA in cells treated with PGE 2 was divided by the relative expression in vehicle-treated cells. Data are given as mean (SEM). Significant difference was determined using one-way analysis of variance of delta cycle threshold values using Tukey posttest analysis. For endometrial biopsy specimens from across the menstrual cycle, results are given as quantity relative to a comparator, a sample of RNA from the liver. Significant difference was determined using the Kruskal-Wallis nonparametric test with the Dunn multiple comparison posttest (Instat; GraphPad Software, Inc., San Diego, CA).
IL-8 mRNA was present at low levels in endometrium from the proliferative, early secretory, and mid secretory stages of the menstrual cycle. A nonsignificant increase in IL-8 mRNA expression was observed in the late secretory phase. By the menstrual stage, IL-8 concentrations had increased significantly compared with endometrium from the proliferative (P < 0.01), early secretory (P < 0.001), and mid secretory (P < 0.05) stages (Figure 1A. The amount of IL-8 protein secreted from endometrial biopsy specimens cultured in vitro for 24 hours demonstrated a similar pattern (Figure 1B). Endometrium from the menstrual stage secreted significantly higher concentrations of IL-8 protein than did tissue from the early and mid secretory phases (P < 0.05). There was no significant decrease in IL-8 protein between the menstrual and proliferative stages. Immunolocalization of IL-8 demonstrated positive cytoplasmic staining in glandular epithelial, surface epithelial, stromal, and perivascular cells in endometrium from the menstrual phase of the cycle (Figure 1, C and D). In contrast, during the mid secretory phase of the cycle, stromal staining was negligible and glandular epithelial cells were only faintly positive (Figure 1, E and F). Semiquantitative scoring of staining intensity revealed that the strongest staining was in the glandular epithelial and perivascular cells (Figure 1G). IL-8 perivascular staining was significantly increased during the menstrual phase of the cycle when compared with the proliferative (P < 0.05), early secretory (P < 0.01), and mid secretory (P < 0.05) stages (Figure 1G). There was a nonsignificant increase in IL-8 staining in glandular epithelial and stromal cells during the menstrual phase (Figure 1G).
To assess the angiogenic potential of IL-8 produced by the endometrium, branching of human umbilical vascular endothelial cells (HUVECs) was quantified after various treatments. Compared with cells treated with unconditioned medium, cells treated with conditioned medium from menstrual tissue incubated for 24 hours in vitro demonstrated a significant increase in HUVEC capillary branch point formation (Figure 2A). No significant increase in angiogenesis was observed with conditioned medium from mid secretory phase explants. These endometrial explants are likely to produce several angiogenic factors. To assess the contribution of IL-8 alone, HUVECs were also treated with recombinant human IL-8. The mean (SEM; median) amount of IL-8 secreted by menstrual endometrial explants was 18.94 (7.57; 19.4) ng. Mid secretory endometrium secreted the lowest levels of IL-8: 0.53 (0.13; 0.44) ng. Therefore, HUVECs were treated with control medium, 20 ng or 0.5 ng of human recombinant IL-8. Compared with cells treated with 0.5 ng of IL-8 or control medium, treatment of HUVECs with 20 ng of IL-8 resulted in a significantly higher number of capillary tube branch points (Figure 2B). Mid-secretory levels of IL-8 had no significant effect on branch points when compared with control medium.
To investigate the regulation of endometrial IL-8, human endometrial explants were cultured for 24 hours with vehicle, 100 nmol/L of PGE 2 , or hypoxic conditions. Secretory endometrium from seven women demonstrated a nonsignificant increase in IL-8 expression with PGE 2 treatment under normoxic conditions. Culture of endometrial explants under hypoxic conditions significantly elevated IL-8 mRNA expression (P < 0.05) (Figure 3A). In contrast, neither treatment induced up-regulation of IL-8 in endometrium from the proliferative phase (n = 3) (Figure 3B). This suggests that previous exposure to progesterone is essential for up-regulation of IL-8 by PGE 2 and hypoxia. There was no significant difference in EP2 receptor mRNA expression in response to PGE 2 or hypoxia between explants from the proliferative and secretory phases of the cycle (data not shown).
To establish whether progesterone withdrawal induces IL-8 mRNA expression, proliferative endometrial biopsy specimens were divided into 8 explants (n = 5). All explants were treated with MPA for 24 hours. After progesterone exposure, progesterone withdrawal was simulated in four of the explants by co-treating with RU486, a progesterone-receptor antagonist. Progesterone withdrawal under normoxic conditions did not significantly upregulate IL-8 mRNA expression (Figure 4).
It was postulated that in vivo, progesterone withdrawal in the late secretory phase induces synthesis of prostaglandins and constriction of spiral arterioles, resulting in an episode of transient hypoxia. Therefore, to mimic the in vivo condition more accurately, two endometrial explants were exposed to hypoxic conditions at simulated progesterone withdrawal. Addition of hypoxic conditions induced significant induction of IL-8 mRNA expression 48 hours after progesterone withdrawal (P < 0.05) (Figure 4A).
To assess the contribution of prostaglandins after progesterone withdrawal, explants were concomitantly treated with MPA (progestogen), RU486 (progesterone-receptor antagonist), and indomethacin (a COX enzyme inhibitor). Addition of indomethacin attenuated upregulation of IL-8 mRNA after progesterone withdrawal under hypoxic conditions (Figure 4A).
To further investigate the role of progesterone and hypoxia in regulating endometrial IL-8 expression, endometrial biopsy specimens from seven women obtained before and 3 to 6 months after LNG-IUS insertion were examined. The LNG-IUS markedly down-regulated the progesterone receptor in all components of the endometrium, resulting in a human model of progesterone deficiency. At comparison of endometrium obtained during the proliferative, early secretory, and mid secretory stages with paired samples obtained after 3to 6-month exposure to LNG-IUS (n = 7), significant up-regulation of IL-8 mRNA expression was observed after LNG-IUS exposure (P < 0.05) (Figure 4B). This increase in endometrial IL-8 after LNG-IUS insertion was also identified at the protein level. Increased IL-8 immunohistochemical staining was visible in the decidualized stromal cells present after LNG-IUS exposure (Figure 4, C and D). Endometrial biopsy specimens obtained during the late secretory and menstrual phases (n = 2) demonstrated no significant change in IL-8 mRNA expression on exposure to the LNG-IUS (data not shown). This suggests that endometrium already exposed to progesterone withdrawal in vivo has no further capacity for IL-8 induction on insertion of LNG-IUS.
To delineate the mechanisms by which PGE 2 and hypoxia induce IL-8 expression, an Ishikawa endometrial epithelial cell line stably expressing the EP2 receptor was used. This cell line was used to mimic primary endometrial epithelial cells, which express receptors for PGE 2 . Cells were exposed to treatment with vehicle or 100 nmol/L of PGE 2 for up to 48 hours under normoxic and hypoxic conditions. Treatment with PGE 2 under normoxic conditions (Figure 5A) demonstrated a significant increase in IL-8 mRNA expression, with maximal up-regulation after 8 hours (P < 0.01). Hypoxic conditions also significantly increased IL-8 mRNA expression (Figure 5B) but exhibited a more delayed induction, reaching maximum up-regulation after 8 to 24 hours (P < 0.01). When cells were exposed to both PGE 2 and hypoxic conditions for 24 hours (Figure 5C), there was a synergistic increase in IL-8 mRNA expression that was significantly greater than with treatment with PGE 2 in normoxia (P < 0.05) or hypoxia (P < 0.05) alone. Levels of secreted IL-8 protein demonstrated a similar pattern, with a synergistic increase in IL-8 protein secretion with PGE 2 treatment under hypoxic conditions (Figure 5D). In contrast, in human endometrial stromal cells, hypoxic conditions had no significant effect on IL-8 mRNA expression or protein levels at any time examined (data not shown). Treatment with 100 nmol/L of PGE 2 resulted in a significant increase in IL-8 mRNA expression after 48 hours (P < 0.05) and a nonsignificant increase in secreted protein levels at the same time point (data not shown).
To determine the role of NF-κB in up-regulation of IL-8 in the endometrium, cells were infected with a dominant-negative inhibitor of NF-κB (Ad-Iκ-Bα) and cultured for 6 hours either in the presence of vehicle or PGE 2 or under hypoxic conditions. Infection of cells with Ad-Iκ-Bα resulted in significant reduction of PGE 2 -induced IL-8 mRNA expression (P < 0.05) when compared with uninfected cells or cells infected with control Ad-d1730 (Figure 6B). Hypoxia-induced IL-8 mRNA expression was not significantly affected by inhibition of NF-κB (Figure 6C).
Echinomycin is a small molecule that inhibits the DNA binding of hypoxia-inducible factor (HIF) to the hypoxic response element sequence but does not affect AP-1 or NF-κB binding (Figure 6D-F). Cells concomitantly treated with PGE 2 and 5 nmol/L of echinomycin demonstrated a significant (P < 0.05) but not absolute reduction in IL-8 mRNA expression when compared with cells treated with 100 nmol/L of PGE 2 alone (Figure 6E). Hypoxiainduced IL-8 mRNA expression was abolished when cells were concomitantly treated with 5 nmol/L of echinomycin (P < 0.05) (Figure 6F).
HIF-1α knockdown was confirmed at Western blot analysis (see Supplemental Figure S1A at http://ajp.amjpathol.org). There was a marked decrease in HIF-1α protein in cells transfected with shRNA against HIF-1α before hypoxic incubation versus untransfected cells or those transfected with a scrambled shRNA sequence. Specificity of the knockdown was confirmed by examination of lamin A/C mRNA expression, which was not significantly different with transfection of any construct (Figure S1B). IL-8 expression was increased with PGE 2 or hypoxic incubation. Transfection of cells with a scrambled sequence did not significantly change IL-8 mRNA expression. In agreement with pharmacologic inhibition of HIF-1α binding, the hypoxic increase in IL-8 was significantly abrogated when HIF-1α was silenced before treatment (P < 0.05) (Figure 6H). PGE 2 -induced IL-8 mRNA expression was nonsignificantly decreased when HIF-1α was silenced, when compared with untransfected cells.
In the present study, significant menstrual up-regulation of endometrial IL-8 mRNA and protein was observed. The timing of this elevation in IL-8 expression is consistent with the onset of endometrial repair. The data support the hypothesis that progesterone withdrawal followed by increased PGE 2 and hypoxic conditions up-regulates endometrial repair factor expression. Furthermore, NF-κB and HIF-1 are two transcription factors that have a role in the induction of IL-8 for menstrual repair. Cross-talk between these factors presents a mechanism for the synergistic increase in IL-8 observed when PGE 2 and hypoxia are present simultaneously, as occurs in the perimenstrual endometrium.
Previous studies have found an increase in IL-8 mRNA and protein expression during the late secretory phase of the menstrual cycle. However, those studies did not examine tissue from the menstrual phase; thus, the maximal increase in IL-8 during this stage was not demonstrated. The finding of significant elevation of IL-8 protein during menstruation is in agreement with the findings of Jones et al, who reported undetectable levels of IL-8 mRNA during the menstrual cycle until a dramatic up-regulation at menstruation. As endometrial repair has been shown microscopically to commence on cycle day 2, the finding of maximal IL-8 levels during menstruation is consistent with a role in endometrial repair. A recent study of the menstrual endometrium revealed an increase in genes associated with extracellular matrix biosynthesis in stromal cells from the functional layer when compared with those from the basal layer. Overexpression of these genes, which includes IL8 (>4-fold increase), suggests that fragments of the functional layer of endometrium participate in endometrial repair. IL-8 is a potent chemokine, and is reported to control the migration and activation of leukocytes during menstruation. A host of chemokines are present in the premenstrual endometrium, including monocyte chemotactic protein-3, eotaxin, fractaline, and 6Ckine (chemokine with 6 cysteines). By using a gene array approach and validation with RT-PCR, Jones et al demonstrated that of all of the chemokines assessed, only IL8 was significantly increased in menstrual phase endometrium. Inflammatory cells produce and secrete proteases, such as matrix metalloproteinases, that have the ability to break down the extracellular matrix. Therefore, the maximal expression of IL-8 at menstruation described herein is consistent with a role in chemotaxis and inflammatory cell accumulation in the endometrium, key events in the initiation of menstruation. In addition, leukocytes form an essential component of the endometrial repair process. Neutrophil depletion using the antibody RB6 8C5 markedly delayed endometrial repair in the mouse model of menstruation. In addition to its role in neutrophil chemotaxis, IL-8 has important angiogenic properties and induces mitogenesis of vascular smooth muscle cells. IL-8 interacts with two chemokine receptors, CXCR1 and CXCR2. Both are expressed in the endometrium throughout the menstrual cycle. Therefore, it was postulated that IL-8 has a functional role in human endometrial angiogenesis and repair. The present study demonstrated that menstrual phase endometrial explants have the ability to produce factors with significant angiogenic potential. In addition, the elevated levels of IL-8 present during menstruation have increased angiogenic potential when compared with levels secreted during the mid secretory phase. Numerous angiogenic factors are present in the endometrium during menstruation, including vascular endothelial growth factor, the angiopoietins, and plateletderived growth factor. All likely have a role in vascular proliferation and differentiation, enabling rapid repair of damaged blood vessels. An element of functional redundancy of these factors is to be expected to ensure efficient endometrial repair. Although IL-8 may not be essential for angiogenesis during endometrial repair, the IL-8 protein levels present during menstruation are sufficient for an active contribution to this physiologic process.
Postmenstrual repair was traditionally considered estrogen-dependent. However, using scanning electron microscopy, Ludwig and Spornitz demonstrated that epithelial cell proliferation and migration commenced on day 2 of the menstrual cycle and that full coverage of the uterine lumen was achieved by day 6. Because estrogen levels remain low throughout the menstrual phase, these observations suggest that initiation of repair may be estrogen-independent. The murine model of menstruation also supports the hypothesis that estrogen is not essential for endometrial repair. Ovariectomized mice were maintained on a soy-free diet and treated with an aromatase inhibitor to remove all estrogenic influence. When assessed morphologically, no significant difference in the rate of endometrial repair was observed in the complete absence of estrogen. Notwithstanding the limitations of the mouse model of simulated menstruation, these results support findings in the human endometrium that suggest that estrogen is not necessary for repair, although it may contribute to the process. Therefore, it was postulated that progesterone withdrawal rather than an increase in estradiol is the stimulus for endometrial repair factor expression.
Progesterone withdrawal in vivo causes significant up-regulation of endometrial IL-8 mRNA expression after 48 hours. However, the mechanisms by which progesterone withdrawal manifests this effect remain undefined. Progesterone withdrawal during the late secretory phase of the menstrual cycle results in up-regulation of COX-2, an enzyme responsible for prostaglandin synthesis. PGF 2α is a potent vasoconstrictor. Premenstrual increases in PGF 2α and other vasoconstrictors such as endothelin-1 result in constriction of spiral arterioles. This causes a transient episode of hypoxia in the functional layer of the endometrium (Figure 7). The hypothesis that hypoxia exists during the perimenstrual phase was derived from classic experiments in the rhesus monkey. Direct observation of changes in intraocular endometrial implants demonstrated vasoconstriction of the spiral arterioles and a decrease in blood flow. Hypoxia has also been demonstrated in the mouse model of menstruation using pimonidazole. Furthermore, although some controversy remains about the presence of hypoxia in the human endometrium, late secretory and menstrual endometrium exhibits positive nuclear immunohistochemical staining for HIF-1α and CAIX, two markers of hypoxia. Therefore, it is proposed that hypoxia is involved in the initiation of postmenstrual repair factor expression after progesterone withdrawal.
Herein, it has been demonstrated that PGE 2 and hypoxia independently up-regulate IL-8 mRNA expression in endometrial epithelial cells and in endometrial explants that have had previous progesterone exposure. Endometrial tissue from the proliferative stage, that is, with no significant in vivo progesterone exposure, demonstrated no such increase in IL-8 expression with PGE 2 or hypoxia. There was no significant difference in EP2 mRNA expression between explants from the proliferative and secretory phases of the cycle. In addition, previously published data on the endometrial expression of the EP2 receptor demonstrated no significant variation across the menstrual cycle. These data suggest that the variation observed in explants from various phases of the cycle in response to PGE 2 and hypoxia is not due to differing levels of EP2 receptor expression. When proliferative explants were subjected to an in vitro model of progesterone withdrawal using the progesterone-receptor antagonist mifepristone, there was no up-regulation of IL-8 under normoxic conditions. Under in vitro conditions, endometrial architecture is disturbed, and up-regulation of COX-2 and subsequent synthesis of PGF 2α are unlikely to result in vasoconstriction and local tissue hypoxia. To overcome the limitations of the in vitro culture system, explants were placed in a hypoxic chamber (0.5% O 2 ) at the time of progesterone withdrawal to more accurately simulate the in vivo environment. The addition of hypoxic conditions induced a significant increase in IL-8 mRNA expression 48 hours after progesterone withdrawal, which suggests that hypoxia is necessary for the increase in endometrial repair factors at menstruation. To delineate the contribution of prostaglandins after progesterone withdrawal, the COX inhibitor indomethacin was added to the in vitro progesterone withdrawal system. This abrogated the up-regulation of IL-8 mRNA expression, indicating that both prostaglandins and hypoxia are required after progesterone withdrawal for up-regulation of repair factor expression.
To determine whether a similar human model of progesterone deprivation up-regulated IL-8 expression, endometrial biopsy specimens from women obtained before and after insertion of LNG-IUS were examined. This IUS markedly down-regulates the progesterone receptor in all endometrial compartments, resulting in a progesterone-deficient environment that simulates the in vitro model used in the present study. The added advantage of this in vivo human model is that the endometrial architecture remains intact, enabling the physiologic processes of chemoattraction and vasoconstriction. Previous studies of long-term progestogen exposure have demonstrated reduced endometrial perfusion and profoundly decreased vasomotion, which may induce a relative endometrial hypoxia. The results demonstrated that IL-8 mRNA expression in normal endometrium during the proliferative, early, and mid secretory phases is low. Paired samples obtained four to six months after LNG-IUS insertion demonstrated significantly increased IL-8 mRNA expression in all seven women. Levels after IUS insertion were comparable to those observed during the normal menstrual phase. The increased IL-8 mRNA expression in this LNG-IUS human model of progesterone withdrawal is comparable to the finding of significantly elevated IL-8 mRNA expression in endometrial samples from women obtained 48 hours after withdrawal of vaginal progesterone administration compared with mid secretory control endometrium.
After progesterone withdrawal during the late secretory phase, both PGE 2 and hypoxia are present in the luminal portion of the endometrium. Therefore, the effect of both PGE 2 plus hypoxic conditions on IL-8 expression in endometrial cells was examined. An Ishikawa endometrial epithelial cell line was used for these studies because primary human glandular endometrial epithelial cells have a limited capacity to proliferate in culture. Treatment with PGE 2 and hypoxia induced a synergistic increase in IL-8 mRNA and protein compared with either treatment alone, which suggests an interaction between the two pathways of IL-8 stimulation. Another endometrial proangiogenic factor, CYR61, has a similar regulation pattern. Endometrial cells treated with hypoxia and PGE 2 demonstrated a synergistic increase in CYR61 mRNA and protein levels. Mechanistic studies have described CYR61mediated induction of IL-8 receptors CXCR1 and CXCR2. Hence, there is evidence that hypoxia and PGE 2 initiate a perimenstrual angiogenic and tissue repair response by activation of CYR61-and IL-8-mediated signaling.
HIF-1 and NF-κB are two nuclear transcription factors present in the endometrium during the perimenstrual phase. The hypoxic response element and the NF-κB binding site have both previously been identified in the IL-8 promoter. Both HIF-1 and NF-κB up-regulate IL-8 mRNA expression in cells from other tissue sites in the body. An adenoviral dominantnegative inhibitor of NF-κB (Ad-Iκ-Bα) maintains NF-κB in a cytoplasmic location, preventing transcription of its target genes. On infection of endometrial epithelial cells with Ad-Iκ-Bα, there was a significant decrease in PGE 2 -mediated IL-8 mRNA up-regulation. Concomitant treatment with hypoxia and echinomycin revealed a significant reduction in hypoxia-mediated IL-8 mRNA expression. These results suggest that PGE 2 -mediated IL-8 up-regulation is NF-κB-dependent and that hypoxia-mediated IL-8 up-regulation is HIF-1mediated. Echinomycin also reduces c-Myc and AP-1 binding by 30% and 50%, respectively, and these transcription factors may also contribute to the decrease in IL-8 production. However, specific inhibition of HIF-1α with shRNA also demonstrated a significant reduction in hypoxia-mediated IL-8 expression. This supports the presence of an interaction between NF-kB and HIF-1α to regulate IL-8 expression. There is mounting evidence for cross-talk between NF-κB and HIF-1 in other tissue sites. Therefore, the presence of both of these transcription factors and possible cross-talk between them may explain the synergistic up-regulation of IL-8 mRNA observed in endometrial cells exposed to PGE 2 and hypoxic conditions simultaneously.
Aberrations in endometrial repair factor expression may lead to prolonged heavy menstrual bleeding. In women with menstrual blood loss in excess of 90 ml the PGF 2α -PGE 2 ratio is significantly decreased and prostaglandin F 2α receptor expression is also decreased. Excessive PGE 2 production at the expense of PGF 2α may result in less constriction of the spiral arterioles and an absent or decreased perimenstrual hypoxic insult. If endometrial repair factor expression depends on the interaction between PGE 2 and hypoxia-induced pathways, it can be speculated that endometrial repair processes may be defective in these women as a result of an altered hypoxic episode.
In summary, IL-8 mRNA and protein are increased in the human endometrium at menstruation. The present data support the hypothesis that progesterone withdrawal, followed by increased PGE 2 and hypoxic conditions, up-regulates endometrial repair factor expression. Endometrial IL-8 mRNA up-regulation may be mediated by NF-κB and HIF-1. Cross-talk between these two transcription factors presents a mechanism for the synergistic increases in IL-8 observed in endometrial cells when PGE 2 and hypoxia are present together. Further studies are required to determine whether hypoxic conditions and subsequent repair factor expression are aberrant in women with heavy menstrual bleeding.
7. SalesK.J.MaudsleyS.JabbourH.N.Elevated prostaglandin EP2 receptor in endometrial adenocarcinoma cells promotes vascular endothelial growth factor expression via cyclic 3′,5′adenosine monophosphate-mediated transactivation of the epidermal growth factor receptor and extracellular signal-regulated kinase 1/2 signaling pathwaysMol Endocrinol1820041533154515044590 8. KaneN.JonesM.BrosensJ.J.SaundersP.T.KellyR.W.CritchleyH.O.Transforming growth factor-{beta}1 attenuates expression of both the progesterone receptor and dickkopf in differentiated human endometrial stromal cellsMol Endocrinol22200871672818032694 9. KongD.ParkE.J.StephenA.G.CalvaniM.CardellinaJ.H.MonksA.FisherR.J.ShoemakerR.H.MelilloG. Echinomycin, a small-molecule inhibitor of hypoxia-inducible factor-1 DNA-binding activityCancer Res6520059047905516204079 10. MizukamiY.LiJ.ZhangX.ZimmerM.A.IliopoulosO.ChungD.C.Hypoxia-inducible factor-1independent regulation of vascular endothelial growth factor by hypoxia in colon cancerCancer Res6420041765177214996738 11. SowterH.M.RavalR.R.MooreJ.W.RatcliffeP.J.HarrisA.L.Predominant role of hypoxia-inducible transcription factor (HIF)-1alpha versus HIF-2alpha in regulation of the transcriptional response to hypoxiaCancer Res6320036130613414559790 12. JobinC.HaskillS.MayerL.PanjaA.SartorR.B.Evidence for altered regulation of I kappa B alpha degradation in human colonic epithelial cellsJ Immunol15819972262348977194 13. HenriksenP.A.HittM.XingZ.WangJ.HaslettC.RiemersmaR.A.WebbD.J.KotelevtsevY.V.SallenaveJ .M.Adenoviral gene delivery of elafin and secretory leukocyte protease inhibitor attenuates NFkappa B-dependent inflammatory responses of human endothelial cells and macrophages to atherogenic stimuliJ Immunol17220044535454415034071 14. DenisonF.C.RileyS.C.WathenN.C.ChardT.CalderA.A.KellyR.W.Differential concentrations of monocyte chemotactic protein-1 and interleukin-8 within the fluid compartments present during the first trimester of pregnancyHum Reprod131998229222959756313 15. CritchleyH.O.KellyR.W.KooyJ.Perivascular location of a chemokine interleukin-8 in human endometrium: a preliminary reportHum Reprod91994140614097989497 16. CritchleyH.O.WangH.KellyR.W.GebbieA.E.GlasierA.F.Progestin receptor isoforms and prostaglandin dehydrogenase in the endometrium of women using a levonorgestrel-releasing intrauterine systemHum Reprod131998121012179647549 17. MilneS.A.PerchickG.B.BoddyS.C.JabbourH.N.Expression, localization, and signaling of PGE(2) and EP2/EP4 receptors in human nonpregnant endometrium across the menstrual cycleJ Clin Endocrinol Metab8620014453445911549693 18. AriciA.SeliE.SenturkL.M.GutierrezL.S.OralE.TaylorH.S.Interleukin-8 in the human endometriumJ Clin Endocrinol Metab831998178317879589693 19. MilneS.A.CritchleyH.O.DrudyT.A.KellyR.W.BairdD.T.Perivascular interleukin-8 messenger ribonucleic acid expression in human endometrium varies across the menstrual cycle and in early pregnancy deciduaJ Clin Endocrinol Metab8419992563256710404837 20. JonesR.L.HannanN.J.Kaitu'uT.J.ZhangJ.SalamonsenL.A.Identification of chemokines important for leukocyte recruitment to the human endometrium at the times of embryo implantation and menstruationJ Clin Endocrinol Metab8920046155616715579772 21. LudwigH.SpornitzU.M.Microarchitecture of the human endometrium by scanning electron microscopy: menstrual desquamation and remodelingAnn NY Acad Sci622199128462064187 22. Gaide ChevronnayH.P.GalantC.LemoineP.CourtoyP.J.MarbaixE.HenrietP.Spatiotemporal coupling of focal extracellular matrix degradation and reconstruction in the menstrual human endometriumEndocrinology15020095094510519819954 23. SalamonsenL.A.ZhangJ.BrastedM.Leukocyte networks and human endometrial remodellingJ Reprod Immunol5720029510812385836 24. Kaitu'u-LinoT.J.MorisonN.B.SalamonsenL.A.Neutrophil depletion retards endometrial repair in a mouse modelCell Tissue Res328200719720617186309 Maybin et al. Page 23 Published as: Am J Pathol. 2011 March ; 178(3): 1245-1256. Sponsored Document Sponsored Document Sponsored Document Sponsored Document Sponsored Document Sponsored Document Maybin et al. Page 24
Table 1 Circulating Estradiol and Progesterone Concentrations at Endometrial Biopsy Histologic stage of cycle Age, mean, years E2, mean (range), pmol/L P4, mean (range), nmol/L Menstrual (n = 8) 41 192.25 (55-514) 3.71 (1.24-10.59) Proliferative (n = 16) 42 441.18 (79-1105) 2.81 (0.97-7.1) Early secretory (n = 10) 42 497.50 (289-841) 59.60 (23.2-112.91) Mid secretory (n = 11) 40 638.00 (242-1949) 64.30 (25.47-114.53) Late secretory (n = 6) 42 318.22 (59.09-819) 8.22 (1.06-16.95) Published as: Am J Pathol. 2011 March ; 178(3): 1245-1256.
Published as: Am J Pathol. 2011 March ; 178(3): 1245-1256.Sponsored DocumentSponsored DocumentSponsored Document
Published as: Am J Pathol. 2011 March ; 178(3): 1245-1256.Sponsored DocumentSponsored Document
Published as: Am J Pathol. 2011 March ; 178(3): 1245-1256.
We thank all of the women who participated in the study and clinical research nurses
Published as:
Sponsored Document Sponsored Document
Refer to Web version on PubMed Central for supplementary material.
The augmentation index predicts cardiovascular mortality and is usually explained as a distally reflected wave adding to the forward wave generated by systole. We propose that the capacitative properties of the aorta (the arterial reservoir) also contribute significantly to the augmentation index and have calculated the contribution of the arterial reservoir, independently of wave reflection, and assessed how these contributions change with aging. In 15 subjects (aged 53 Ϯ 10 yr), we measured pressure and Doppler velocity simultaneously in the proximal aorta using intra-arterial wires. We calculated the components of augmentation pressure in two ways: 1) into forward and backward (reflected) components by established separation methods, and 2) using an approach that accounts for an additional reservoir component. When the reservoir was ignored, augmentation pressure (22.7 Ϯ 13.9 mmHg) comprised a small forward wave (peak pressure ϭ 6.5 Ϯ 9.4 mmHg) and a larger backward wave (peak pressure ϭ 16.2 Ϯ 7.6 mmHg). After we took account of the reservoir, the contribution to augmentation pressure of the backward wave was reduced by 64% to 5.8 Ϯ 4.4 mmHg (P Ͻ 0.001), forward pressure was negligible, and reservoir pressure was the largest component (peak pressure ϭ 19.8 Ϯ 9.3 mmHg). With age, reservoir pressure increased progressively (9.9 mmHg/decade, r ϭ 0.69, P Ͻ 0.001). In conclusion, the augmentation index is principally determined by aortic reservoir function and other elastic arteries and only to a minor extent by reflected waves. Reservoir function rather than wave reflection changes markedly with aging, which accounts for the age-related changes in the aortic pressure waveform. arteries; blood pressure; wave reflection A GREATER UNDERSTANDING OF the mechanisms responsible for the morphology of the aortic pressure waveform may further our understanding of studies such as Anglo-Scandinavian Cardiac Outcomes Trial (ASCOT) and Second Australian National Blood Pressure Trial (ANBP2), where there were lower event rates with newer antihypertensive regimes (e.g., angiotensinconverting enzyme-and calcium channel blocker-based therapy) than older regimes (-blocker-and diuretic-based therapy), despite an equivalent brachial blood pressure (8, 9, 34).
The augmentation index (AIx) is a measure of the pressure increment from the shoulder of the systolic waveform, normalized to pulse pressure, and is considered a proxy measure of wave reflection (21,23,26). Aortic (or central) AIx can be readily estimated from the pulse waveform measured directly at the carotid or radial arteries with (3,4) or without (5) the use of a transfer function. However, studies examining the relationship between the AIx and cardiovascular events or mortality have produced disparate conclusions (6,7,10,18,31).
One explanation of these contradictory findings is that the AIx depends on both the timing and magnitude of reflected waves and that differential changes in each of these factors (e.g., with age) may account for inconsistencies. An alternative explanation is that the AIx is not simply or even largely a result of discrete reflected waves (33). Indeed, a metaanalysis of the published literature has failed to identify the changes in reflection timing on which the current hypotheses for pressure augmentation are based (1).
Recent studies have proposed that the capacitive or reservoir function of the aorta and large elastic arteries plays a major (and neglected) role in determining the morphology of the pulse waveform (30) and that the pressure waveform can be explained in terms of a reservoir pressure related to arterial compliance and an "excess" or wave-related pressure because of traveling waves, as first proposed by Lighthill (17). In anesthetized dogs, reflected waves were negligible in the proximal aorta after allowance for reservoir pressure, implying as originally suggested by Womersley that the design of the arterial tree minimizes backward wave reflections (36). Therefore, we hypothesized that once reservoir pressure is taken into account, the magnitude of reflections would be minimal in the human aorta.
We tested this hypothesis by using a "wave reservoir" model, which calculates reservoir pressure on the basis of the pressure and flow velocity waveforms. This approach was used to examine the relative contributions of longitudinal waves and the reservoir to the morphology of the aortic pressure waveform and to the magnitude of the AIx. Furthermore, we go on to explore how the reservoir and wave pressure components change with aging to determine the shape of the arterial pressure waveform.
Eighteen subjects in whom the probability of coronary artery disease was considered relatively low were recruited from a group of patients scheduled for coronary angiography. All subjects rested in bed for 1 h before angiography. Exclusion criteria included previous coronary intervention, valvular pathology, regional wall motion abnormality, and non-sinus rhythm. All subjects had good left ventricular function with ejection fraction exceeding 55%. All subjects were required to refrain from coffee and alcohol for at least 12 h and fasted for at least 9 h before the beginning of the study. Subjects who smoked were asked to refrain from smoking for at least 24 h before the study. All subjects received heparin (5,000 units intravenously) before the hemodynamic measurement phase; no other drugs were administered. Each subject gave written informed consent for participation in this study, which was approved by our local ethics committee. All studies were performed in accordance with our institution's guidelines. Of the 18 subjects screened, 15 were found to have no significant coronary artery disease on angiography (characteristics of these subjects are shown in Table 1), and these individuals proceeded to have measurements of aortic flow velocity and pressure. A 0.014-in.-diameter pressure wire (Wavewire) and a Doppler-flow wire (Flowwire, Volcano Therapeutics, formerly Jomed) were positioned 5 cm from the aortic valve using fluoroscopy. Each wire was carefully positioned to ensure that its sensor tip was aligned and that there was a stable signal. Concurrent analog output feeds were taken from an electrocardiogram and Wavewire and Flowwire consoles into a National Instruments DAQ-Card AI-16E-4 and acquired at 1 kHz using Labview. The recorded data were analyzed off-line using custom-written Matlab software (Mathworks, Natick, MA). The blood pressure and Doppler velocity recordings were filtered using a leastsquares (Savitzky-Golay) polynomial smoothing filter (25) and ensemble-averaged using the ECG R wave as a fiducial point.
AIx and augmentation pressure calculations. We used a conventional approach that defines the beginning of augmentation as the shoulder or inflection point (using the fourth-derivative method). At that instant in time, a reference value was taken for each of the two or three pressure components (forward wave pressure, backward wave pressure, and when included, reservoir pressure, Fig. 3). The highest subsequent value of that pressure component minus its reference value at the time of inflection is defined as the pressure augmentation by that component. The AIx was calculated as previously described (21).
As separated pressure is potentially sensitive to the imputed timings, we also calculated the area under the curve (or the pressure integral) over the complete cardiac cycle for each component, as a more time-independent measure.
Calculation of the arterial reservoir pressure and pressure separation. During left ventricular contraction, blood enters the aorta faster than it can leave; this increasing volume distends the aorta, increasing the reservoir pressure (12). The reservoir pressure generated is therefore determined by the instantaneous difference between inflow and outflow and the arterial compliance. The relationship between reservoir pressure and flow into the reservoir can also be viewed as the transverse impedance (19). During diastole, there is no inflow through the aortic valve and reservoir pressure declines quasiexponentially as blood leaves the aorta. We calculated the reservoir pressure using pressure and flow velocity as described by Wang et al. (30). Following the calculation of the reservoir pressure (Fig. 1, label 1), the pressure attributable to longitudinal waves was calculated by subtracting the reservoir pressure from the measured pressure (Fig. 1, label 2). Forward (Eq. 1) and backward (Eq. 2) pressures were then calculated using wave intensity analysis. This time domain approach gives essentially identical results to wave separation performed using frequency domain (impedance)-based approaches (28). An analysis was undertaken for wave-only and wave-reservoir models (Fig. 1, label 3), where dP is the incremental change in the measured or wave pressure, and dU the incremental change in blood velocity. is the density of blood (taken as 1,050 kg/m 3 ), and c is the wave speed calculated using the single-point equation (11).
Data are means Ϯ SD or n (%). P forward ϭ ͫ 1 2 ͑dP ϩ cdU͒ ͬ
The reflection coefficient (calculated as peak Pbackward/peak Pforward), separated pressures, augmentation pressure, and AIx were calculated by both ignoring and accounting for the arterial reservoir.
Reproducibility. The reproducibility of hemodynamic measurements was assessed by examining separate 30-s recordings of blood pressure and velocity for each patient. The standard deviation of the difference between these replicate recordings in the aorta was Ϯ5.9 mmHg for mean blood pressure (5% within subject coefficient of variation) and Ϯ0.31 ms Ϫ1 for mean Doppler velocity (16% within subject coefficient of variation).
StatView 5.0 (SAS Institute, Cary, NC) was used for statistical analyses. Continuous variables are reported as means Ϯ SE and categorical variables as n (in %). Comparisons were made using Student's t-test for continuous data, and a 2 test was used for categorical variables. Associations were examined using linear regression. P Ͻ 0.05 was taken as statistically significant.
The effect of the reservoir pressure on the calculated separated pressures. An example of a pressure waveform separated into forward and backward components, with and without inclusion of the reservoir pressure, is shown in Fig. 2. Accounting for the reservoir pressure markedly reduced peak backward (reflected) pressure by 88% (P Ͻ 0.001), peak forward pressure by 33% (P Ͻ 0.001), and the reflection coefficient by 81% (P Ͻ 0.001, Table 2).
Calculation of constituent components of the augmentation pressure. Augmentation pressure (the rise in pressure from the inflection point to peak systolic pressure) was 22.7 Ϯ 13.9 mmHg, and the AIx was 33 Ϯ 10%.
When the pressure constituents of the AIx were calculated ignoring the arterial reservoir, the peak backward (reflected) pressure component of the augmentation pressure was larger than the peak forward pressure component (16.2 Ϯ 7.6 vs. 6.5 Ϯ 9.4 mmHg, P Ͻ 0.001, Table 2) and backward pressure comprised 71% of the AIx (Fig. 3, left). However, when the pressure constituents of the AIx were calculated accounting for the arterial reservoir, the calculated backward (reflected) pres-sure was no longer the principal constituent of augmented pressure. Backward (reflected) pressure only accounted for 5.8 Ϯ 4.3 mmHg of augmentation pressure and only 25% of the AIx (a reduction of 64% compared with findings ignoring the reservoir pressure). The reservoir pressure was the largest component of the augmentation pressure (19.8 Ϯ 9.2 mmHg, P Ͻ 0.001, Table 2) and AIx (87%, Fig. 3, right), and there was little or no contribution from the forward-pressure component.
Reservoir phenomenon and wave speed. We examined the relationship between the magnitude of the reservoir pressure and aortic pulse wave velocity measured using the foot-to-foot technique over a 50-cm length of aorta (distal from the aortic root). Reservoir pressure was closely correlated to pulse-wave velocity squared (r ϭ 0.67, P Ͻ 0.001), which is not surprising since the square of the wave speed is inversely related to the distensibility of the aorta. With the use of regression analysis, pulse-wave velocity was found to increase by 3.42 m/s for every 10 mmHg increase in reservoir pressure. Values are means Ϯ SE. Pulse pressure, systolic pressure, and augmentation pressure were separated into their respective forward and backward Ϯ reservoir pressure. The pulse pressure components were calculated after subtraction of diastolic pressure. After accounting for the reservoir pressure, both forward and backward pressure was significantly reduced. Statistical comparisons were made using a paired Student's t-test. Changes in magnitude and timing of pressure constituents with age. We applied the same analysis over the entire pressure waveform to assess the changes in forward, backward, and reservoir pressures that occur with age. The whole group included subjects with characteristically different-shaped pressure waveforms (i.e., type A, type B, and type D beats) (2). The augmentation pressure and AIx both increased with aging. When the aortic reservoir was ignored, both the forward and backward pressures increased with aging (Table 3). However, when the aortic reservoir was accounted for, the forward pressure continued to increase with aging (Table 3, and Fig. 4A) but the backward pressure was no longer found to increase (Table 3, and Fig. 4B). The reservoir pressure was found to markedly increase with aging (Table 3, and Fig. 4C). These findings are essentially identical when applied using either pulse pressure or systolic pressure waveforms.
In this study, we have used a novel "wave reservoir" model to analyze the components making up the aortic pressure waveform in humans. In contrast to widely held assumptions, after accounting for the reservoir, the contribution from reflected waves is small, although not absent, and reservoir pressure is the dominant contributor to the AIx and augmentation pressure. Increases in both the arterial reservoir pressure and forward pressure and not in distal reflection explained the change in the aortic pressure waveform with aging.
Our study confirms previous observations regarding an increase in the AIx and augmentation pressure with aging but does not address the potential utility of the AIx as a clinical predictor of disease. Importantly, however, it provides a novel mechanistic insight into the factors responsible for the shape of the aortic pressure waveform in humans by demonstrating that while the onset of pressure augmentation (the shoulder) is determined by the arrival of the backward-traveling (reflected) wave, the degree of augmentation is principally determined by the arterial reservoir.
The concept of an arterial reservoir used in this study bears some similarities to the two-element windkessel concept introduced by Frank (12) and later developed into a three-element model (32). The two-element windkessel is now rarely used in hemodynamic modeling because of its zero-dimensional nature (i.e., it assumes an infinite wave speed) and its inability to accurately predict the pressure waveform in systole. Unlike the two-element windkessel, we do not propose that local loading involves instantaneous integration across the total compliance of the entire system but rather that the proximal aorta can be considered as one of a number of distributed elements, i.e., a transmission line model. Transmission line models give a more comprehensive description of the origins of the arterial pressure and flow waveforms but at the expense of computational and interpretative complexity. The wave-reservoir concept is a reduced model that draws a distinction between local (trans-Fig. 3. Calculation of the components of augmentation pressure ignoring (left) and accounting for (right) the aortic reservoir pressure. Pressure was separated first using conventional separation technique, which ignores the reservoir pressure, and then using the wave reservoir technique, which accounts for aortic reservoir pressure. Augmentation pressure was calculated as the rise in pressure between the inflection point and peak pressure. With the use of the wave-only analysis (left; which ignores the aortic reservoir), the augmentation pressure is composed predominantly from backward-traveling pressure with a small contribution from the ongoing forward-traveling pressure. When the reservoir is accounted for (right), pressure augmentation primarily arises from reservoir pressure, with a far smaller contribution arising from backward-traveling pressure. The forward-traveling pressure was found to no longer contribute. The pulse pressure and systolic pressure waveforms were separated into wave and reservoir components before and after accounting for the reservoir. Before the accounting for the reservoir pressure, both forward and backward pressure increased with increasing age. After the accounting for the arterial reservoir, reservoir pressure increased rapidly with age (note that aging explained 69% of the variance in reservoir pressure), exceeding the age-related increases in forward pressure. However, backward pressure was found to no longer increase with age. Observations were very similar when using either pulse pressure or systolic pressure waveforms.
verse) and distal (longitudinal) influences by employing a lumped model for the former while accommodating longitudinal wave travel in the latter. Wave speed is not assumed to be infinite in the wave-reservoir model, except in the local segment, which is assumed to be hydrodynamically compact (i.e., short in relation to the wave speed), permitting the application of a windkessel-type analysis.
As a result of the limitations of the windkessel-only model applied to the whole circulation, the traveling-wave paradigm and its associated analytical techniques have been widely adopted. These approaches successfully describe the shape of the pressure and flow waveforms, but pressure waveform separation results in a biologically implausible phenomenon: simultaneous "self canceling" forward-and backward-traveling waves in diastole. This is an inevitable consequence of the linear assumptions employed in wave separation and the neglect of a reservoir. During diastole, inflow into the proximal aorta is nearly zero, while pressure falls in a quasiexponential fashion. Consequently, any linear separation technique will result in forward and backward pressure with nearly equal magnitudes in diastole.
We propose that this phenomenon of self-canceling flow waves and additive pressure waves arises as a result of neglecting the increase in volume resulting from radial distension, which is substantial in the aorta and elastic arteries (24). Indeed, around 40% of stroke volume ejected in systole is stored in the distended elastic arteries (30). The aortic reservoir pressure, by definition, is proportional to the volume of blood stored in the aorta, which in turn depends on the compliance of the aorta and the impedance to outflow. Downstream impedance mismatching and wave reflection in will therefore contribute to the magnitude of the reservoir pressure. In effect, the arterial reservoir acts like a water tower (or multiple interconnecting water towers) storing volume in ejection and damping the pressure pulse and discharging the stored volume once ejection ceases. A subtraction of the reservoir pressure ac-counts for the potential energy stored in the reservoir and permits linear wave separation techniques to be applied.
Wave reflection, previously thought to be the major constituent of augmentation pressure and diastolic pressure, arises at sites of impedance mismatch (e.g., at branches) (13,22). While such discrete reflections do occur and may have substantial magnitude locally, the arrangement of the arterial tree markedly attenuates the backward travel of waves such that a limited amount of reflection is evident in the proximal aorta once the arterial reservoir is accounted for (30,33,35). In this regard, our observations are consistent with a previous invasive study of healthy children (27) and other studies in anesthetized dogs (16).
While wave reflection is not completely abolished after reservoir subtraction, our observations indicate that the AIx is not predominantly a measure of wave reflection but rather is largely due to the compliant properties of the aorta and other elastic arteries (15,20). As the aorta becomes stiffer (i.e., its compliance falls) with increasing age or disease, pressure in the arterial reservoir rises more rapidly for a similar increase in volume and reaches higher levels (Fig. 4, and Table 3) (29). This results in an increase in the AIx or augmentation pressure. This phenomenon, namely a change in local properties causing a change in the aortic waveform, has been observed in experiments that demonstrated acute changes in the pressure augmentation (akin to the pattern seen in aging) by applying a relatively noncompliant Teflon graft to the elastic portion of the aorta of healthy pigs (14) and dogs (15), without changes being made to the distal reflection sites (14).
In this study the mean age of our patients was 54 yr, ranging between 35-73 yr. While it was possible to demonstrate a clear difference in the timing of wave reflection between young, middle-aged, and elderly adult subjects, it is possible that the timing and magnitude of the reflected waves may differ in subjects even younger than those studied here. It is also possible that, in much younger subjects, with a fall in pressure after the shoulder (so-called C-type waveforms) or in subjects with markedly elevated heart rates, there may be greater difficulty in fitting the monoexponential segment of the reservoir pressure, although the concept of the reservoir phenomenon would still hold true. Ethically, it was only possible for us to make hemodynamic measurements in subjects already undergoing coronary angiography on clinical grounds, and subjects younger than 35 yr rarely require this invasive procedure.
The arterial reservoir makes a large contribution to the aortic blood pressure waveform in humans and is the principal component of the AIx. In contrast, wave reflection only made a minor contribution to the AIx. This reservoir pressure is the aggregate pressure resulting from the net difference between the total arterial system inflow and outflow divided by an effective arterial compliance. Reservoir pressure increases markedly with aging, probably as a result of decreased compliance, and this is the major factor accounting for the associated change in morphology of the aortic pressure waveform. Modifying the behavior of the arterial reservoir rather than changing wave reflection may be a useful target for future therapy.
This work was funded by a grant from the
None of the authors has a conflict of interest to disclose.
Su Y, Blake-Palmer KG, Fry AC, Best A, Brown AC, Hiemstra TF, Horita S, Zhou A, Toye AM, Karet FE. Glyceraldehyde 3-phosphate dehydrogenase is required for band 3 (anion exchanger 1) membrane residency in the mammalian kidney. Am J Physiol Renal Physiol 300: F157-F166, 2011. First published October 27, 2010; doi:10.1152/ajprenal.00228.2010.-The mammalian kidney isoform of the essential chloride-bicarbonate exchanger AE1 differs from its erythrocyte counterpart, being shorter at its N terminus. It has previously been reported that the glycolytic enzyme GAPDH interacts only with erythrocyte AE1, by binding to the portion not found in the kidney isoform. (Chu H, Low PS. Biochem J 400:143-151, 2006). We have identified GAPDH as a candidate binding partner for the C terminus of both AE1 and AE2. We show that full-length AE1 and GAPDH coimmunoprecipitated from both human and rat kidney as well as from Madin-Darby canine kidney (MDCK) cells stably expressing kidney AE1, while in human liver, AE2 coprecipitated with GAPDH. ELISA and glutathione S-transferase (GST) pull-down assays using GST-tagged C-terminal AE1 fusion protein confirmed that the interaction is direct; fluorescence titration revealed saturable binding kinetics with Kd 2.3 Ϯ 0.2 M. Further GST precipitation assays demonstrated that the D 902 EY residues in the D 902 EYDE motif located within the C terminus of AE1 are important for GAPDH binding. In vitro GAPDH activity was unaffected by C-terminal AE1 binding, unlike in erythrocytes. Also, differently from red cell Nterminal binding, GAPDH-AE1 C-terminal binding was not disrupted by phosphorylation of AE1 in kidney AE1-expressing MDCK cells. Importantly, small interfering RNA knockdown of GAPDH in these cells resulted in significant intracellular retention of AE1, with a concomitant reduction in AE1 at the cell membrane. These results indicate differences between kidney and erythrocyte AE1/GAPDH behavior and show that in the kidney, GAPDH is required for kidney AE1 to achieve stable basolateral residency.
ANION EXCHANGER 1 (AE1), also known as band 3, is a Na ϩindependent Cl Ϫ /HCO 3 Ϫ -transporting member of the SLC4 gene family of anion exchangers. It plays critical roles in the regulation of intracellular and systemic pH, intracellular Cl Ϫ levels and cell volume (reviewed in Ref. 2). AE1 is a polytopic plasma membrane protein that in mammals is expressed in erythrocytes (eAE1) and kidney (kAE1). kAE1 is normally located at the basolateral side of the acid secreting ␣-intercalated cell (␣-IC) of the collecting duct (reviewed in Ref. 36). Under the control of separate promoters, both AE1 isoforms are encoded by SLC4A1. The kAE1 promoter lies in intron 3, with an initiation codon in exon 5, resulting in humans in the absence of the first 65 amino acids that are present in human eAE1 (21). Mutations in SLC4A1 affecting eAE1 and/or kAE1 are associated with hereditary spherocytosis (HS; a dominantly inherited disorder) and distal renal tubular acidosis (dRTA), respectively (reviewed in Ref. 44). Notably, however, in most cases single mutations resulting in HS do not also produce dRTA, and vice versa. This suggests that for disease-causing mutations affecting the shared portion of AE1, the mechanisms involved in these two conditions must be different.
AE1 is composed of a large cytosolic N-terminal domain, a central transmembrane region that is predicted to span the lipid bilayer 12-14 times and is responsible for catalyzing one-forone exchange of Cl Ϫ for HCO 3 Ϫ , and a short cytosolic Cterminal tail. In humans, eAE1 is better characterized than kAE1. eAE1's N-terminal domain plays a cytoskeletal scaffolding role through binding to several proteins including ankyrin, protein 4.2, and protein 4.1, thereby contributing to maintenance of red cell shape and flexibility (6,26,36,39). The N terminus of eAE1 has also been reported to interact with various glycolytic enzymes, including GAPDH, aldolase, and phosphofructokinase-1, but the physiological significance of these associations remains unclear (7,10,17,18,24,28). Some, including GAPDH, have been shown to bind to eAE1 within the initial N-terminal portion that is missing from kAE1 and are reported not to interact with the N terminus of kAE1 (7,41,42).
Functions of the C-terminal tail of AE1 (AE1C) are less well understood than those of the other two domains. To date, the only reported binding partner for the AE1C domain is carbonic anhydrase II (CAII), but this remains controversial (reviewed in Ref. 2). We have demonstrated that kAE1, lacking the last 11 residues, corresponding to the R901X mutation in dRTA patients (19), loses its normal basolateral targeting in Madin-Darby canine kidney (MDCK) cells without loss of anion exchange function (12,37,38). This suggests that some basolateral targeting information is contained in the 11 residues at the extreme end of the AE1C domain, and this targeting is known to be dependent on the Y 904 residue within this region (12,36).
However, the molecular basis of kAE1 targeting remains to be elucidated. To further explore the function of the C-terminal domain, we sought binding partners for AE1C. Parallel yeast two-hybrid assays were conducted using either AE1C wildtype (AE1C-WT) or AE1C lacking the last 11 residues (AE1C-⌬11) as bait to screen a human kidney cDNA library. We report here the identification of a new C-terminal binding partner, GAPDH, for kAE1. We demonstrate that this partnership contributes to the stable basolateral residency of kAE1 and that its characteristics differ from those of the N-terminal eAE1/GAPDH interaction.
To express the C terminus of AE1 in bacterial cells, the coding sequence for the final 36 residues of human AE1 (AE1C-WT) 876LIFRN-VELQCLDADDAKATFDEEEGRDEYDEVAMPV911 (residues at the start and end of this domain are numbered; underscored amino acids were mutated separately, as detailed below) was cloned into the vector pGEX-4T1. Site-directed mutagenesis of this construct was performed using a QuikChange Site-Directed Mutagenesis Kit (Stratagene). Base substitutions were separately introduced into codons 902-903 (GATGAA¡GCTGCA), codon 904 (TAC¡GCC), or codons 905-906 (GACGAA¡GCCGCA), which resulted in the amino acid changes D 902 E¡AA (AA1), Y 904 ¡A (YA), or D 905 E¡AA (AA2), respectively. All inserts in constructs were amplified in-house by PCR and sequence-verified before use.
The resulting glutathione S-transferase (GST)-tagged AE1C-WT (GST_AE1C-WT) and mutant fusion proteins (GST_AE1C-AA1, GST_AE1C-YA, GST_AE1C-AA2, GST_AE1C-⌬11) were separately expressed in Escherichia coli BL21 cells and purified using glutathione Sepharose beads (Amersham Biosciences). To remove the GST tag from GST_AE1C-WT, purified fusion protein was incubated with thrombin (Sigma) at room temperature for 8 h, and AE1C-WT was then HPLC purified as previously reported (31).
To express intact kAE1 in MDCK cells, cDNA encoding kAE1 was subcloned from pHM6-kAE1 (12) into the vector pEGFP-C2 to create an N-terminal green fluorescent protein (GFP)-tagged kAE1 construct. This tag does not affect kAE1 trafficking or activity (5). The construct was subsequently cloned into the ⌬pMEP vector (15) containing a metallothionine promoter for stable expression in mammalian cell lines and sequence-verified. MDCK cells were transfected with the ⌬pMEP-GFP-AE1 vector using a Cell Line Nucleofection Kit L (Amaxa/Lonza) and, following selection in hygromycin B, were FACS sorted for GFP fluorescence. Cell lines expressing eGFP-kAE1 were maintained in media containing 200 g/ml hygromycin B.
Two peptides, corresponding to the coding sequence of the last 27 residues of human AE1 but differing in the phosphorylation state of Y 904 (nonphosphorylated: AE1C-Y 904 and phosphorylated: AE1C-pY 904 ) were synthesized and HPLC-purified to Ͼ95% by the University of Bristol Peptide Synthesis Facility.
Using Matchmaker Two-Hybrid System 3 (Clontech), AE1C-WT or AE1C-⌬11 was employed as bait to screen the Pre-Transformed Human Kidney Matchmaker cDNA Library (Clontech). Human kidney cDNA library clones were supplied in the pACT2 vector, containing a GAL4 activation domain, and were transformed into yeast strain Y187 (Clontech).
Positive colonies were selected by blue growth on SD/-Ade/-His/-Leu/-Trp/X--Gal plates. Library inserts were recovered from yeast by PCR and retransformed into Y187 in the pGADT7 vector (Clontech). They were then mated back to the original AE1 bait strains and to the control strain provided by the manufacturer to confirm specificity. Appropriate colonies were sequenced and identified through BLAST searches (http://ncbi.nlm.gov/blast).
Coimmunoprecipitation. Immunoprecipitation assays, using human or rat kidney or human liver samples obtained from the Cambridge Human Tissue Bank (Protocols 99/078 and 03/279), were carried out essentially as previously described (30,31). Briefly, 80 -120 g of each membrane sample were prepared and solubilized in buffer containing 10 mM Tris•HCl (pH 7.4), 1 mM EDTA, 1 mM DTT, 10% glycerol, 1.5% n-nonyl--D-glucopyranoside (n-NDG), and Protease Inhibitor Cocktail (Roche). All steps were carried out at 4°C unless otherwise stated. Twenty microliters of specific rabbit polyclonal antiserum directed against the C terminus of AE1 (residues 900 -911; gift of E. Martinez-Anso, Pamplona, Spain) or 40 g of purified goat polyclonal ␣-AE2 antibody (sc-46710) raised against an epitope within the first 50 amino acids of AE2, which is absent from AE1 (Santa Cruz Biotechnology; personal communication) were added to the recovered kidney and liver supernatants, respectively, before overnight incubation. Fifty microliters of ␣-rabbit IgG-agarose beads (for ␣-AE1) or protein G-Sepharose beads (for ␣-AE2) were added, and incubated for 1-4 h. The beads were then washed three times with buffer A [20 mM Tris•HCl (pH 7.4), 5 mM NaN3, and 0.3% n-NDG]; three times with buffer A containing 500 mM NaCl; and finally three times with buffer A. Bound proteins were eluted from beads by incubating in SDS sample buffer [0.175 M Tris•HCl (pH 6.8), 5.14% SDS, 18% glycerol, 0.3 M DTT, 0.006% bromophenol blue] for 5 min at 95°C (kidney) or 1 h at room temperature (liver), and supernatants were subjected to SDS-PAGE. Western blotting was performed with an ␣-GAPDH mouse monoclonal antibody (Abcam) or the relevant precipitating antibody, according to standard methods.
Immunoprecipitation assays using MDCK cells stably expressing kAE1 (38) were carried out similarly, except cells were treated with or without 200 M pervanadate for 30 min to maximize the phosphorylation state of AE1 (43), then lysed in buffer containing 150 mM NaCl, 20 mM Tris•HCl (pH 7.4), 10% glycerol, 1% NP-40, 10 mM sodium orthovanadate, 2 mM PMSF, Protease Inhibitor Cocktail set V (Calbiochem), and 1% Phosphatase Inhibitor Cocktail 2 (Sigma). Supernatants were precleared and transferred to either protein Gagarose beads preloaded with the mouse monoclonal ␣-AE1 Nterminal antibody Bric170 (IBGRL, Bristol) or protein A-agarose beads preloaded with ␣-rbAE1Ct, a rabbit polyclonal ␣-AE1 Cterminal antibody raised against residues 881-900 (40), for 4 h followed by three washes with cell lysis buffer. Proteins immunoprecipitated with Bric170 or ␣-rbAE1Ct were probed on blots with the ␣-GAPDH antibody, ␣-rbAE1Ct, or ␣ϪAE-PhosY 904 (a rabbit polyclonal antibody specific to AE1 pY 904 ) (43).
Full-length rabbit muscle-type GAPDH (Sigma), which shares 95% identity with human liver-type GAPDH identified by yeast two-hybrid assay, was first dissolved in 0.05 M Na2CO3/ NaHCO3 (pH 9.6) to a final concentration of 1 mg/ml, which was then immobilized onto a 96-well plate at 37°C for 1 h. Uncoated GAPDH was removed by three washes with TBST (TBSϩ0.5% Tween 20), and nonspecific sites were blocked with blocking buffer (TBSTϩ5% BSA) for 1 h at 37°C. Recombinant GST_AE1C-WT protein dissolved in blocking buffer (range 0.01-2 g/ml) was applied and incubated at 37°C for 1 h. GST replaced the GST fusion protein in parallel wells to evaluate nonspecific binding. After washing the wells six times with TBST, goat polyclonal ␣-GST antibody (Amersham Biosciences) diluted 1:1,000 in blocking buffer was added and incubated for 1 h at 37°C followed by six washes with TBST. For detection, horseradish peroxidase-conjugated ␣-goat IgG antibody (DAKO) was applied and incubated for 1 h at 37°C. Following six washes in TBST, bound proteins were visualized using ABTS (22% mg/vol in 50 mM sodium citrate, pH 4.0) containing 0.05% H2O2. Å405 values were measured using a microplate reader (Anthos HTII).
One hundred micrograms purified GSTtagged WT or mutant AE1C fusion protein was first immobilized onto glutathione Sepharose beads, followed by incubation overnight with GAPDH in PBST (PBSϩ2% Triton X-100) at 4°C. GST replaced GST fusion proteins as a negative control. Beads were collected and washed three times with PBST, then three times with PBST containing 500 mM NaCl; and finally three times with PBST. Bound proteins were eluted from beads by boiling in SDS sample buffer for 3 min at 95°C, and supernatants were subjected to SDS-PAGE. Western blotting was performed using the ␣-GAPDH antibody, and data obtained from three separate assays were quantified densitometrically using ImageJ software.
Fluorescence titration was performed in a LS55 Luminescence spectrometer (PerkinElmer Instruments) at 25°C with excitation wavelength of 295 nm (2.5-nm bandwidth) and emission of 340 nm (5-nm slit width). Then, 0.5 ml of 2.8 M GAPDH in PBS was titrated with either HPLC-purified AE1C-WT or synthetic AE1 C-terminal peptides (AE1C-Y 904 or AE1C-pY 904 ) over the range 0.16 -7.28 M. After each addition, the solution was allowed to equilibrate for 1 min before recording Trp fluorescence emission. Quenching of Trp fluorescence was analyzed using the F-F 0/F0 ratio (where F 0 and F are fluorescence intensities at 340 nm in the absence and presence of the AE1 peptides, respectively) plotted against peptide concentration using Origin (OriginLab software).
GAPDH activity was measured as described (23) at 340 nm and 37°C. Reaction mixtures contained 2.2 mM glyceradehyde-3-phosphate, 0.25 mM NAD ϩ , 20 mM sodium phosphate (pH 7.0), 100 mM sodium pyrophosphate (pH 8.5), 3 M DTT, and 6.6 nM GAPDH. To investigate the potential effects of AE1 on GAPDH activity, HPLC-purified AE1C-WT (25 M) was preincubated with the 6.6 nM GAPDH for 15 min at room temperature before inclusion in the reaction mixture. Assays were performed in quadruplicate.
Cell culture, GAPDH knockdown, and biotinylation. MDCK cells were cultured in DMEM (Sigma) supplemented with 10% FBS, penicillin (100 U/ml)/streptomycin (100 g/ml), and L-glutamine (2 mM) at 37°C with a 5% CO 2 atmosphere in a humidified incubator.
For RNAi experiments, all media were supplemented with 1 mM pyruvate as described elsewhere (46). Endogenous GAPDH expression in MDCK cells, stably expressing eGFP-kAE1 and grown in six-well plates, was depleted with a small interfering RNA (siRNA) oligonucleotide based on the canine GAPDH sequence (NM_001003142, sense strand: 5=-CCAAATATGACGACATCAA-3=) (Thermo Scientific Dharmacon). Three micrograms oligonucleotide, or a decoy siRNA (Silencer Negative Control siRNA, Ambion), were transfected into ϳ5 ϫ 10 5 cells using a Amaxa Nucleofector Kit L according to the manufacturer's instructions. Twenty-four hours later, transfection was repeated, cells were immediately seeded onto Transwell filters (Sigma), and kAE1 expression was induced with 2 M CdCl2 and 100 M ZnCl2. Cells were examined 2-3 days later.
GAPDH knockdown was evaluated by Western blotting of cell lysates. To examine levels of kAE1 at the plasma membrane, biotinylation assays were performed using a Cell Surface Protein Isolation Kit (Pierce) according to the manufacturer's instructions. Biotinylated surface proteins were separated by SDS-PAGE followed by Western blot analysis using antibodies against AE1 (Bric170) and GP135 (gift of F. Buss, Cambridge, UK), an apical membrane protein employed as a loading control.
ATP measurement. ATP levels in GAPDH-depleted or control cells were measured as described (16). Cells were first lysed in buffer containing 50 mM Tris•HCl (pH 7.4), 150 mM NaCl, 1% NP-40, 0.5% sodium deoxycholate, 5 mM EDTA, and Protease Inhibitor Cocktail (Roche), treated with equal volumes of trichloroacetic acid (40 g/l), and pH corrected to 7.0 with 1 M Tris. ATP content was calculated from a luciferin-luciferase assay (ATP Bioluminescence Assay Kit, Roche) using a GLOMAX 96 microplate luminometer (Promega). ATP levels were measured in triplicate using two separate knockdown preparations.
Immunofluorescence. Cells were fixed in 4% paraformaldehyde for 15 min and permeabilized with PBS containing 0.1% Triton X-100 for 5 min. All steps were then carried out at room temperature using a buffer containing PBS and 1% BSA. Following 15 min in the buffer, cells were incubated with rat monoclonal ␣-ZO-1 (Santa Cruz Biotechnology), or mouse monoclonal ␣-E-cadherin (BD Transduction Laboratories), ␣-tubulin (Sigma) antibodies, or Alexa Fluor 594 phalloidin (Molecular Probes) for 1 h. Goat ␣-rat (Santa Cruz Biotechnology) or goat ␣-mouse (Molecular Probes) Alexa Fluor 568 secondary antibody was applied for 1 h. Following mounting in Vectashield medium (Vector Laboratories), bound antibody was visualized using a LSM510 Confocal laser scanning microscope. Replacement of the primary antibody with an appropriate serum (rat or mouse) (Sigma), or omitting the primary antibody, was used as a negative control.
Data were analyzed using STATA 11 IC (College Station, TX) and are presented as means Ϯ SE. Differences were compared using the independent or paired Student's t-test as appropriate.
Parallel yeast two-hybrid assays using AE1C-WT or AE1C-⌬11 as bait to screen a human kidney cDNA library yielded 2 clones of the 40 sequenced, which were both 100% identical to the full coding sequence for the human liver-type glycolytic enzyme GAPDH, with no other matches. Subse- quent specific mating tests showed that they interacted with AE1C-WT but not AE1C-⌬11 (Fig. 1A), implicating the final 11 residues of AE1 (RDEY 904 DEVAMPV) as the binding moiety.
GAPDH coimmunoprecipitates with intact AE1 in both human and rat kidneys and with AE2 in the liver. For in vivo verification, we first performed coimmunoprecipitation assays from human kidney membrane, as well as from rat kidney where immunocolocalization of the two proteins has previously been demonstrated (14). We initially confirmed that eAE1 was undetectable in kidney membranes by Western blotting, (Supplementary Fig. S1A; all supplementary material for this article is available online at the journal web site) and also ascertained from the supplier that the AE2 epitope sequence is not present in AE1 (see Coimmunoprecipitation).
As shown in Fig. 1B, ␣-AE1 antiserum was able to coprecipitate GAPDH and AE1, indicating their potential association in human kidney (top left). GAPDH and AE1 similarly coimmunoprecipitated from rat kidney membrane (top right). Specificity of these assays was confirmed by the absence of GAPDH when the precipitating antibody was omitted (Ϫ lanes). In addition, with the use of an ␣-AE2 antibody, GAPDH coimmunoprecipitated from human liver membrane with AE2 (bottom). Probing with the precipitating antibody confirmed the presence of AE1 or AE2 in the precipitated complex (Supplementary Fig. S1, B and C). AE1 is not expressed in the liver, whereas AE2 expression is ubiquitous. Importantly, both kAE1 and AE2 lack the N-terminal D 6 DYED and E 19 EYED motifs found in eAE1 that have been identified as sites for AE1/GAPDH binding in red blood cells (10). Hence, the association between GAPDH and kAE1 or AE2 indicates the presence of previously unrecognized binding site(s). Examination of the protein sequences shows only one similar motif in each of kAE1 and AE2, within their C termini: D 902 EYDE in AE1 and D 1232 EYNE in AE2. The C termini of AE1 and AE2 share 86% sequence similarity, including this motif (Fig. 1C). In addition, these findings suggest that association between GAPDH and the anion exchange protein family is a more general phenomenon than previously reported.
Direct binding of GAPDH to AE1C-WT confirmed by GST pull-down analysis, ELISA, and fluorescence titration. Since coimmunoprecipitation suggests, but does not prove, a direct interaction between two proteins, we next performed a GST pull-down assay to confirm that GAPDH and AE1C-WT interact directly in vitro. GST-tagged AE1C-WT fusion protein was first expressed and purified (Supplementary Fig. S2, lane 2). Fusion protein immobilized on glutathione beads was incubated with GAPDH. Western blot analysis of washed and eluted proteins using ␣-GAPDH antibody clearly displayed the 36-kDa GAPDH band only in the GST_AE1C-WT sample and not with GST alone (Fig. 2A), demonstrating a specific and direct interaction between GAPDH and AE1C-WT proteins.
This direct interaction was further confirmed by ELISA, which as shown in Fig. 2B, demonstrated AE1C-WT binding to GAPDH in a specific, concentration-dependent, and saturable manner. The binding affinity of the AE1C-WT/GAPDH interaction was calculated in vitro using fluorescence titration (Fig. 2C). The GAPDH monomer contains three Trp residues, whereas no Trp residues are present in AE1C-WT. GAPDH fluorescence was measured in the presence of increasing concentrations of AE1C-WT when the proteins were excited at 295 nm. The observed fluorescence change was dependent upon the amount of AE1C-WT, as there was no change detectable when buffer alone was added (data not shown). The quenching of GAPDH Trp fluorescence yielded a K d value of 2.3 Ϯ 0.2 M.
The D 902 EYDE motif in the AE1C domain is dominant in GAPDH binding. By analogy with the N-terminal motif in human eAE1 described above, D 902 EYDE (which is missing from AE1C-⌬11) is the likeliest candidate motif for kAE1's GAPDH binding. To investigate this, we expressed, purified, and confirmed the purity and specificity of various GST-tagged AE1C variants: AE1C-⌬11 to mimic the known disease-causing truncation; AE1C-AA1, AE1C-YA, and AE1C-AA2 mutants to disrupt the putative acidic motif, where AA1 and AA2 represent D 902 E¡AA and D 905 E¡AA, respectively, and AE1C-YA is Y 904 ¡A (Supplementary Fig. S2, lanes 3-6). We performed parallel GST pull-down analyses using each of the purified fusion proteins incubated with GAPDH.
Fig. 2. GST pull-down, ELISA, and fluorescence titration analyses for binding of AE1C-WT to GAPDH. A: direct interaction between AE1C-WT and GAPDH was first confirmed by GST pull-down assay. Immobilized GST_AE1C-WT, but not GST alone, was able to pull down GAPDH, indicating binding. This direct interaction was further confirmed by ELISA (B) and fluorescence titration (C). B: ELISA plates coated with GAPDH were incubated with increasing concentrations of GST_AE1C-WT fusion protein () or GST alone (grey symbols). Specific binding of AE1C-WT to GAPDH is shown. C: GAPDH was titrated with increasing concentrations of AE1C-WT. The GAPDH tryptophan residues were excited at 295 nm, and fluorescence change due to binding of AE1C-WT was monitored at 340 nm, yielding saturable binding with Kd ϭ 2.3 Ϯ 0.2 M.
GST_AE1C-WT and GST alone replaced the mutant AE1C fusion proteins as positive and negative controls, respectively.
Supernatants eluted from beads were analyzed by Western blotting using ␣-GAPDH antibody (Fig. 3A, representative of 3 replicate experiments). Blotting for GST confirmed equivalent protein loading in all lanes. Densitometric analysis (Fig. 3B), correcting for GST band intensity, demonstrated that the DEYDE motif is indeed important for GAPDH binding, with a major reduction in the amount of GAPDH pulled down by AE1C-⌬11 compared with WT (67 Ϯ 8% reduction, P ϭ 0.001 vs. WT). A small amount of binding was preserved by the proximal portion of the C terminus, probably reflecting the presence of a second acidic patch (LDADD). Looking at the DEYDE motif in more detail, we found the first DE pair plus the Y in the D 902 EYDE motif to be more important than the second, since the amount of GAPDH associated with AE1-AA2 was comparable to WT levels, whereas AE1C-AA1 or AE1C-YA both reduced binding by similar amounts as the truncation (72 Ϯ 2.3 and 60 Ϯ 5.3% reductions, respectively, P Ͻ 0.001 for either vs. WT).
Phosphorylation of Y 904 of AE1 does not affect GAPDH binding. The C-terminal Y 904 within the DEYDE motif can be phosphorylated in both eAE1 and kAE1 (43,45), and phosphorylation of the C terminus has been implicated in acute internalization of kAE1 in an in vivo cell culture system (43). In red cell studies, Campanella et al. (7) have shown that phosphorylation of Y 8 and Y 21 , located within the two binding motifs at the N terminus of eAE1 and not found in kAE1, displaced GAPDH from erythrocyte membranes. We therefore asked whether phosphorylation of Y 904 would affect the GAPDH/kAE1 interaction. MDCK cells stably expressing kAE1 and treated with pervanadate were used as the model system (31,43). After pervanadate treatment, very little kAE1 was immunoprecipitated using Bric155 (which only recognizes the unphosphorylated C terminus) compared with untreated cells or when the N-terminal antibody Bric170 was used (recognizing both phosphorylated and unphosphorylated C terminus; Supplementary Fig. S3). This confirmed that the majority of kAE1 is phosphorylated under these conditions. Figure 4A shows that Y 904 -phosphorylated AE1 was immunoprecipitable using Bric170, as detected by pY 904 -specific antibody (␣AE1PhosY 904 ; middle), compared with the phosphorylation-independent ␣-C-terminal AE1 antibody (␣-rbAE1Ct; top). Second, at steady state (i.e., without pervanadate), the ␣AE1PhosY 904 antibody immunoblot revealed that no immunoprecipitable AE1 detected by ␣-rbAE1Ct (top, Ϫ lane) was phosphorylated at Y 904 (middle, Ϫ lane). Third, GAPDH coimmunoprecipitated with kAE1 regardless of whether Y 904 was phosphorylated (bottom). These experiments were repeated using ␣-rbAE1Ct as the immunoprecipitating antibody and then probed with a monoclonal GAPDH antibody, and results were similar (data not shown).
In addition, repeating the in vitro fluorescence titration assays using synthetic AE1 C-terminal peptide with or without Y 904 phosphorylation (AE1C-pY 904 and AE1C-Y 904 ) displayed very similar binding kinetics to GAPDH (Fig. 4B), indicating maintenance of the interaction. Taken together, these results suggest that phosphorylation of Y 904 does not diminish the capacity of GAPDH to bind to the AE1C domain. This implies a major difference between the GAPDH/AE1 interactions in red blood cells and the kidney.
In an in vitro assay of GAPDH function, specific GAPDH enzymatic activity in the absence of AE1C-WT was 111 Ϯ 15 mol NADH produced•min Ϫ1 •mg GAPDH Ϫ1 . Figure 4C shows that addition of an excess of AE1C-WT (25 M) had no significant effect (120 Ϯ 15; P ϭ 0.69). This result again differentiates the kidney/GAPDH interaction from that in erythrocytes, where catalytic activity of GAPDH was potently inhibited by addition of either erythrocyte membrane or the isolated N terminus of eAE1 (10,41).
siRNA knockdown of GAPDH leads to loss of basolateral AE1 in polarized cells. As displayed in Fig. 5B, bottom, polarized MDCK cells stably expressing eGFP-tagged WT kAE1 demonstrate normal basolateral localization of the fusion protein, consistent with the GFP tag having no effect on kAE1 trafficking (5). Using a specific canine siRNA, ϳ90% knockdown of GAPDH expression was achieved in these stable cells, whereas use of a decoy oligo bearing no homology to the canine genome had no effect (Fig. 5A). All media in these experiments were supplemented with pyruvate, as is routinely used to avoid energy losses following GAPDH depletion (3,22,46). No significant viability difference was observed between control and knockdown cells. Amounts of ATP were the same in knockdown and control cells, respectively: 2.74 Ϯ 0.38 and 2.64 Ϯ 0.13 nmol/mg protein (P ϭ 0.76), similar to those reported for cultured proximal renal tubular cells (1). Total levels of kAE1 were similar in knockdown and control cells (Fig. 5A). Immunolocalization demonstrated that following GAPDH knockdown, overall polarization was preserved in cell monolayers, as evidenced by preserved junctional ZO-1 staining (Fig. 5B). However, marked intracellular retention of kAE1 was observed in the knockdown cells (Fig. 5B, top), which we confirmed by surface biotinylation: Fig. 5C shows that minimal AE1 remained in the biotinylated (surface membrane) fraction of GAPDH-depleted cells compared with control cells. As a nonradioactive alternative to pulse-chase analysis, we performed a time course series at 4, 8, and 16 h after induction of AE1 expression, which in Fig. 5C also demonstrates that surface AE1 was low in knockdown cells at all time points. Since our preliminary studies indicated that plasma membrane expression of eGFP-kAE1 reaches steady state at 12-16 h (Supplementary Fig. S4), these data imply that GAPDH is more likely to be involved in AE1 translocation to the membrane rather than simply in its retention, since in the latter case, levels of AE1 in knockdown and control cells would have been more similar to each other at the early time points. Taken together, these data indicate a requirement of GAPDH for kAE1 to achieve normal membrane residency in renal epithelial cells.
To exclude the possibility that GAPDH knockdown exerted general effects on membrane targeting or cytoskeletal integrity, we also stained cells with ␣-E-cadherin, ␣-tubulin, and phalloidin. All three showed normal appearances in the GAPDHdepleted cells (Fig. 5, D and E). The former indicates that more than one basolateral targeting pathway is present, and the latter two confirm preservation of actin filament and microtubule structures in these cells.
Although earlier work has shown that the distal portion of the cytoplasmic C-terminal tail of AE1 contains targeting information for both delivery of the protein to its proper functional location in polarized renal epithelia (which in health is chiefly on the basolateral membrane of ␣-IC) and for regulation of this localization (11,12,37,38,43), little has to date been reported concerning interacting proteins for this domain of the molecule. Our data introduce GAPDH not only as a novel binding partner for the C-terminal tail of AE1 but as a component of AE1's ability to achieve basolateral residency in polarized cells.
In erythrocytes, the potential for an association of GAPDH with AE1 was in fact first proposed over 40 years ago (32). Only recently, however, were molecular studies employed that suggested docking motifs for GAPDH on human eAE1, involving tyrosine-containing sequences D 6 DYED and E 19 EYED located separately in two tandem binding regions within the first 23 residues in the N terminus (10). However, this presents challenges for three reasons. First, although in murine studies, colocalization and association of GAPDH with AE1 have been observed in both rat and mouse erythrocyte membranes (8,10,14), and in vitro evidence also suggests association of GAPDH with rat erythrocyte membrane preparations (4), neither of the two putative N-terminal GAPDH-binding motifs is conserved in the N terminus of either rat or mouse eAE1. Indeed, the absence of sequence conservation in this region of eAE1 is a common feature of many of the mammalian AE1 polypeptides. Second, a potential interaction between GAPDH and AE1 in the kidney was suggested by Ercolani et al. (14) through observation that GAPDH colocalizes with kAE1 at the rat kidney ␣-IC surface in a manner very similar to that observed in erythrocytes, which is supported by our successful coimmunoprecipitation of the two proteins from rat as well as human kidney despite the absence of the N-terminal region and earlier experimental evidence that the isolated N terminus of kAE1 did not interact with GAPDH (7,10). Third, Tsai et al. (41) reported that peptides composed of varying numbers of Nterminal residues of human eAE1, and therefore containing both D 6 DYED and E 19 EYED motifs, showed only 1-3% of the binding affinity for GAPDH displayed by whole erythrocyte ghosts containing intact eAE1. These observations, taken together with the several methods to confirm specificity that we report here, indicate that additional docking site(s) for GAPDH must exist within the AE1 polypeptide. The D 902 EYDE motif located in the last 11 residues forms an ideal candidate and is highly conserved across many mammalian AE1 orthologs and paralogs including those of the mouse and rat (Fig. 1E), with complete conservation of the D 902 EY residues that we have shown to be important. In contrast, the eAE1 D 6 DYED and E 19 EYED motifs are not found outside human eAE1.
A further suggestion of the biological importance of the C-terminal interaction between GAPDH and anion exchanger is provided by the ability of AE2 and GAPDH also to coimmunoprecipitate. AE2, which shares a C-terminal DEYNE motif highly similar to that in AE1C, is also a polarized epithelial membrane resident with a much more widespread tissue expression pattern. Thus the link between the anion exchanger and GAPDH, previously attributed to a part of the Fig. 5. Knockdown of GAPDH decreases kAE1 protein on the surface of polarized MDCK cells. Cells were transfected with 2 rounds of small interfering (si) RNA against GAPDH (KD) or scrambled sequence (C) over 2-3 days. A: Western blot analysis of cell lysates: densitometry confirmed ϳ90% knockdown of GAPDH protein compared with control cells. Levels of kAE1 were similar; -actin was a loading control. B: localization of full-length enhanced green fluorescent protein (eGFP)-tagged kAE1 in polarized MDCK cells by confocal microscopy (left low power, right high power). GAPDH-depleted cells retained polarization, as shown by the tight junctional marker ZO-1 (top, XZ sections, red stain), but AE1 was significantly intracellular (top, XY sections, green) compared with GAPDH-replete cells (bottom, XY sections). C: cell surface proteins were isolated following biotinylation at various times post-AE1 induction. At all time points, membrane levels of AE1 were severely depleted in knockdown cells (KD lanes) compared with control cells (C lanes). GP135, an apical membrane protein detected on the same blot, was used as a loading control. D: basolateral localization of E-cadherin was preserved in both KD (top) and control cells (bottom). E: preservation of cytoskeletal integrity: F-actin stained by phalloidin demonstrated no significant differences between KD (1-3) and control cells (4 -6). 1 and 4, Cross sections through microvilli; 2 and 5, below terminal web; 3 and 6, basal. Similarly, there were no differences in tubulin staining between KD (7) and control (8) cells. Bars ϭ 10 m.
AE1 protein known not to be present outside the red blood cell, is also true for other tissues.
In erythrocytes, which are end-differentiated, the functional significance of the AE1/GAPDH interaction is not clear. One hypothesis (7,8) is that linkage of glycolytic enzymes to the erythrocyte membrane, partly through AE1, might function to assist in efficient deposition of ATP synthesized by glycolysis in a "membrane pool." This compartmentalized energy could then be used to fuel neighboring ion pumps. Although this could also be happening in the kidney and in AE2-expressing epithelia, our demonstrations that the kAE1/GAPDH interaction neither regulates GAPDH glycolytic activity (unlike in red blood cells where in fact, binding actually diminishes GAPDH's glycolytic activity), nor is it altered by phosphorylation, imply a different role for the AE1/GAPDH interaction in the kidney. This is logical given that kAE1 membrane residency can be altered by endocytic recycling, whereas mature erythrocytes lack cellular machinery for endocytosis such that eAE1 remains in the membrane once inserted.
Defective trafficking of kAE1 has been observed in in vitro investigation of a variety of C-terminal mutants including both naturally occurring dRTA-causing mutations and specifically engineered ones. These abnormalities include nonpolarized and/or apical membrane localization or intracellular retention, depending on the sequence alteration and the degree of polarization of the cellular system under study, and reveal that more than one targeting motif is present, even within the final 11 residues. For example, the kAE1⌬11 mutant was shown to either locate to both apical and basolateral membranes or accumulate uniquely in the apical membrane, depending on whether transfection of AE1 was transient or stable (12,37), while replacement of Y 904 by phenylalanine or alanine demonstrated prominent intracellular retention (12,37), as did deletion of the last four amino acids of kAE1 (our unpublished observations). All these studies confirm that the C terminus of kAE1 is critical for correct movement to its final destination, whereas the marked intracellular retention of WT kAE1 at all time points that we observed following reduction in GAPDH levels in this study delineates the first functional protein interaction described to date.
While phosphorylation of AE1 and GAPDH binding can be seen as a means of spatial-temporal regulation of activity in red blood cells, our results suggest that in the kidney, movement of the whole complex presents an alternative mechanism. Indeed, since GAPDH remains bound to the kAE1 C terminus in the presence of phosphorylated Y 904 , there is potential for GAPDH localization to be regulated in tandem with kAE1 upon phosphorylation of kAE1. In support of this, similar kAE1:GAPDH band intensity ratios pre-and post-pervanadate treatment were observed on blots in all coimmunoprecipitation assays (Fig. 4 and data not shown).
Our finding of the involvement of GAPDH in the normal cellular behavior of kAE1 adds to the expanding repertoire of possible functions for GAPDH emerging from other mammalian studies, in addition to its essential role in the glucose metabolic pathway. These include participation in intracellular membrane transport, and fusion, microtubule bundling, and kinase activity (reviewed in Ref. 29). In the renal epithelial cell system we have employed, we have not found evidence of the latter two properties. First, cellular morphology was well preserved following GAPDH knockdown (Fig. 5E), and sec-ond, since kAE1 phosphorylation leads to its internalization (43), the GAPDH knockdown would have been predicted to lead to membrane retention of kAE1, the opposite of our observation.
There is, however, other supportive evidence for GAPDH's involvement in cargo transport, in vertebrate photoreceptors (9), and in nuclear translocation in conjunction with Siah, an ubiquitin-E3-ligase (3). In addition, GAPDH has been implicated in the endoplasmic reticulum to Golgi transport through direct interaction with both Rab-2 and microtubules (33)(34)(35), although the latter is not borne out by our or others' studies (13,25). Our results support a role for GAPDH in forward cargo transport, since knockdown of GAPDH resulted in less surface appearance of kAE1. However, since E-cadherin was normally located in knockdown cells, distinct pathways are required for basolateral membrane targeting of different membrane residents.
The specific pull-down assays employing the R901X (AE1C-⌬11) tail construct suggested preservation of some GAPDH binding (Fig. 3), whereas the original specific yeast two-hybrid test was positive only for the full-length construct. This difference could possibly result from differing detection limits between the very disparate techniques and systems. Although we cannot rule out the existence of an additional, low-affinity binding site in the C terminus of kAE1 or elsewhere on the molecule, it should be noted that the C-terminal tail of human AE1 contains a high percentage of acidic amino acids, with a pI value of ϳ3.8. The residual in vitro binding of R901X may therefore be explained by the presence of more proximal acidic residues, with the main binding being acid/ tyrosine dependent, as shown by the results with DE902-3AA and Y904A substitutions. To minimize possible nonspecific electrostatic interactions as pointed out previously (20, 27), we have as described been careful to perform our assays with high-ionic-strength buffers incorporating detergents.
Our data concerning direct binding of the C terminus of AE1 to GAPDH do disagree to some degree with those of Chu and Low (10), who reported, as we have here, that the C terminus of AE1 did not inhibit GAPDH activity, but they did not achieve GST pull-down. We cannot address the difference between our findings and theirs as, unfortunately, supporting data and experimental methods for this latter point were not included in their report.
In summary, our data demonstrate for the first time that nonerythrocytic anion exchange and GAPDH are partnered, and in the kidney at least, that this is functionally important via AE1's C terminus.
AJP-Renal Physiol• VOL 300 • JANUARY 2011 • www.ajprenal.org
We thank
This work was supported by the
No conflicts of interest, financial or otherwise, are declared by the authors.
Complement factor H-related protein 5 (CFHR5) nephropathy is a familial renal disease endemic in Cyprus. It is characterized by persistent microscopic hematuria, synpharyngitic macroscopic hematuria and progressive renal impairment. Isolated glomerular accumulation of complement component 3 (C3) is typical with variable degrees of glomerular inflammation. Affected individuals have a heterozygous internal duplication in the CFHR5 gene, although the mechanism through which this mutation results in renal disease is not understood. Notably, the risk of progressive renal failure in this condition is higher in males than females. We report the first documented case of recurrence of CFHR5 nephropathy in a renal transplant in a 53-yearold Cypriot male. Strikingly, histological changes of CFHR5 nephropathy were evident in the donor kidney 46 days post-transplantation. This unique case demonstrates that renal-derived CFHR5 protein cannot prevent the development of CFHR5 nephropathy.
Complement dysregulation is associated with several distinct patterns of glomerular pathology. Common to glomerular abnormalities associated with defective control of the alternative pathway is deposition of C3 in the absence of significant immunoglobulin (1,2). This pathological appearance typifies a number of conditions associated with genetic or acquired complement dysregulation, including dense deposit disease, C3 glomerulonephritis (C3GN) and CFHR5 nephropathy. 'C3 glomerulopathy' has recently been proposed as a new term under which this heterogeneous group of disorders can be classified (1).
C3GN is a feature of CFHR5 nephropathy, a familial renal disease characterized by persistent microscopic hematuria, synpharyngitic macroscopic hematuria and progressive renal failure (3). C3GN may be associated with membranoproliferative or mesangial proliferative features. Endemic in Cyprus, affected individuals have a heterozygous internal duplication in the CFHR5 gene. Previously, mutations in complement factor H, CD46 (membrane cofactor protein) and factor I were identified among patients with biopsy-proven C3GN (2), but CFHR5 nephropathy is the first description of C3GN associated with a mutation in the CFHR5 gene. CFHR5 is a member of the complement factor H (CFH) family, a group of highly related proteins encoded by genes located within the regulator of complement activation (RCA) gene cluster on chromosome 1. Comprising CFH, CFH-like protein (CFHL-1) and complement factor H-related proteins 1-5, the proteins are composed of individual domains termed short consensus repeats (SCRs), which display varying degrees of amino acid sequence similarity to each other. CFHR5 is a 65 kDa protein composed of nine SCRs, and the internal duplication in exons 2 and 3 characteristic of CFHR5 nephropathy results in an expressed protein with duplicated SCRs 1 and 2 respectively. Although the role of CFHR5 is not yet fully understood, its complement regulatory activity in vitro (4) and co-localization with renal complement deposits in vivo (5) suggest that it may play a role in complement regulation within the kidney. Furthermore, the mutant protein has been shown to have reduced affinity for glomerular-bound complement, raising the possibility of impaired targeting to
complement within the kidney (3). We report the case of a 53-year-old gentleman with end-stage renal failure (ESRF) secondary to CFHR5 nephropathy, who underwent renal transplantation from a deceased donor and was found to have evidence of disease recurrence in a transplant biopsy 46 days later.
A previously healthy British male with Cypriot ancestry was referred at the age of 36 with persistent microscopic hematuria, episodes of macroscopic hematuria coinciding with upper respiratory tract symptoms, and renal impairment (serum creatinine 178 lmol/L). He was not aware of any family history of renal disease at that time although relatives with CFHR5 nephropathy have subsequently been identified. Physical examination and blood pressure were normal. However urinalysis demonstrated 1+ blood and 1+ protein, and a renal biopsy was performed. Light microscopy revealed 30% glomerular obsolescence with a fibrous crescent in one glomerulus, and a few tubular red cell casts. Immunoperoxidase staining showed capillary wall C3 but notably was negative for IgA. Subendothelial electron dense deposits and rare subepithelial 'humps' were seen on electron microscopy (EM), as were mesangial deposits associated with an increase in mesangial cells and matrix. A further decline in renal function 6 years later led to a second biopsy (Figure 1). This demonstrated large segmental scars, capillary wall thickening with double contours and mesangial cell interposition. Granular capillary wall C3 was again evident in the absence of immunoglobulin staining. On EM there were subendothelial and mesangial deposits, and rare subepithelial deposits. These histological features are consistent with C3 glomerulonephritis.
Over the following 4 years he suffered progressive renal impairment and required renal replacement therapy at the age of 47. Six years after commencing hemodialysis, dur-ing which he had further episodes of macroscopic hematuria, he received a deceased donor renal transplant. The donor was a 65-year-old female with no significant past medical history, who had a creatinine of 82 lmol/L. HLA matching demonstrated a 2,1,1 mismatch, but there were no recipient class I or II HLA antibodies at the time of transplantation, and this proceeded with no complications. Immunosuppression included alemtuzumab and corticosteroids peri-operatively, with tacrolimus monotherapy continued at discharge. Although his renal function initially improved, his creatinine stabilized at 186 lmol/L, and in the presence of persistent microscopic hematuria and following a single episode of macroscopic hematuria (with a urine protein: creatinine ratio of 56 mg/mmol), he underwent a renal transplant biopsy, 46 days after transplantation (Figure 2).
Light microscopy showed occasional neutrophils in the capillary loops of one glomerulus and a small area of tubulointerstitial fibrosis. Immunoperoxidase staining showed capillary wall granular C3 and complement component 9 (C9), whilst electron microscopy showed increased mesangial matrix associated with scattered mesangial and subendothelial deposits (with new basement membrane beneath in two areas; Figure 2C). Complement component 4d (C4d) staining was negative. The findings in the renal transplant were therefore consistent with a recurrence of his original disease. Serum C3 was 0.73 g/L (normal range 0.7-1.7 g/L) at the time of biopsy, whilst CFH and factor I levels were 320 mg/L (207% of normal control) and 32 mg/L (177% of normal control), respectively.
A second transplant biopsy performed 3 months later on account of a rise in the serum creatinine to 250 lmol/L, showed granular (mainly mesangial) C3 staining associated with subendothelial, mesangial and subepithelial deposits. Indirect immunofluorescence also showed the presence of glomerular complement components 5b (C5b)-9 (Figure 2D). Polymerase chain reaction (PCR) using genomic DNA isolated from peripheral blood monocytes Vernon et al. and serum CFHR5 western blot analysis (Figure 3) revealed the presence of the heterozygous internal duplication in the CFHR5 gene, confirming the diagnosis of CFHR5 nephropathy.
To our knowledge this is the first description of recurrence of CFHR5 nephropathy in a transplant. The recurrence of CFHR5 nephropathy in an unrelated kidney demonstrates that local synthesis of normal CFHR5 by the kidney is not sufficient to prevent disease. However, we are aware of two other incompletely characterized cases with the CFHR5 mutation and renal disease, in which good allograft function was evident one decade after deceased donor renal transplantation. Firstly, a Cypriot male with a renal biopsy demonstrating C3GN reached ESRF at the age of 46. A deceased donor renal transplant was performed at the age of 48 and he died 12 years later from a myocardial infarction. The serum creatinine 3 months before his death was 89 lmol/L and the graft was never biopsied. He was a member of a family described previously and genotyping of his daughter and the offspring of his cousin demonstrated that he was an obligate carrier of the CFHR5 mutation (individual III-1 from family 1 in reference (3)). Secondly, a 40-year-old Cypriot individual with ESRF and small kidneys, who presented with macroscopic hematuria at the age of ten but had been lost to follow up, received a deceased donor renal transplant at the age of 43. The transplant functioned well for 10 years until it was lost following atheroembolic complications of diagnostic coronary angiography. He returned to hemodialysis and died 6 years later. Subsequent molecular testing on stored genomic DNA demonstrated the presence of the CFHR5 mutation. There was no record of either native or allograft renal biopsy. These cases, together with the observations that the original disease takes many decades to cause ESRF, suggest that graft loss due to recurrence of CFHR5 nephropathy is not inevitable. Clearly larger studies will be required to establish the clinical course of CFHR5 nephropathy in renal transplantation. In summary, we describe the first reported case of recurrence of CFHR5 nephropathy in an unrelated renal transplant. Notably histological recurrence was demonstrable only 46 days after transplantation.
American Journal of Transplantation 2011; 11: 152-155
American Journal of Transplantation 2011; 11: 152-155
The authors would like to thank the staff of the Histopathology department and
The authors of this manuscript have conflicts of interest to disclose as described by the American Journal of Transplantation. These are as follows:
This manuscript was neither prepared nor funded in any part by a commercial organization.
Consultant: Bio Nano Consulting (Paid Director); Ownership: ReOx Ltd.-Director and stockholder; Honoraria: Ipsen, Roche; Scientific Advisor: Roche Foundation for Anaemia Research (RoFAR); Other: Registrar (unpaid)-Academy of Medical Sciences, Board Member (unpaid)-Medical Education England, Chair Physiological Sciences Funding Committee (Paid)-Wellcome Trust.
This article assesses Charles Tilly's Durable Inequality and traces its influence. In writing Durable Inequality, Tilly sought to shift the research agenda of stratification scholars. But the book's initial impact was disappointing. In recent years, however, its influence has grown, suggesting a more enduring legacy.
It is an honor to be part of this volume and to appraise the influence of Charles Tilly's Durable Inequality. When I was just a doctoral student, attempting to navigate the rough waters of a department whose approach to comparative historical work ranged from the indifferent to the hostile, I went on a research trip to "interview" the archives, seeking a way to formulate a dissertation project on working class formation in nineteenth century America. Carol Connell, one of Tilly's students from his time at the University of Michigan, was then a junior professor at Stanford and she encouraged me to start my trip with a visit to Tilly's workshop at the University of Michigan. When I got there, I found a better vision for how academic work could be done, as well as both practical help and moral support from Tilly. Then, after I began an assistant professor position at Berkeley, I went to spend a semester's leave at Tilly's center at the New School, where I was able to participate in the sort of workshop setting I had so admired in Ann Arbor. Those experiences have profoundly shaped not only my work but also my teaching and mentorship. It was a great gift and I'm happy to have the opportunity to credit Charles Tilly's influence and help.
In areas like the study of contentious politics, those of us influenced by Tilly's work have the advantage of being able to track the development of his thinking and theorizing from structure to agency in his many books and articles. We can assess Dynamics of Contention, for example, in light of not only his theoretical work on collective action, From Mobilization to Revolution, but also his many empirical studies of contention in France and Britain. However, in the case of Durable Inequality (Tilly 1999), we do not have this large body of earlier work. Oh, we do have an article on proletarianization (Tilly 1984), which in retrospect foreshadows some of the approach. But we have nothing like the early, careful, empirically rich studies of collective action, where Tilly articulated alternative theoretical arguments and provided the kind of persuasive evidence that led so many of us to believe that he had better answers for how to understand contentious politics. Instead, Durable Inequality is Tilly's first major work on inequality and it draws very little on empirical work that Tilly did before his turn to agency.
Durable Inequality opens with a critique of the vast majority of research on stratification. It is an appreciative critique in that it notes that previous research has clearly identified the empirical reality that analysts of inequality should be able to explain (for example, by showing how much male female pay differences spring not from unequal pay within the same jobs but from job segregation). However, it is also a sharp critique for Tilly indicts stratification scholars for having done a poor job of actually explaining the extensive inequality they document. They have done a poor job, he suggests, because they have relied on an individualistic framework that identifies in one way or another "powerful agents or institutions that sort individuals whose attributes vary significantly into positions whose rewards differ greatly" (Tilly 2000b: 783-4). Such explanations are unsatisfactory he contends because categorical differences in advantages among human beings are much larger than individual differences in "attributes, preferences, or performances". And they endure much longer. So any satisfying explanation for inequality must begin by confronting the fact that categorical differences in advantages swamp individual differences. 1The bulk of Durable Inequality sketches a set of processes that Tilly suggests underlie the many and varied forms of inequality that historians, sociologists and anthropologists have uncovered and described. Two powerful processes are fundamental in his view: exploitation and opportunity hoarding. Exploitation is the process by which powerful connected people have control over resources and use those resources to enlist others in production of value while excluding them from the full value added by their efforts, using any of a number of means, such as legislation, work rules, and outright repression. Opportunity hoarding occurs when members of a categorically based network confine the use of the value-producing resource to others in the in-group. Tilly is careful to note that elites engage in opportunity hoarding but most of his examples are of non-elites who make peace-more or less-with a categorical distinction and look for ways to advance within it rather than breaking down the distinction. Behind his understanding of exploitation, as both Erik Olin Wright (2000) and Michael Mann (1999) have argued, prowls Marx's labor theory of value. And his notion of opportunity hoarding owes much to Weber's idea of social closure.
Two more processes help to cement inequality in Tilly's model: emulation (in which established organizational models are copied in new settings) and adaptation, or the creation of everyday procedures and practices that people use to cope with and so reproduce the categorical distinctions in their daily interactions. Here, of course, are echoes of the new institutionalism.
Tilly's model rests on these four processes. The model is both relational and interactional and it identifies organizations, broadly defined, as the key sites for the creation and maintenance of durable categorical inequalities. Tilly insists that the approaches of prior scholars who focus on human capital or on labor market scarcities or on discrimination, are not wrong so much as simply drawing attention to the by-product of the four processes he identifies.
With a dazzling array of examples and stunning erudition, Durable Inequality attempts to flesh out this abstract model and to make it plausible enough to inspire stratification scholars to shift their research program away from examining individuals and social mobility to instead focus on studying different combinations of mechanisms, settings and categories. Yet Durable Inequality did not have that immediate effect.
The initial reaction to the book was respectful and there was widespread agreement that it was an important theoretical contribution. But several objections were also raised, one of which is especially relevant for thinking about Tilly's role and influence in the history of social science: what constitutes an explanation? Michael Mann (1999) articulated a critique that I believe many others share: mechanisms aren't causality he implied, noting that although Tilly claimed to explain inequality, his focus on mechanisms left the cause of inequality unaddressed. 2 Mann wrote, "Since in [Tilly's] theory causes must obviously concern very long run processes whereby exploitation and hoarding are "installed", we require historical analysis of their emergence." He didn't find an account that explained why modern societies are dominated by the categories discussed in the book (gender, racial, ethnic, occupational), and thus came away from the book believing that it's conceptual reach far exceeded Tilly's grasp. Others complained that it underplayed agency and focused too much on the role of organizations in producing inequality (Laslett 2000;Morris 2000).
However, the most distressing critique for Tilly's himself was almost certainly that the book did not have a big impact on stratification scholars. Tilly tried in different venues to engage his critics and to more clearly lay out the research implications of his work-in Comparative Studies in Society and History (Tilly 2000a) and also in Contemporary Sociology (Tilly 2000b) -but to little immediate effect. In 2001, he published one of the best pieces on Durable Inequality in the journal, Anthropological Theory (Tilly 2001). He hoped, it seems to me, that having had less success than he would have liked with sociologists, he might find a more receptive audience with anthropologists.
I don't know if that will prove to be the case or not, but I do know that the impact in sociology has grown over time, and that recent work suggests that the influence of Durable Inequality on stratification researchers is growing to be something closer to what Tilly wished.
In 2005, in the Journal of Social History, the American historian Michael Katz along with Mark Stern and Jamie Fader analyzed a century of data on gender inequality in the United States and, inspired by Tilly, asked directly about the mechanisms that have reproduced it in modern U.S. history. They argue that durable gender inequality across the 20th century was paradoxical rather than relentless, as Tilly's book suggests. By this they mean that Tilly's model does not account for the paradox of group mobility that coexists with structural inequality and ends up reproducing it. They also suggest that a similar process of internal differentiation has been characteristic of other categorically unequal groups in American history, most notably African Americans and many ethnic minorities. Katz and his collaborators thus assert that their work offers a crucial supplement to Durable Inequality.
Categorically Unequal (2007) builds even more ambitiously on the edifice of Durable Inequality. It offers a systematic account of the American stratification system in the twentieth century and relies heavily on the framework elaborated in Durable Inequality to get at the mechanisms that have sustained racial, gender, and class inequalities in the United States. But it also tackles the origins of inequality in a most un-Tilly like fashion, drawing on social cognition studies and neuroscience to argue that humans have a natural tendency to think in categorical terms and, most controversially, that distinctions based on age, gender, race, and ethnicity are more or less hard-wired. This is a step that Tilly steadfastly rejected (2001: 363 1998: 64-5), and there are tensions in Massey's analysis that follow from his effort to marry Tilly's relational analysis with a social psychology of categorization. For example, some of the distinctions he sees as hard-wired, like age, never became an organizing principle of durable inequality in the US, while other distinctions that have no obvious basis in hardwiring, like class, do. Yet, Massey's use of Tilly 's central mechanisms (exploitation, opportunity hoarding, emulation, adaption) to illuminate the varying institutional processes that make different categorical inequalities more and less persistent over time does yield an important empirical finding, to wit, that progress in breaking down one dimension of categorical inequality often goes together with increased disadvantage for other categorically unequal groups. Most significantly, Massey shows that the shrinking of the class divide in the period form 1933 to 1974 went hand in hand with racial and gender inequality while progress on the gender divide over the last 30 years has been coupled with an increasing pervasiveness of class division. This finding, of course, runs smack up against Katz et al's suggestion that paradoxical inequality operates similarly for women, African Americans, and most ethnic groups in the United States, and also challenges Katz et al's argument that when group mobility coexists with structural inequality it can often end up reproducing the very categorical inequality it seems to challenge. Both Massey's and Katz et al's analysis of inequality rely on the kind of data that has been central to the scholarly exploration of stratification for decades: individual and occupational data from surveys or from the census. Such data is often the only kind available, especially for historical studies. However, as Tomaskovic-Devey et al. (2009) and Avent-Holt and Tomaskovic-Devey (2010) have recently noted, survey and census data on individuals is abstracted from its organizational context and is thus problematic for investigating the interactional and relational contexts of inequality that is fundamental to the theory advanced in Durable Inequality.
Sociologists have known since Bielby and Baron (1986), for example, that the finer the data we are able to gather on jobs within workplaces (i.e. within their organizational context), the greater the inequality we are likely to detect. In two recent articles that significantly advance Tilly's relational model, Tomaskovic-Devey et al. (2009) and Avent-Holt and Tomaskovic-Devey (2010) use data collected on organizations in the United States and Australia-including large and small forprofit firms, non-profit workplaces, and government agencies-to investigate inequality patterns in wages. Reasoning that wage distributions within firms result from actors negotiating and contesting who should receive greater and lesser rewards for their work, both articles examine whether earnings inequality is amplified when categorical distinctions are mapped onto positional differences inside the organizations. Looking specifically at wage disparity between two different groups: mangers and core production workers on the one hand, and core workers and the lowest paid workers on the other, they find that both in the U.S. and Australia, inequality is greater when categorical differences like gender, education, race (in the U.S.) and linguistic group (Australia) can be used by categorically advantaged groups to extract greater rewards for the positions they hold in the organization. They further find that while the generic processes in which actors attempt to exploit and hoard opportunities from categorically distinct others in work organizations are similar in the United States and Australia, status distinctions tend to produce larger wage inequalities in the US, largely because of its extremely decentralized wage bargaining system. They interpret the larger wage inequalities in the US as evidence that Tilly's generic relational model is enhanced when it is contextualized, as he would no doubt have agreed.
These demonstrations of the explanatory power of Durable Inequality in the work of Tomaskovic-Devey and his collaborators have appeared in key journals of mainstream stratification research and can be expected to generate new analyses and further questions. Thus, although Durable Inequality may have been less influential initially than some of Tilly's other works, that is clearly beginning to change. Another important development is that scholars are beginning to connect Tilly's ideas about the processes that drive inequality with the kinds of contentious politics that might potentially undermine them. For example, Gibson and Woolcock's (2008) study of development projects in rural Indonesia borrows from Durable Inequality to reconceptualize the meaning of "empowerment." They argue that empowerment can most usefully be thought of in terms of marginalized groups developing routines of contestation that expose and weaken at least one of the four causal mechanisms that drive durable inequality. They then use this reconceptualization to explain outcomes of struggles over local power relations.
In this article, I have concentrated on scholarship that engages deeply with the specific processes and mechanisms laid out in Durable Inequality. The book's influence, however, can also be seen in the increasing number of studies that find its title to be a useful metaphor or that cite it when emphasizing the stubborn persistence of categorical differences and the boundary maintenance such persistence entails (Sampson and Morenoff 2006;Sampson and Sharkey 2008;Schneider 2008;Light 2009;Stainback and Tomaskovic-Devey 2009). These, too, attest to an enduring legacy.
Let me conclude by relating something Tilly told me when I spent a leave at the New School. I asked him how he was so productive. He told me that he procrastinated on his "A-list" priorities by working on his "B" and "C-list" priorities. He added that although that was key to his ability to publish so many books and articles, he worried that it would mean that at the end of his life he would have never gotten to his "A-list." I like to think that his turn from structure to action in the 1990s was not only a matter of stocktaking and reformulation after years of denial as other contributions collected here so eloquently suggest, but also of finally allowing himself to tackle his "A" priorities. After all, he wrote a phenomenal amount in the 1990s, like someone not only with a lot to say but also with the joy that must come from finally letting yourself embark upon your top priorities if you have always put them on hold in the past. The result was a huge treasure trove of ideas and research questions. It's gratifying to see that at least when it comes to his work on Durable Inequality, others have finally taken up his ideas and are pursuing the kind of research questions he urged us to ask. In short, I like to think that this recent work using, extending, and challenging the theory in Durable Inequality is advancing one of Tilly's "A" priorities.
Open Access This article is distributed under the terms of the Creative Commons Attribution Noncommercial License which permits any noncommercial use, distribution, and reproduction in any medium, provided the original author(s) and source are credited.
In the interest of full disclosure, he makes this critique of my own work with my colleagues at Berkeley, Inequality By Design(Fischer et al. 1996). It is appreciative in that he evaluates our effort as besting Hernstein and Murray's in The Bell Curve, but sharp in that he argues that we fall into the trap of focusing on individual difference. Moreover, he writes that to the extent we do offer an institutional account, it's untidy and does not clearly identify causal mechanisms.
Tilly's answer to Mann and to the others who make similar critiques of the relational realism of his late work would no doubt have been, "how is why!"Tarrow (2008) documents this with respect to Tilly's work on contentious politics. See alsoTilly and Goodin 2006
Humanity has entered a new phase of sustainability challenges, the Anthropocene, in which human development has reached a scale where it affects vital planetary processes. Under the pressure from a quadruple squeeze-from population and development pressures, the anthropogenic climate crisis, the anthropogenic ecosystem crisis, and the risk of deleterious tipping points in the Earth system-the degrees of freedom for sustainable human exploitation of planet Earth are severely restrained. It is in this reality that a new green revolution in world food production needs to occur, to attain food security and human development over the coming decades. Global freshwater resources are, and will increasingly be, a fundamental limiting factor in feeding the world. Current water vulnerabilities in the regions in most need of large agricultural productivity improvements are projected to increase under the pressure from global environmental change. The sustainability challenge for world agriculture has to be set within the new global sustainability context. We present new proposed sustainability criteria for world agriculture, where world food production systems are transformed in order to allow humanity to stay within the safe operating space of planetary boundaries. In order to secure global resilience and thereby raise the chances of planet Earth to remain in the current desired state, conducive for human development on the long-term, these planetary boundaries need to be respected. This calls for a triply green revolution, which not only more than doubles food production in many regions of the world, but which also is environmentally sustainable, and invests in the untapped opportunities to use green water in rainfed agriculture as a key source of future productivity enhancement. To achieve such a global transformation of agriculture, there is a need for more innovative options for water interventions at the landscape scale, accounting for both green and blue water, as well as a new focus on cross-scale interactions, feed-backs and risks for unwanted regime shifts in the agro-ecological landscape.
Humanity has reached the global phase of sustainability challenges. Growing evidence over the past decade shows that mankind is causing undesired environmental impacts at regional to planetary scales (Steffen et al. 2004). It is increasingly clear that humanity embarked on the ''great acceleration'' of the human enterprise in the mid 1950s when the industrial metabolism reached a critical scale, with negative impacts on the environment accelerating, and causing, for the first time in human history, ecological impacts at a global level. These environmental trends show an abrupt ''hockey-stick'' shape, where centuries of relatively slow and linear change (equivalent to the shaft), abruptly shift in a negative direction over a short time period (equivalent to the blade), creating multiple hockey stick shapes for fundamental ecosystem functions such as climate regulation, land productivity, freshwater flows, and biodiversity, and for ecosystem services, such as fisheries and terrestrial foods.
In the last decade, the understanding of the integrated risks from these multiple, accelerated pressures has improved, and has also been observed in terms of impacts on large scale ecosystems on the planet (Carpenter et al. 2009;Eisenman and Wettlaufer 2009;Reid et al. 2009;Shakhova et al. 2010). It has now been established that the Earth system functions as an integrated, self-regulating and complex system (Lovelock 2006), but there is still limited knowledge of the earth system forces in play, particularly in terms of feedback dynamics, when multiple environmental hockey sticks play-as they certainly often do-as a well-trained hockey team at a planetary level. The ecological resilience of the Earth system to the human induced energy imbalance in the climate system due to emissions of carbon dioxide, is an important example of this interplay between biophysical processes on the planet that are subject to exponential negative trends. Roughly 50% of our global emissions of carbon dioxide are absorbed by terrestrial land systems and the oceans, providing a gigantic buffer to a planetary disturbance of the climate system. At the same time, the ecological trendlines for the world's oceans and land areas show negative ''hockey stick'' patterns for critical parameters affecting the longterm ability of these systems to sequester carbon, e.g., trends of rapid rise in eutrophication, overfishing, and land degradation.
Conceptually, therefore, we can talk of a new phase in the quest for sustainable development. We have entered the global era of sustainability, with evidence that humanity is hitting hard wired processes at the planetary level. This means that human development cannot-as one may argue it could until 50 years ago-be separated from the global commons, such as the climate system, the global freshwater cycle, and the global nutrient cycles. It also means that an integrated social-ecological approach to human development is required, which moreover integrates the importance of resilience and the capacity of systems to remain in a-for humans-desired state without crossing thresholds and falling, often abruptly and irreversibly, into undesired states.
It is thus in the Anthropocene that the challenge of feeding humanity needs to be resolved. The number of hungry in the world remains at approximately one billion people, and in order to feed the world in 2050 global food production may have to increase by at least 70% (IAASTD 2009). This requires nothing less than a new green revolution, which moreover will have to occur mostly in the world's most social-ecologically vulnerable regions. Finite land and freshwater resources will form the limiting factors for this challenge. This paper explores the boundaries that define the challenge of meeting the freshwater needs to feed a world in 2050 in the Anthropocene.
This ''new'' social-ecological challenge is complex, and can-in a simplified form-be conceptualised as a ''quadruple squeeze'' on humanity's ability to secure long term sustainable development on planet Earth (Fig. 1). The first squeeze consists of the demographic growth requirements, which arises from the projected expansion of the current 6.8 billion people in the world to a population that is projected to surpass nine billion in only 40 years (UN DESA 2009). Moreover, the planetary pressure from the demographic squeeze is characterized by a 20/80 dilemma, with the old ''industrialized'' and rich economies which represent only a minority on the planet (the 20% of the world's population that predominantly suffer from what Ashok Khosla, president of the IUCN, has defined as ''affluencia'', Ashok Khosla, personal communication) having caused the bulk of the historic acceleration of environmental pressures on the planet, while the poor majority (80% of the world's population suffering from ''povertitis'', Ashok Khosla, personal communication) are most vulnerable to the impacts of global environmental degradation, and are-at least to a significant extent-on a positive development trajectory towards improved human welfare and economic growth (Rockstro ¨m et al. 2010a). However, the trend so far, is that this positive development momentum occurs in an unsustainable way, contributing to a major acceleration, petrifying the hockey stick pattern, of negative pressures on the planet.
The second squeeze consists of the global anthropogenic climate crisis, which, despite being only one among several global environmental challenges, occurs globally, affects essentially all other biophysical systems on the planet, and may trigger fundamental shifts in preconditions for human development. The climate ''squeeze'' is characterised by a dilemma represented here by 550/450/350. The policy interpretation of the IPCC 4th assessment report (AR4) is that a stabilisation of the concentration of CO2 at 450 ppm may provide a good enough chance of avoiding a global average temperature increase exceeding 2°C (considered as a threshold for dangerous climate change) (WBGU 2009). Projections indicate that the world is rapidly moving
Human growth 20/80 dilemma Ecosystems 60 % loss dilemma Climate 550/450/350 dilemma Surprise 99/1 dilemma Human growth 20/80 dilemma Ecosystems 60 % loss dilemma Climate 550/450/350 dilemma Surprise 99/1 dilemma towards concentrations of 550 ppm and beyond (International Energy Agency 2008; Richardson et al. 2009). Post-IPCC AR4 science suggests that the systems on Earth may be more sensitive to anthropogenic warming than previously thought, e.g., when including feedbacks caused by surface albedo change from melting ice sheets, which indicate that a stabilisation at 350 ppm in fact may be necessary to reduce risks of dangerous climate change (Hansen et al. 2008). We have today reached 390 ppm. Anthropogenic climate change is a major disturbance regime on the planet, which one would have hoped occurred at a state of high planetary resilience. Unfortunately, evidence indicates that this is not the case, and that we in fact face a global ecosystem crisis-the third squeeze-simultaneously with the global climate crisis. The UN Millennium Ecosystem Assessment (MEA 2005) showed that humans have accelerated the degradation of ecosystems during the past 50 years, deteriorating the capacity of 60% of key ecosystem functions and services to continue delivering human wellbeing and resilience in the future. Two key functions are the capacities of ecosystems to function as sinks of carbon and to regulate water flows in landscapes. Some 50% of the global emissions of greenhouse gases (GHGs) are absorbed by marine and terrestrial ecosystems, a capacity that may be on the decline (Canadell et al. 2007).
The fourth planetary ''squeeze'' is the growing insights of the universality of surprise in ecosystem change. We have developed our predominant social and economic paradigms on the erroneous notion that ecosystems change occurs in incremental, generally linear, and thereby predictable (and controllable) ways. Instead, empirical evidence suggests the reverse. Ecosystems change in nonlinear ways as a response to disturbance regimes, often abruptly and irreversibly. Multiple stable states separated by thresholds characterise systems ranging from lakes to savannahs (Scheffer et al. 2001). Critical tipping elements have been identified in the climate system (Lenton et al. 2008), water related regime shifts in agricultural systems (Gordon et al. 2008), and key tipping points in the Earth system (Schellnhuber 2009). The surprise and non-linear reality generates a dilemma that 99% of change in ecosystems tends to occur from 1% of events (S. Carpenter, personal communication), such as major shifts in forests or marine systems after fires and storms, etc. Stewardship of ecosystems that adapt to surprise requires redundancy and buffering capacity in order to build resilience to shocks and disturbance, which reduces the operating space for human development (Chapin et al. 2010).
The quadruple squeeze creates a complex social-ecological cocktail of planetary interactions that pose critical challenges for human development. We may have entered a new geological era, the Anthropocene, where humanity now constitutes the major driving force of change at the planetary scale (Steffen et al. 2007). Based on scientific evidence, sustainability or collapse is a question that now must be posed (Constanza et al. 2007).
A large number of poor people depend on agriculture for livelihood security and agriculture plays a key role in economic development (World Bank 2005) and poverty reduction (Irz and Roe 2000). For instance, in sub-Saharan Africa agriculture accounts for 35% of GDP and employs more than 70% of the population (World Bank 2000). Presently around 3 billion people live in the tropics and sub-tropics (World Bank 2008), and approximately a fifth of the world population lives in water-constrained, agricultural areas (Rockstro ¨m and Karlberg 2009). More than 20 years ago, Falkenmark (1986) found a correlation between poverty and water stress. Since then it has also been shown that the countries suffering from the largest prevalence of malnutrition, according to the UN Millennium Development Project, are commonly located in the semi-arid and dry sub-humid regions of the world (SEI 2005). Clearly, there are strong linkages between poverty, hunger and water, which make many poor people vulnerable to changes in water availability and distribution in time and space. It is important to understand these ''inherent'' vulnerabilities when addressing the need for a new green revolution, and the threats posed by the risk of reinforced water scarcity induced by anthropogenic global environmental change.
Agriculture in the tropical zone is largely dependent on the green water resource, i.e., the water that infiltrates and is stored in the soil profile (Falkenmark and Rockstro ¨m 2004). Blue water additions for irrigation, i.e., water in rivers, lakes and groundwater, only play a small part in the total water balance. Absolute water scarcity in the tropics is rarely the reason for the low yields commonly experienced in the area, although drought is often blamed for crop failure and crop reductions (Rockstro ¨m et al. 2010b). Rainfall in the tropical zone is erratic and rainfall variability generates dry-spells almost every rainy season in the semi-arid and dry-subhumid regions (Barron et al. 2003). Poor water resources management commonly results in an inability to bridge these intra-seasonal dry-spells, and this causes the lion share of all yield reductions commonly ascribed to drought. In other words, a large share of the yield reductions could be avoided with better water management-a great opportunity for agriculture in the tropical drylands (Rockstro ¨m et al. 2010b).
A global assessment of land and water availability for food production indicates that there is in fact enough water globally to produce sufficient food both today and also in 2050 (Falkenmark et al. 2009). However, in some areas of the world, there are local deficits to meet the national water requirements for food production (Fig. 2). These counties will, as a first option, have to rely on food imports, and secondly, when the economic situation in the country does not allow for trade to cover the food deficit, will have to resort to national solutions such as unsustainable expansion of agriculture or food aid.
An assessment of water vulnerability and food production for the African continent depicts an alarming picture for the future. In the year 2000, the water requirement to feed the 800 million inhabitants on the continent was approximately 1000 km 3 year -1 , assuming a food water requirement of 1300 m 3 cap -1 year -1 (Falkenmark and Rockstro ¨m 2004). Our estimates show that on agricultural land in Africa, the estimated water availability in the year 2000 for food production was in the same order of magnitude as water requirements, i.e. also around 1000 km 3 year -1 (calc. based on Rockstro ¨m et al. 2009a). In other words, according to these estimates it is possible to produce enough food to meet requirements from a land and water perspective; however, this analysis does not account for inequalities in resources distribution and wealth for example. A projection to 2050 shows a less optimistic picture for the continent. Assuming water productivity improvements that reduce food water requirements to 1000 m 3 cap -1 year -1 (Rockstro ¨m et al. 2009a), the total water requirement for food for the African continent is expected to double compared with the year 2000 due to population growth. At the same time, the amount of water available for food production is expected to increase by 500 km 3 year -1 due to irrigation expansion and intensification of water use on grazing lands, also accounting for climate change (calc. based on Rockstro ¨m et al. 2009a). The net effect is that Africa will not be able to meet the total continental water requirement for food by 2050.
As the climate becomes more extreme in the future many poor that depend on agriculture as their main source of income may become even more vulnerable. In large parts of the tropical zone, there is more than a 90% chance that the summer-averaged temperature will exceed the highest temperature on record in 2090, which is expected to significantly reduce crop yields (Battisti and Naylor 2009). It is also in the tropical zone that the countries classified as most vulnerable to climate-related water challenges can be found (Sullivan and Huntingford 2009). Due to climate change, global crop production is expected to decrease by around 10% by 2050 (Rost et al. 2009). However, this figure does not include the effect of increased CO 2 fertilisation (Tubiello and Ewert 2002;Long et al. 2006;Challinor and Wheeler 2008), temperature stress (Battisti and Naylor 2009) and increased tropospheric ozone which also has been shown to negatively affect yields (Emberson et al. 2009). The consequences of these multiple and simultaneous drivers of change in agriculture are poorly understood. To conclude, it is clear that a more erratic future climate is likely to increase the water related vulnerability of farmers in the tropical zone, many of which are impoverished already today.
Resilience provides the capacity of a system to cope with shocks and undergo change while retaining essentially the same structure and function (Walker et al. 2009). Resilience is here broadly understood as the ability of socialecological systems (SES) to persist in a desired state, the capacity to adapt within a given state, and to transform into new development trajectories in situations of crisis (Folke and Rockstro ¨m 2009). Even minor disturbances can push the system into a new regime if its resilience is low. Characteristic of a regime shift is that returning to the Fig. 2 Countries with a surplus of water (export), water deficit countries that can compensate their lack of water for food production with trade (import), and water deficit countries that will have to rely on national solutions to meet their remaining deficits. Based on data from Rockstro ¨m et al. (2009a)original regime is either very difficult or even impossible. A resilience approach entails identifying alternate system regimes and the thresholds between them, and the internal slow variables within the SES that interact and can cause the system to shift into an unproductive state due to external shocks (Walker et al. 2009).
Agriculture and water related regime shifts were described by Gordon et al. (2008). The study suggests that commonly it is possible to identify a productive (desirable) regime and a non-productive (undesirable) regime in agroecosystems. In tropical agricultural systems some farmers are currently locked into unproductive states, in which they lack capital to invest in agricultural inputs such as fertilisers and good crop varieties, and they commonly exhaust asset holdings accumulated from good years during periods of drought (Enfors and Gordon 2008). It has been suggested that investments in agricultural water management interventions would not only result in increases of average yields, but would also lead to reduced risks for crop failure (Rockstro ¨m et al. 2010b). As the return of investments become more reliable, farmers may be more likely to invest in additional inputs such as fertilisers, better crop varieties and improved management strategies. Local communities show a high dependency on local ecosystem services both as a supplement for crop production generating off-farm incomes (Cooper et al. 2008), and increasingly so in times of failing on-farm yields (e.g., gathering of wild growing fruits and vegetables; charcoal) (Enfors and Gordon 2008).
The challenge for world food production over the coming 40-50 years is to achieve a major production increase while building resilience in the face of the pressures from the quadruple squeeze. A first attempt to translate the global sustainability challenge in the Anthropocene was recently made with the introduction of the ''planetary boundaries'' concept, aimed at providing a safe operating space for humanity at the planetary scale (Rockstro ¨m et al. 2009b, c). The concept evolves from linking global change with resilience science, and provides a framework to identify physical boundaries for key earth system processes associated with thresholds that may jeopardise the desired stability of the planet to continue providing favourable conditions for human development. This analysis identified nine key earth system processes (climate change, stratospheric ozone depletion, ocean acidification, land use change, freshwater depletion, rate of biodiversity loss, interference with the global N and P cycles, chemical pollution and aerosol loading). Together, the nine boundaries associated with these processes provide a first attempt of defining a global safe operating space for humanity in the Anthropocene, with the aim of avoiding large scale undesired ecological surprise.
Importantly, several of the proposed planetary boundaries are coupled to world agriculture. Agricultural expansion is the largest anthropogenic transformation of land use on the planet, currently covering some 35% of the total land area (*12% for cropland) (Foley et al. 2005;Ramankutty et al. 2008). Moreover, agriculture is estimated to be a major source of GHG emission, accounting for *30% of global annual emissions (including the effects of deforestation), and constitutes, over the past 50 years, the largest driver behind loss of biodiversity, ecosystem change, and increase in freshwater use (MEA 2005).
A first attempt to define the specifications for a new revolution in agriculture as defined by these planetary boundaries is presented in Box 1, and will require major shifts in agricultural production systems world-wide. Agriculture must transform from being a source to a sink of GHGs. The food production increase essentially will have to occur on current cropland, except for certain regions in Africa, central Europe and parts of Latin America, where there still appears to exist land that can be converted to agriculture in a sustainable way. Blue water extraction for irrigation is limited, and thus any yield improvements have to be linked with corresponding improvements in water productivity. Plant nutrient cycles (N and P) must be managed more efficiently and management interventions in agriculture have to be implemented in such a way that the Box 1 Global specification of sustainability requirements for world agriculture in order to stay within the Planetary Boundaries
Climate change To stay within 350 ppm requires an agricultural system that goes from being a source to a global sink.
Land use change Cropland can only expand from 12 to 15%. Higher yields have to be produced on current croplands by increasing productivity.
Freshwater use Keep global consumptive use of blue water \4000 km 3 /year. We are at 2,600 km 3 /year today, and thus irrigation expansion is limited.
Interference with global N and P cycles Reduce to 25% of current N extraction from atmosphere. Not increase P inflow to oceans. Rate of biodiversity loss Reduce loss of biodiversity to \10 E/MSY from current 100-1000 E/MSY (E/MSY = extinctions per million species per year).
rate of species loss does not exceed the biodiversity threshold. This requires nothing less than a triply green revolution of agriculture (as compared to Gordon Conways call for a doubly green revolution (Conway 1997)): more yields (green) have to be produced, in a green (sustainable) way, by improving especially the green water productivity. It will thus not suffice to ''minimise environmental impacts'' of conventional, fossil-fuel based and resource depleting farming systems. Instead, the green sustainability leg of such a revolution would require that agricultural development contributes to allow humanity to stay within the safe operating space provided by the planetary boundaries. Second, since the main water source for agricultural production is green water, this entails managing the green water resource more efficiently than is currently done.
A triply green revolution in agriculture thus entails more efficient green water use on current croplands. Water is the bloodstream of the landscape and efficient integrated water and land resources management (IWLRM) requires governance across scales. A stronger emphasis on green water management entails a down-scaling of the focus of water resource management, from the current emphasis on river basins, to a stronger focus on management of smaller meso-scale catchments (10-1000 km 2 ) (Rockstro ¨m et al. 2010b). It is here that green water flows are actively involved in producing bio-resource values for human wellbeing, and providing regulatory flows-both green and blue-across scales as water flows through the landscape. Managing green water flows for increased resilience at the meso-scale are determined by several processes (Fig. 3). Increased green water use for biomass production, be it for food, fiber, fuel, or forestry, has to be balanced against downstream environmental water flow requirements (e.g., King et al. 2003), to reduce the risk for unwanted downstream impacts (e.g. Calder 1999). Today, there is a lack of research describing the downstream consequences of upstream implementation of agricultural water management interventions (Karlberg et al. 2009).
Land-use change is also a decision about water (Fig. 3). It is estimated that the land-use conversions to agriculture that took place during the last 300 years have resulted in a decrease in green water flow and an increase in recharge and stream-flow (Scanlon et al. 2007). A study in West Africa showed that clear-cutting of tropical forests increases the annual stream-flow by 35-65%, depending on the basin, although forests occupy less than 5% of the total basin area (Li et al. 2007).
A relatively poorly understood area is how vegetation impacts on rainfall amounts via moisture feed backs (Savenije 1996). It seems that systems partly shape the micro-climate that they exist in, and a change in land-use system could therefore also impact on the local climate of the region. In order to develop relevant policy for efficient green water use in agriculture, there is a need for IWLRM to develop into more holistic assessments of blue and green water flows at the landscape scale.
Several agricultural water interventions focus specifically on green water management and have been shown to significantly improve crop yields (Fig. 4). Rockstro ¨m 0 50 100 150 200 250 E t h i o p i a K e n y a ( M a c h a k o s ) K e n y a ( R a c h u n y o ) T a n z a n i a Z a m b i a S y r i a ( l o w p r e c . ) S y r i a ( h i g h p r e c . ) K e n y a B u r k i n a F a s o I n d i a Crop yield improvements (%) Non fertilised Fertilised CA WH WSD Fig. 4 Crop yield improvements at different locations, ranging from conservation agriculture (CA), ex-situ water harvesting systems (WH) and watershed development programmes (WSD), (i.e., programmes that combine conservation agriculture with supplementary irrigation). Sources: Conservation Agriculture: Rockstro ¨m et al. (2009d). Water harvesting systems: Syria, Oweis (1997); Kenya, Barron and Okwach (2005); Burkina Faso, Fox and Rockstro ¨m (2003). Watershed development programmes (one example): Wani et al. (2008). Prec. = precipitation (2003) showed that in low-yielding agriculture (i.e., yields below 3 ton ha -1 ), yield improvements also lead to subsequent improvements in water productivity due to larger soil surface coverage. Conservation agriculture is a type of soil and water conservation system which replaces conventional ploughing with lower intervention practices for soil management, and which combined with mulch management, builds organic matter and improves soil structure (Derpsch 1998;Landers et al. 2001). Apart from improving the water holding capacity of the soil, the increased organic matter content in conservation agriculture systems also result in more carbon being stored in the system. Supplementary irrigation with locally collected run-off, so called ex-situ water harvesting systems, is used to bridge dryspells, which frequent the tropical drylands (Siegert 1994;Fox and Rockstro ¨m 2000). When these water interventions are combined with fertilisation, the resulting yield and water productivity is even higher (Fig. 4). Productive sanitation systems, i.e., the safe reuse of human urine and faeces as a fertiliser for increased food production, could be combined with agricultural water management interventions to improve crop yields in a sustainable way by reducing the nutrient losses of the food production-human consumption chain.
A triply green revolution of agriculture has to be accomplished within a social-ecological framework, integrating agricultural management with stewardship of landscape capacity to generate ecosystem functions and services. This requires special attention to cross-scale effects, thresholds, and feedback interactions, in order to develop resilient, multi-functional landscapes for the generation of food and other ecosystem services. New innovative approaches to achieve this goal are urgently needed.
Placing the global freshwater challenge within the context of global food security and the impacts of accelerated global environmental change, raises the urgency of developing strategies to build resilience in water resource governance and management. Taking a social-ecological and resilience perspective on the challenge of human development in the Anthropocene, indicates that the water and food nexus is subject to a ''quadruple squeeze'' from demographic pressure, the global climate crisis, the global ecosystem crisis, and the growing insight of the universality of non-linear dynamics in ecosystem change.
As a consequence of the growing social-ecological pressure on several key earth system processes, nothing less than a triply green revolution will be needed to produce food for a growing world population. Food production will have to increase at record pace, which can only be accomplished through major investments in both green and blue water resource management, and will require a sustainability transformation that enables global agriculture to produce food within the safe operating space of the planetary boundaries.
Farming systems in the world are not configured to deliver a triply green revolution. Neither ecological nor conventional agricultural systems fulfil the green criteria suggested in this article. What is urgently needed is a new definition of sustainable agricultural systems as well as new, innovative technologies that enable major productivity improvements. Elements of knowledge to develop these triply green systems are presented in this articleparticularly the opportunities of major improvements in water productivity and yield increase through various agricultural water interventions. A major research initiative is needed to develop, test and promote agricultural systems that contribute to a sustainable future for human-kind in the long-term.
Ó Royal Swedish Academy of Sciences 2010 www.kva.se/en
The pleiotropic effects of creatine (Cr) are based mostly on the functions of the enzyme creatine kinase (CK) and its high-energy product phosphocreatine (PCr). Multidisciplinary studies have established molecular, cellular, organ and somatic functions of the CK/PCr system, in particular for cells and tissues with high and intermittent energy fluctuations. These studies include tissue-specific expression and subcellular localization of CK isoforms, high-resolution molecular structures and structure-function relationships, transgenic CK abrogation and reverse genetic approaches. Three energy-related physiological principles emerge, namely that the CK/PCr systems functions as (a) an immediately available temporal energy buffer, (b) a spatial energy buffer or intracellular energy transport system (the CK/PCr energy shuttle or circuit) and (c) a metabolic regulator. The CK/PCr energy shuttle connects sites of ATP production (glycolysis and mitochondrial oxidative phosphorylation) with subcellular sites of ATP utilization (ATPases). Thus, diffusion limitations of ADP and ATP are overcome by PCr/Cr shuttling, as most clearly seen in polar cells such as spermatozoa, retina photoreceptor cells and sensory hair bundles of the inner ear. The CK/PCr system relies on the close exchange of substrates and products between CK isoforms and ATPgenerating or -consuming processes. Mitochondrial CK in the mitochondrial outer compartment, for example, is tightly coupled to ATP export via adenine nucleotide transporter or carrier (ANT) and thus ATP-synthesis and respiratory chain activity, releasing PCr into the cytosol. This coupling also reduces formation of reactive oxygen species (ROS) and inhibits mitochondrial permeability transition, an early event in apoptosis. Cr itself may also act as a direct and/or indirect anti-oxidant, while PCr can interact with and protect cellular membranes. Collectively, these factors may well explain the beneficial effects of Cr supplementation. The stimulating effects of Cr for muscle and bone growth and maintenance, and especially in neuroprotection, are now recognized and the first clinical studies are underway. Novel socio-economically relevant applications of Cr supplementation are emerging, e.g. for senior people, intensive care units and dialysis patients, who are notoriously Cr-depleted. Also, Cr will likely be beneficial for the healthy development of premature infants, who after separation from the placenta depend on external Cr. Cr supplementation of pregnant and lactating women, as well as of babies and infants are likely to be of benefit for child development. Last but not least, Cr harbours a global ecological potential as an additive for animal feed, replacing meat-and fish meal for animal (poultry and swine) and fish aqua farming. This may help to alleviate human starvation and at the same time prevent overfishing of oceans.
Creatine (Cr) has emerged as a safe nutritional supplement not only to increase muscle mass and performance, prevent disease-induced muscle atrophy and improve rehabilitation, but also to strengthen cellular energetics in general (see Salomons and Wyss 2007). The latter represents the physiological basis for the beneficial effects of Cr supplementation in the treatment of multiple pathologies that display bioenergetic dysregulation, such as myopathies or neurodegenerative diseases (see Andres et al. 2008). Although Cr effects are likely due to pleiotropic cellular functions, its main role is in the creatine kinase (CK/PCr) system for temporal and spatial energy buffering. Interdisciplinary approaches in the frame of a system bioenergetics have been successfully applied in the past and will further be necessary to understand the CK/PCr system in more detail (Saks et al. 2006a;Saks 2007). This review first summarizes the fundamental knowledge that has been accumulated on the complex CK/PCr system over the last three decades, and in a second part gives an overview on the pleiotropic effects of Cr related to Cr supplementation as an adjuvant therapy in various pathologies. Exciting new discoveries related to anti-oxidant and anti-apoptotic effects, as well as protection of membranes by PCr are also discussed with respect to cell protection by Cr.
The CK/PCr system for temporal and spatial buffering and regulation of cellular energetics Creatine and creatine kinase Although ATP represents the universal energy currency in all organisms and cells, ATP levels are not simply upregulated in cells with high and/or intermittently fluctuating energy demand. Elevation of the intracellular ATP concentration [ATP], as an immediate energy reserve, followed by its hydrolysis, would lead to a massive accumulation of ADP plus P i and also liberate H ? , acidifying the cytosol. Since this would inhibit ATPases, such as the myofibrillar acto-myosin ATPase and consequently muscle contraction, and many other cellular processes, nature has evolved a means to deal with the problem of the immediate replenishment of ATP stores. The so-called phosphagens evolved as high-energy compounds that are ''metabolically inert'' and as such do not interfere with primary metabolism. One of them, PCr, together with its corresponding kinase, CK, first appeared at the dawn of eukaryotic evolution some one billion years ago (Bertin et al. 2007).
CK catalyses the reversible reaction: PCr 2À þ MgADP À + H þ / ÀCK À .MgATP 2À + Cr and can thus either utilize PCr (with a higher DG free energy change than ATP) to regenerate ATP or alternatively capture immediately available cellular energy, storing up to 10 times the amount in the ATP pool. The CK system stabilizes cellular [ATP] at approximately 3-6 mM, depending on the cell type, at the expense of [PCr], and thus maintains the intracellular ATP/ADP ratio at a very high level and consequently keeps the DG free energy change of ATP hydrolysis as high as possible. This guarantees an efficient use of ATP for all types of cellular functions; that is, the energy gained per ATP hydrolysed is kept at a physiological maximum. Resynthesis of ATP by the CK reaction, for example upon activation of muscle contraction, also removes ADP and H ? as products of ATP hydrolysis, so that the net product of ATPase plus the CK reaction is liberation of Pi as a metabolic signal. Thus, the CK acts not only acts as an energy buffer but also as a metabolic regulator (for review, see Wallimann et al. 1992Wallimann et al. , 2007)).
CK isoforms and their molecular structure CK, which is crucially involved in a plethora of bioenergetic processes, is particularly important and is expressed at high levels in cells with high energy requirements such as skeletal, cardiac and smooth muscle, kidney, brain and neuronal cells, retina photoreceptor cells, spermatozoa and sensory hair cells of the inner ear (Wallimann et al. 1992(Wallimann et al. , 2007)). The most important feature for the cellular functions of the CK/PCr system is the presence of tissue-and cellspecific CK isoforms with defined subcellular locations. All CK isoforms are encoded by separate nuclear genes and, in most tissues, a single cytosolic CK isoform is coexpressed together with a single mitochondrial CK isoform (mtCK). Cytosolic muscle-type CK (M-CK) and brain-type CK (B-CK) form homo-dimers or hetero-dimers, e.g. MM-CK in skeletal muscle, MM-, MB-and BB-CK in heart, or BB-CK in brain, kidney, spermatozoa, skin and many other tissues. MtCK is situated in the outer mitochondrial compartment and occurs as sarcomeric mtCK (smtCK) expressed mainly in muscle tissue and as ubiquitous mtCK (umtCK) expressed in a large number of other cells and tissues. Both form homo-dimers and homo-octamers, with the latter being the predominant oligomeric form in vivo.
Importantly, CK isoforms are localized differentially on a subcellular level and these specific locations are essential for the functioning of the CK network (Wallimann et al. 1992(Wallimann et al. , 2007)). As for many other cellular processes, ''location, location is the name of the game'' (Hurtley 2009). The proposed CK/PCr energy shuttle (Wallimann 1975;Saks et al. 1978;Bessman and Geiger 1981;Bessman 1986Bessman , 1987;;Wallimann et al. 1992;Schlattner et al. 2006a, b;Wallimann et al. 2007) connects sites of ATP production (glycolysis and mitochondrial oxidative phosphorylation) with subcellular sites of ATP utilization (ATPases). The molecular bases for this spatial energy buffering are functionally coupled, subcellular CK micro-compartments, at sites where ATP production and ATP consumption are tightly connected to CK/PCr action. At these subcellular sites, CK reactions may run in different (forward or backward) directions, but on the global cellular or organ level the CK system appears to be as if in equilibrium (see Fig. 1).
The molecular structures of CK isoforms have been solved at atomic level (Fritz-Wolf et al. 1996;Rao et al. 1998;Eder et al. 1999Eder et al. , 2000a;;reviewed by McLeish and Kenyon 2005) and their biochemical characteristics, e.g. enzyme catalysis, oligomerization, membrane interaction and binding to subcellular structures (e.g. Rossi et al. 1990;Eder et al. 2000b;Hornemann et al. 2000;Schlattner et al. 2000;Schlattner and Wallimann 2000;Schlattner et al. 2004) and specific interaction with subcellular partners and domains involved in such interactions (Kraft et al. 2000;Hornemann et al. 2003) have been studied in detail over the past decades. These studies have revealed important aspects of structure-function relationships and molecular physiology of CK that has allowed an understanding of the CK/PCr system and its eminent physiological role (Schlattner et al. 2006a;Wallimann et al. 2007).
Cr derived either from endogenous synthesis in the body or taken up from alimentary sources, e.g. meat and fish, is transported into muscle and other target cells with high and fluctuating energy requirements by a specific creatine transporter (CRT) (Speer et al. 2004;Straumann et al. 2006; see Fig. 1 with respective labels and numbering used in the following text). Imported Cr is charged to the highenergy compound PCr by the action of either strictly soluble, cytosolic CK (CK-c, (3)), by CK coupled to glycolysis (CK-g, ( 2)) or by mtCK coupled to oxidative phosphorylation (OP) (mtCK, (1)). In a resting cell, this results, at equilibrium, in a distribution of the total Cr pool into approximately two-thirds [PCr] and one-third [Cr] and in a very high ATP/ADP ratio (C100:1). A fraction of cytosolic isoforms of CK are specifically associated with ATP-consuming processes (CK-a, ( 4)), such as the myofibrillar acto-myosin ATPase, the SR Ca 2? -ATPase, the plasma membrane Na ? /K ? -ATPase, the ATP-gated K ?channel or ATP-requiring constituents for cell signalling.
Within these functional micro-compartments, CK regenerates the utilized ATP, drawing from the large PCr pool. These micro compartments with associated CK
CRT 1 2 3 4 L O S O T Y C A I R D N O H C O T I M Cr ATP Cr ATP ATP ATP ATPase ANT G OP a -K C c -K C g -K C mt CK ADP PCr ADP ADP ADP ANT cytosolic ATP/ADP ratio glycolysis oxidative phosphorylation cytosolic ATP-consumption ATP-supply ATP-consumption cytosolic CK (CK-a, see (4)) specifically associated with subcellular sites of ATP utilization (ATPase, e.g. ATP-dependent or ATP-gated processes, ion-pumps etc.) also forms tightly coupled microcompartments regenerating the ATP utilized by the ATPase reaction in situ on the expense of PCr. The proposed CK/PCr energy shuttle or circuit connects, via highly diffusible PCr and Cr, subcellular sites of ATP production (glycolysis and mitochondrial oxidative phosphorylation)
with subcellular sites of ATP utilization (ATPases). This model is based on functionally coupled, subcellular CK microcompartments, where ATP production and ATP consumption are tightly connected to CK/PCr action (Wallimann 1975;Wallimann and Eppenberger 1985;Schlattner et al. 2006a;Wallimann et al. 2007) represent the ATP/PCr-consuming side of the CK/PCr system (4). At the ATP/PCr-generating side of the system, there are the glycogenolytic/glycolytic CK-g microcompartments (2) and the mtCK microcompartment connected to OP and energy channelling reactions inside the mitochondrion (1) (Schlattner et al. 2006b). MtCK is specifically located in the intermembrane space of mitochondria with preferential access to ATP generated by OP via adenine nucleotide translocator (ANT) of the mitochondrial inner membrane. This mitochondrial ATP is trans-phosphorylated into PCr that then leaves the mitochondria (Dolder et al. 2003). This route of ATP generation is most important for refilling the PCr energy store in oxidative tissues, e.g. upon extensive stimulation of muscle contraction and thus is relevant for recovery after exhausting exercise (Quistorff et al. 1993).
A large cytosolic PCr pool of up to 30 mM is built up by CK using ATP predominantly from OP (1) as in the heart, or from glycolysis (2) plus OP (1) as in skeletal muscle. PCr is then used to buffer global cytosolic (3) and local (4) ATP/ ADP ratios. This would represent the temporal buffer function of the system. This function has been confirmed using reverse genetic approaches. For example, by introducing a phosphagen kinase gene, the CK orthologue arginine kinase (AK), into Escherichia coli or yeast, a phospho-arginine pool was built up that improved recovery of the bacteria and yeast from transient pH stress (Canonaco et al. 2002(Canonaco et al. , 2003)). By a similar strategy, yeast cells into which the genes for the enzymes for Cr biosynthesis (AGAT and GAMT) plus CK were introduced, proved to be resistant to metabolic stressors, such as low pH and starvation, by stabilizing the ATP levels during the transient stress period to pre-stress levels (Canonaco et al. 2002).
In cells that are polarized and/or have very high or localized ATP consumption (4), the differentially localized CK isoforms, together with easily diffusible PCr and Cr, maintain a high-energy PCr/Cr-circuit between ATP-providing (1, 2) and ATP-consuming processes (3, 4). Thus, the energy producing and consuming terminals of the shuttle are connected via PCr and Cr, with no obligatory need for ATP or ADP to diffuse, for example, from mitochondria (1) to the sites of ATPases (4) or backwards, respectively. Metabolite channelling (Schlattner et al. 2011) occurs where CK is associated with ATP-providing (1, 2) or ATP-consuming processes (4), that may be represented by ATPases, such as the actin-activated myosin MgATPase for muscle contraction (Wallimann et al. 1984;Ventura-Clapier et al. 1987;van Deursen et al. 1993) and actin-based cell motility (Kuiper et al. 2009), ATPdependent ion-pumps and transporters, such as the SR Ca 2? pump (Rossi et al. 1990;de Groof et al. 2002), the Na ? /K ? -ATPase (Guerrero et al. 1997), the gastric H ? /K ? -ATPase (Sistermans et al. 1995b), as well ATP-gated ion-channels (Dzeja and Terzic 1998), ATP-requiring metabolic enzymes and protein kinases, such as AMP-activated protein kinase (AMPK) (Ceddia and Sweeney 2004) or Akt/PKB (Deldicque et al. 2007(Deldicque et al. , 2008) ) involved in cell signalling (Saks et al. 2006b;Wallimann et al. 2007).
CK in specialized polarized and epithelial cells CK is not only prominent in sarcomeric skeletal and cardiac muscle, where MM-CK is co-expressed with smtCK and where these isoforms are located at specific subcellular sites (Wegmann et al. 1992), but also in smooth muscle, brain and other non-muscle tissues where BB-CK is coexpressed with umtCK (Wallimann et al. 1992). For example, BB-CK is highly expressed in spermatozoa, retina photoreceptor cells (Wallimann and Hemmer 1994) as well as in the sensory hair cells present in the inner ear (Shin et al. 2007). The common denominator of these cells is that they are highly polar, elongated cells and that mitochondria are located at a distance from intracellular sites of ATP consumption. Therefore, these cells are the best models to investigate how mitochondrial-generated high-energy phosphates reach the sites of ATP consumption when separated by a long diffusion distance.
In sea urchin sperm, 100% of the energy required for sperm tail movement is generated in one single large mitochondrion located in the mid-piece behind the sperm head (Fig. 2). Mitochondria from sea urchins and other organisms harbour high concentrations of octameric mtCK (Wallimann et al. 1986a;Tombes and Shapiro 1987;Kaldis et al. 1996b), whereas the sperm tails contain a ''cytosolic'' CK which in sea urchins is a contiguous CK trimer (Quest et al. 1997), but in other organisms consists of brain-type BB-CK dimers (Wallimann et al. 1986a;Kaldis et al. 1996b). Most of the sperm tail CK is distributed along the entire sperm tail and/or associated with the cell membrane and the dynein ATPase (Quest et al. 1997).
In vivo experiments examining flagellar wave bending of living sea urchin sperm, which after activation were incubated with dinitrofluorobenzene (DNFB), a specific inhibitor of CK, showed that increasing concentrations of DNFB attenuated firstly flagellar wave bending, and the amplitude of bending, at the very distal end of the sperm tail. As DNFB concentrations were increased, flagellar wave attenuation was gradually and progressively affected towards the more proximal regions and the mid-piece (Tombes et al. 1987). Similar results had been obtained with chicken sperm (Wallimann et al. 1986a). This was a first and very elegant visual demonstration showing that with progressive inhibition of CK, the diffusion distances for ATP from the mitochondrion to the very end of the sperm tail are indeed limited and that after inhibition of tail CK they are no longer compensated for by PCr/Cr diffusion.
Using in vivo 31 P-NMR saturation transfer, high-energy phosphate concentrations as well as the rate of flux through the CK reaction were measured before and after activation of intact live sea urchin sperm in sea water (van Dorsten et al. 1996(van Dorsten et al. , 1997)), combined with concomitant measurement of O 2 consumption (ATP production). Knowledge of the metabolite concentrations together with their diffusion constants then allowed the calculation of diffusion fluxes of the respective metabolites (Kaldis et al. 1997). As shown in Fig. 2, PCr and Cr display a significantly higher diffusion flux compared with ATP and ADP. Most remarkably, the diffusion flux of ADP from the distal sperm tail end back to the mid-piece mitochondrion is slower by three orders of magnitude compared with Cr (Kaldis et al. 1997). Thus, the CK/PCr shuttle is compensating for the diffusion limitations of mainly ADP and somewhat less so of ATP.
Similar findings on the importance of the CK/PCr system working as a spatial buffer have recently been presented for elongated photoreceptor cells of the retina (Linton et al. 2010). CK isoforms are expressed in all cells of the retina, but the highest levels of CK are found in the polar photoreceptor cells (Wallimann et al. 1986b), where cytosolic BB-CK and umtCK are located in the inner segments and some BB-CK in the outer segments of the photoreceptor cells of certain species, e.g. in bovine rod outer segments (Wegmann et al. 1991;Hemmer et al. 1993). It was postulated that this compartmented intracellular distribution of CK isoforms could induce an energy flux carried by the CK/PCr system. It would emanate from the mitochondria clustered centrally within the photoreceptor cells and propagate in two directions, towards the synapse as well as into the outer segment towards the photoactive membrane stacks in the rod outer segments (Wallimann et al. 1986b).
Most recently, it has been shown by biochemical and electrophysiological measurements that the CK/PCr system is indeed fundamental to energy distribution in photoreceptors. In darkness, PCr emanating from the central mitochondria of the photoreceptor cells flows into the synaptic terminal, where the ATP required for sustained glutamate release (dark current) is regenerated by cytosolic CK that is localized at this shuttle terminal (Linton et al. 2010). Since we found BB-CK, albeit at lower levels, also in the bovine rod outer segments (Hemmer et al. 1993), it is conceivable that such a vectorial PCr energy transport not head tail acrosome nucleus mitochondrion axoneme with dynein motor protein 5 µm proximal mitochondrial diffusion fluxes distal dynein ATPase [ n o i t c u d o r p P T A µmol min -1 g -1 ] ATP consumption 3 6 µmol min -1 g -1 3,6 µmol min g P i P i ADP ADP 2,4 0,008 r C r C r C P r C P 23,4 13,4 K C K C t m P T A P T A 1,8 (Kaldis et al. 1997). The diffusion flux of ADP from the sperm tail end towards the mitochondrion located at the mid piece is more than 2,000 times slower compared with that of Cr, whereas the diffusion flux of ATP from the region of the mitochondrion towards the sperm tail is roughly seven times slower than that of PCr. Thus, the PCr/ Cr-shuttle is a physiological adaptation to overcome the diffusional limitations of adenosine nucleotides, especially of ADP, to facilitate long-distance energy transport, as well as high-throughput fluxes of cellular energy. A similar system has been proposed and proved to work also in the sensory hair cells of the inner ear (Shin et al. 2007) and in the polar photoreceptor cells of the retina (Wallimann et al. 1986b;Hemmer et al. 1993;Linton et al. 2010) and in the sensory hair cells of the inner ear (Shin et al. 2007). (Figure adapted from Kaldis et al. 1997) only runs in one direction towards the synaptic terminal, but also in the opposite direction into the rod outer segments. There, bound CK would convert PCr into ATP used for photoreceptor signalling, that is, for stabilizing the [ATP] needed for resynthesis of cGMP from GTP upon photic stimulation (Hemmer et al. 1993).
Cr supply to the retina is important and seems to be supported by a dual system: (a) by uptake from the circulation via CRT expressed in the endothelial cells of the blood/retinal barrier, and (b) by endogenous Cr synthesis in the Mu ¨ller glia cells (Tachikawa et al. 2007). Chorioretinal degeneration in patients with gyrate atrophy, who present with a Cr-deficiency syndrome, can be ascribed both to a reduced Cr supply from circulating blood and a disrupted endogenous Cr supply to the retina from the local Mu ¨ller glia due to inhibition of Cr biosynthesis by hyper-ornithinemia (Sipila 1980). This indicates that the CK/PCr system is physiologically important for vision, even though up to date no phenotype for visual defects has been reported in transgenic CK knockout mice (see below).
Brain-type cytosolic BB-CK has been also localized in the inner ear hair cells (Spicer and Schulte 1992). CK was identified, by a proteomic approach with isolated hair cells from purified vestibular bundles, as the second most prominent protein besides actin, and other proteins of the cytoskeleton and proteins involved in Ca 2? homeostasis, stress response and glycolysis (Shin et al. 2007). Present at a high concentration of approximately 0.5 mM inside the sensory hair cells, the CK enzyme is capable of maintaining constant ATP levels despite a turnover of 1 mM ATP/s. This turnover is imposed by the plasma membrane Ca 2? -ATPase pump that maintains Ca 2? -cycling during activation of these specialized sensory cells. The polarized hair bundle cells cannot rely on ATP diffusion and it was shown that the CK/PCr shuttle is essential for high-sensitivity hearing and vestibular function, e.g. body balance and equilibrium. CK knockout mice presented with hearing loss and a strong vestibular phenotype (Shin et al. 2007). Interesting in this context is the fact that Cr supplementation of healthy wild-type mice significantly attenuates noise-induced destruction of inner and more so of outer hair cells and the concomitant hearing loss (Minami et al. 2007). These data indicate that the maintenance of ATP levels by the CK/PCr system, possibly together with antioxidant properties of Cr, are important for attenuating temporary and permanent noise-induced hearing loss. Thus, Cr supplementation may be recommended as a preventive measure in professionally noise-exposed individuals.
Rather surprisingly, high concentrations of CK isoforms were also found in a variety of epithelial cells that are not known to have high fluctuating energy requirements, but may need energetic support for maintaining high rates of cell divisions, resorption or secretion activities. In stomach epithelium parietal cells, the CK/PCr system appears to work in conjunction with the H ? /K ? -ATPase pump and is involved in gastric acid secretion (Sistermans et al. 1995b). In epithelial cells of the intestine (Sistermans et al. 1995a) CK may be involved in food absorption and transport, as well as cell renewal. In skin, BB-CK and umtCK have been localized in the keratinocytes of the highly proliferative suprabasal layer of the epidermis, as well as in the hair follicles and sebaceous glands (Schlattner et al. 2002), indicating a function of the CK/PCr system for normal skin function (Zemtsov 2007), proliferation and hair growth (Schlattner et al. 2002). During wound healing, CK isoforms were highly upregulated indicating a function of CK and Cr (Schlattner et al. 2002). In accordance with the postulated important functions of the CK system in skin, topical application of a Cr containing lotion directly onto skin has been shown to exert a marked protection from UV-induced oxidative damage and mutagenesis in vitro (Berneburg et al. 2005) and in vivo (Lenz et al. 2005).
There is still a widely expressed concern that Cr supplementation may be deleterious to the kidney. This, unfortunately, has to do with the fact that creatine (Cr) is still mixed-up with creatinine (Crn). Crn is the cyclic degradation product of Cr that is generated by non-enzymatic conversion from Cr, until a roughly 2/3-1/3 chemical equilibrium between Crn and Cr is established. Crn content is measured as a kidney function marker in the serum of patients because it is very prominent and easy to measure chemically. While an accumulation of Crn in the serum normally indicates that kidney function is impaired, this is entirely unrelated to the somewhat higher serum Cr and/or Crn concentrations during Cr supplementation which, in this case, is not indicative of kidney malfunction or any general toxicity. As the Crn concentration may increase somewhat with Cr supplementation (Schedel et al. 1999), due to the chemical equilibrium reaction between Cr and Crn, this is often taken as a false argument for impairment of kidney function, but in fact is a normal consequence of Cr intake. On the contrary, CK and Cr are important for kidney function. CK is highly expressed in kidney epithelial cells (Wallimann and Hemmer 1994) and the CK/PCr system supports Na ? /K ? -ATPase ion pump function in the kidney (Guerrero et al. 1997). In addition, kidney proximal tubule epithelial cells also express Cr transporter (CRT) that is responsible for resorption and salvaging of Cr, a valuable guanidino compound, from the urine (Li et al. 2010). So, one may argue that if Cr would be deleterious for the kidney, why would the kidney absorb Cr from the urine instead of secreting it? Indeed, in a placebo-controlled double-blind clinical study, involving healthy men, Cr supplementation at 10 g/day for 3 months had no deleterious effect on kidney function (Gualano et al. 2008b). In a single case study concerning a man with only one kidney, who presented with a mildly decreased glomerular filtration rate, no Cr-induced deleterious effects of Cr supplementation (20 g/day for 30 days) were observed (Gualano et al. 2010). Even long-term Cr supplementation (4 g/day for 2 years) is today considered safe, as seen in aged Parkinson patients (Bender et al. 2008b). Thus, Cr taken at the recommended dosage in a chemically pure form is not deleterious to kidney function and health.
Functions of cytosolic CK associated with glycolysis A fraction of cytosolic CK is associated with the glycolytic enzyme complex (Fig. 1, (2)) that, in sarcomeric muscle, is concentrated at the myofibrillar I-band (Kraft et al. 2000). It was shown that there, CK is specifically associated with those glycolytic enzymes that are either involved in ATP generation, such as pyruvate kinase (PK) (Dillon and Clark 1990;Kraft et al. 2000) or with the main regulatory enzyme of glycolysis, phosphofructokinase (PFK), which is regulated by rising [ATP] via a negative feed-back mechanism of ATP on PFK (Kraft et al. 2000). The CK-PFK interaction is pH-dependent and stronger at lower pH than at neutrality. This is physiologically relevant since under conditions of muscle activation and working glycolysis, the intracellular pH may drop. Thus, if glycolysis is initiated immediately after contraction to produce ATP, the close structural and functional association of CK with distinct members of the glycolytic micro-compartment makes sense in two ways: (a) glycolytic ATP is immediately removed from this compartment by associated CK and thus inhibition of glycolysis by negative feed-back regulation via ATP accumulation is prevented, and (b) glycolytic ATP can be used concomitantly to reduce the depletion of the PCr pool during contraction (Kraft et al. 2000;Wallimann et al. 2007). The principle of tight functional coupling of glycolysis to CK action has been shown in vitro in a system of reconstituted glycolysis containing CK (Scopes 1973) and also in vivo in an anoxic goldfish model (van den Thillart et al. 1989) where mitochondrial function is basically eliminated due to lack of oxygen. There, clearly, glycolytic ATP was shown to be used to replenish the PCr pool, although not to the maximal extent (van den Thillart et al. 1989;Van Waarde et al. 1990).
The coupling of CK to the glycogenolytic/glycolytic pathway is also supported by gated 31 P-NMR measurements following at very high time-resolution the metabolite fluctuations elicited by single muscle contractions (Chung et al. 1998). After a single contraction the recovery of PCr is much faster than at the end of stimulation. This implies a distinct recovery mechanism in the first phase, which is in line with the contention that a significant proportion of glycogenolytic/glycolytic ATP is immediately used and trans-phosphorylated by CK to replenish PCr (see also Shulman 2005). After the end of stimulation, however, PCr recovery kinetics is consistent with a predominant role of OP (van den Thillart et al. 1989;Van Waarde et al. 1990;Quistorff et al. 1993;Chung et al. 1998). Finally, transgenic CK knock-out mice show an altered glycolytic network in their CK-deficient muscles (de Groof et al. 2001a), which also indicates an intricate interconnection of the two systems. This is corroborated by the fact that muscles of PFK-deficient patients show a dramatic delay in PCr recovery following exercise (Grehl et al. 1998).
Metabolite channelling in the mtCK microcompartment
After the import of nuclear encoded, nascent mtCK through the mitochondrial outer membrane and cleavage of the N-terminal targeting sequence, mtCK first assembles into dimers. These dimers rapidly associate into octamers and although this reaction is reversible, octamer formation is strongly favoured by the high mtCK concentration in the inter-membrane space (Schlattner et al. 2006a). The symmetrical and cube-like mtCK octamers (Fritz-Wolf et al. 1996;Eder et al. 2000a) then directly bind to acidic phospholipids in the mitochondrial membranes (Fig. 3), preferentially to cardiolipin of the inner mitochondrial membrane (Schlattner et al. 2004), and in a calciumdependent manner directly to VDAC (Schlattner et al. 2001). Affinity of both ANT and mtCK for cardiolipin situates them in common cardiolipin patches (Fig. 3), which can also be induced by membrane-bound mtCK (Epand et al. 2007b), thus allowing for a functional interaction between both proteins (Wallimann et al. 1998;Schlattner et al. 2006b). MtCK is found in two locations (Wegmann et al. 1991): (a) in the so-called mitochondrial contact sites (Fig. 3), where mtCK simultaneously binds to inner and outer membrane due to its symmetrical octameric structure (see below) and where it functionally associates with ANT and VDAC (Kottke et al. 1991) and (b) in the cristae (not shown) associated with inner membrane and ANT only (Wegmann et al. 1991; for details, see Schlattner et al. 2006aSchlattner et al. , 2011)). A well-coupled MtCK micro compartment is also maintained by diffusion restrictions for adenylates at the outer membrane VDAC, possibly via direct interaction of VDAC with tubulin (Rostovtseva et al. 2008). The preferred or exclusive substrate and product fluxes in a well-coupled mtCK microcompartment are indicated in Fig. 3 by black arrows (Saks et al. 2000); minor or alternative fluxes at a lower degree of coupling are shown with blue arrows.
According to this scheme, ATP generated by OP via the F 0 F 1 -ATPase is transported through the inner membrane by ANT in exchange for ADP. This ATP may either leave the mitochondrion directly via outer membrane VDAC or is preferentially accepted and trans-phosphorylated into PCr by octameric mtCK in the intermembrane space. PCr then preferentially leaves the mitochondrion via VDAC and feeds into the large cytosolic PCr pool. ADP generated from the mtCK trans-phosphorylation reaction is accepted by ANT and immediately transported back into the matrix to be recharged. In contact sites, this substrate channelling allows for a constant supply of substrates and removal of products at the active sites of mtCK. In cristae, only ATP/ ADP exchange is facilitated through direct channelling to the mtCK active site, while Cr and PCr have to diffuse along the cristae space to reach VDAC (not shown; for mitochondria. The structural basis of these mtCK microcompartments are proteolipid complexes containing either VDAC, octameric mtCK and ANT in the peripheral intermembrane space (as shown) or octameric mtCK and ANT in the cristae (not shown). These proteolipid complexes are maintained by mtCK interactions with anionic phospholipids and VDAC in the outer membrane, and with cardiolipin and thus indirectly with cardiolipin-associated ANT in the inner membrane (see cardiolipoin patches). In cases of a less coupled mtCK microcompartment, e.g. after impairment of mtCK function by oxidative damage, there is partial direct ATP/ADP exchange with the cytosol (blue arrows). (Figure adapted from Kaldis et al. 1997; Meyer et al. 2006; Schlattner et al. 2006a; Schlattner et al. 2011) (The different fluxes are indicated by coloured arrows in the figure)
details, see Schlattner et al. 2006a, b;Wallimann et al. 2007).
Functional coupling of ANT to mtCK leads to a saturation of mtCK with ANT-delivered ATP, and at the same time to a locally high ATP/ADP ratio in the vicinity of mtCK. In combination with cytosolic Cr, entering the intermembrane space via VDAC, these conditions are favourable to drive the synthesis of PCr from ATP by mtCK without loss of energy content. This maintains maximal thermodynamic efficiency for high-energy phosphate synthesis and channelling, which in the form of PCr is exported into the cytosol (Dolder et al. 2003). Such a reaction sequence represents an instructive example of functional coupling and metabolite channelling. The active ATP/ADP exchange maintained by coupled mtCK favours ATP generation by the F 0 F 1 -ATPases and thus proper functioning of the respiratory chain, which could otherwise generate elevated levels of superoxide and reactive oxygen species (ROS) (Schlattner et al. 2011).
The intricate functional coupling of mtCK to the ANT (Vendelin et al. 2004), leading to a saturation of the ANT on the outer site of the inner membrane with ADP, which then is transported back into the matrix to be recharged by the F 0 F 1 -ATPase, efficiently couples substrate oxidation to ATP generation. Such tight coupling, by avoiding futile electron transfer, conceivably also would lower the production of free oxygen radicals (ROS). Indeed, mitochondrial respiration in the presence of Cr needs only micromolar concentrations of ADP to be fully stimulated, whereas in the absence of Cr comparably high concentration of ADP in the millimolar range are needed for a similar respiratory rate (Kay et al. 2000;Saks et al. 2000). This important phenomenon, termed ''creatine-stimulated respiration'' is entirely dependent on the presence of mtCK, for mitochondrial respiration can no longer be stimulated by Cr in intact chemically skinned muscle fibres or mitochondria isolated from cardiac or skeletal muscle of smtCK knockout mice (Kay et al. 2000). Finally, Cr exerts a strong indirect anti-oxidant effect by significantly reducing the intra-mitochondrial production of ROS, as well as elevating and preserving the mitochondrial membrane potential (Meyer et al. 2006). These Cr-meditated events may represent the basis for some of the remarkable neuro-protective effects of Cr that had been discovered recently (reviewed by Andres et al. 2008).
Both, sarcomeric smtCK (Fritz-Wolf et al. 1996) and ubiquitous umtCK (Eder et al. 2000a) show cube-like octameric structures of mtCK with approximately 100 A ˚side lengths, built by four identical mtCK dimers that are arranged by fourfold symmetry around a central channel of approximately 20 A ˚in diameter. These octamers maintain multiple and complex interactions with the phospholipids of the mitochondrial membranes (reviewed by Schlattner et al. 2006b). By their identical top and bottom faces, which each expose four C-termini, MtCK binds strongly to anionic phospholipids, in particular cardiolipin that is abundant in the inner mitochondrial membrane. By virtue of their molecular symmetry, mtCKs are also able to crosslink two membranes. Both membrane-binding and crosslinking characteristics of mtCK have been thoroughly investigated and quantified by a number of biophysical techniques (Rojo et al. 1991;Stachowiak et al. 1996Stachowiak et al. , 1998;;Schlattner et al. 2004). By site-directed mutagenesis, a cluster of positively charged amino acids at the C-termini has been identified as responsible for mtCK's ability to specifically attach to cardiolipin-containing membranes (Schlattner et al. 2004;Schlattner et al. 2006a). Binding of mtCK with mitochondrial membranes takes place in two phases (Schlattner et al. 2004). The first phase of mtCK attachment is mediated by ionic interaction by positively charged amino acid clusters at the C-terminal of mtCK; the second slower phase is mediated by partial insertion of a hydrophobic stretch into the membrane bilayer (Schlattner et al. 2006a;Maniti et al. 2010). The ability of mtCK to bind to and cross-link two membranes explains the contact site formation between mitochondrial inner and outer membranes and the resulting mechanical stabilization of mitochondria as shown with liver mitochondria from transgenic mice expressing mtCK (Speer et al. 2005). Also, the formation of the characteristic crystalline intra-mitochondrial ''railway-track'' inclusions, built of mtCK octamers (Stadhouders et al. 1994) inside of mitochondria of ''ragged red fibres'' from patients with mitochondrial cytopathies, can be explained by membrane binding of mtCK octamers either peripherally between inner and outer membrane or in the cristae between inner membranes. Once bound to membranes, mtCK shows a pronounced tendency to form ordered 2D crystalline arrangements (Schnyder et al. 1994). These resemble the sheet-like crystalline inclusions in such patient's mitochondria (Stadhouders et al. 1994). Interestingly, these pathological intra-mitochondrial mtCK crystals that are formed as a result of a compensatory over-expression reaction to an energy deficit (O'Gorman et al. 1997b), can also be induced after Cr depletion in adult cardiomyocytes by addition of guanidine propionic acid (GPA) (Eppenberger-Eberhardt et al. 1991). These mtCK inclusions, both in Crdepleted cardiomyocytes and in mitochondrial myopathy patients, disappear upon Cr supplementation of the cell culture medium or the patients, respectively (Eppenberger-Eberhardt et al. 1991;Tarnopolsky et al. 2004). Finally, recent data show that mtCK, once bound to cardiolipincontaining membrane vesicles, is able to specifically cluster cardiolipin molecules around its molecular surface (Epand et al. 2007b) and, if cross-linked to a second membrane vesicle, mtCK is able to facilitate lipid exchange between the two membranes (Epand et al. 2007a). This of course seems relevant for the structure and physiology of mitochondrial inner/outer membrane contact sites and the pre-apoptotic process of mitochondrial permeability transition pore (MPTP) function.
Control of mitochondrial permeability transition, stabilization of inner/outer mitochondrial membrane complex and anti-apoptotic effects of creatine As mentioned earlier, mtCK occupies strategically important dual locations in the intermembrane and the cristae space (Kottke et al. 1991). In the periphery of the mitochondrion, mtCK is part of a protein complex that is involved in the so-called mitochondrial permeability transition (MPT) pore complex (O' Gorman et al. 1997a;Dolder et al. 2003). Although the molecular composition of the pore remains an open question, it seems to involve the adenine nucleotide transporter ANT-1 isoform and the voltage-dependent anion carrier VDAC (Zhivotovsky et al. 2009) together with mtCK (Kroemer et al. 2007). MPT represents an early event in apoptosis that often leads to swelling of mitochondria and release of apoptosis-inducing factors that can be initiated experimentally by exposure of mitochondria to atractyloside, an inhibitor of ANT, and/or by elevation of extra-mitochondrial [Ca 2? ] (Azzolin et al. 2010). Under these conditions, isolated mitochondria from the liver of normal wild-type mice, which do not express mtCK in the liver, undergo swelling and apoptosis. This process can be inhibited by cyclosporine, a potent antiapoptotic drug. On the other hand, transgenic mice that have been engineered to express mtCK in their liver are protected from apoptosis and its destructive consequences by simple addition of Cr or its analogue cyclocreatine that is also phosphorylated by the CK reaction (Fig. 4). The extent of protection by Cr is comparable to that of cyclosporine (Dolder et al. 2003). Thus, Cr is not only involved in stimulating mitochondrial respiration but also works as an effective mitochondrial protectant and anti-apoptotic compound (Brdiczka et al. 2006). This may explain some of the cell-protective effects observed with Cr. For example, transgenic mice expressing mtCK in their liver acquire, after Cr supplementation, a remarkable tolerance against hypoxia (Miller et al. 1993) and liver toxins (Hatano et al. 1996), as well as against tumour necrosis factor-induced apoptosis (Hatano et al. 2004). Since PCr has been shown to bind to and protect biological membranes (Saks et al. 1996;Tokarska-Schlattner et al. 2003, 2005a), it is conceivable that the PCr generated by mtCK in the mitochondrial intermembrane space would also bind to mitochondrial membranes and stabilize them against swelling, as this was shown for plasma membranes of erythrocytes (Tokarska-Schlattner et al. 2003, 2005a). Thus, mtCK plus Cr seem to exert cell protection not only by improving cellular energetics, but also by more or less energy-independent actions that also affect apoptosis (O'Gorman et al. 1997a;Brdiczka et al. 2006) (see Table 1).
Phenotypes of CK knockout mice point to important cellular functions of CK and the PCr/Cr system Ablation of a given gene in transgenic knockout animals is a valuable tool to possibly evaluate the functions of this specific gene, or the respective protein coded by this gene, in an animal. The various constitutive CK knockouts, engineered by the Wieringa Group in Nijmegen NL, illustrates the phenotypical defects caused by the deletion of CK. Ablation of smtCK in muscle is a good example for a positive identification of a CK function. The fact that mtCK is required for stimulation of mitochondrial respiration by Cr (Saks et al. 2000) has been unambiguously corroborated by this technology with mtCK-deficient or double CK knockout mice (Kay et al. 2000).
Gene ablation may lead to complex adaptations in the organism to compensate for the loss of function related to the knocked out gene. Some very interesting compensatory events take place in the absence of CK function. In the Fig. 4 Mitochondrial permeability transition is inhibited by CK substrates. At a concentration of 10 mM, the CK substrates creatine (Cr) and cyclocreatine (cCr), Cr-analogon, inhibit MTP in isolated mouse liver mitochondria to a comparable degree as 1 lM cyclosporin A (CSA), the gold standard for MPT inhibition (Dolder et al. 2003). The Cr-analogon guanidinopropionic acid (GPA) that is not accepted as a substrate by the CK reaction has no effect as compared to control without additions (none). Isolated liver mitochondria from transgenic mice expressing uMtCK in their liver were analysed by a swelling (light scattering) assay. They were energized with glutamate/ malate in presence of 2mM Mg 2? and then challenged by 120 lM constitutive CK knockouts, this may lead to a physiological and phenotypical amelioration of the phenotype, thus often hindering the phenotypic expression of a dysfunction related to CK deletion. As an example for phenotypic compensation, knocking out of CK in muscle leads to marked changes in mRNA expression profiles involving nuclear and mitochondrial mRNA species that are relevant for bioenergetics (de Groof et al. 2001b), as well as to altered expression of proteins involved in the glycolytic network and mitochondria (de Groof et al. 2001a). Double knock-out of cytosolic and smtCK in muscle also leads to remarkable compensatory adaptation of muscle structure and metabolism, e.g. in white, glycolytic Type II muscle fibres, which normally do not contain large numbers of mitochondria, the CK double knockout animals show a vastly increased mitochondrial propensity and a positioning of these large numbers of mitochondria in such a way that each myofibril is almost completely surrounded by contiguous rows of mitochondria. Thus, these transgenic ''glycolytic'' muscle fibres look rather like entirely oxidative insect flight muscle (Veksler et al. 1995;Ventura-Clapier et al. 1995;Steeghs et al. 1998;Ventura-Clapier et al. 2004;Novotova et al. 2006). This points to a compensation in the CK knockouts for reducing diffusion distances for ATP from the mitochondria to the contractile apparatus and thus would support the proposed function of the PCr shuttle that normally compensates for the diffusion limitations of adenine nucleotides via shuttling of PCr and Cr. This notion is fully supported by detailed analysis of energy provision in CK knockouts, e.g. by the propinquity of mitochondria to myofibrils enabling ATP/ADP to be channelled directly from mitochondria to myofibrils and back (Kaasik et al. 2003). In addition, the glyocogen content in these muscles is elevated indicating that instead of PCr, glycogen/glucose is taken as a more or less immediate source of energy for muscle contraction. Such
Table 1 Pleiotropic effects of creatine for cell function and cell protection Energy-related effects of creatine
Cr improves cellular energy state (PCr/ATP ratio) and muscle performance (Harris et al. 1992;Greenhaff et al. 1993) Cr facilitates intracellular energy transport (PCr circuit or shuttle) (Wallimann 1975;Saks et al. 1978Saks et al. , 2006b;;Bessman and Geiger 1981;Wallimann and Eppenberger 1985;Bessman 1986;Wallimann et al. 1992Wallimann et al. , 1998;;Kaasik et al. 2003;Wallimann et al. 2007) Cr improves the efficiency of cellular energy utilization (e.g. for Ca 2? -handling) (Rossi et al. 1990;Steeghs et al. 1997;Pulido et al. 1998;van Leemputte et al. 1999) Cr stimulates mitochondrial respiration (improved energy provision) (Kay et al. 2000;Meyer et al. 2006) Cr stabilizes mitochondrial PTP complex and thus acts as mitochondrial protectant (anti-apoptotic) (O'Gorman et al. 1997a;Dolder et al. 2003;Hatano et al. 2004) Anti-oxidant and anti-apoptotic effects of creatine Cr acts as a mild direct anti-oxidant (free radical scavenger) (Lawler et al. 2002) Cr acts as a strong indirect anti-oxidant in mitochondria (where ROS production is lowered by tight coupling of respiration/ATP production to ATP export) (Meyer et al. 2006;Sestili et al. 2006) Cr reduces oxidative damage to DNA, specifically to mtDNA (Guidi et al. 2008) Cr up-regulates enzymes for oxidative stress defence (Young et al. 2010) Cr strongly protects in vivo from mitochondrial toxins (Rotenone & Paraquat) (Hosamani et al. 2010) Cr stabilizes mitochondrial PTP complex and thus acts as mitochondrial protectant (anti-apoptotic) (O'Gorman et al. 1997a;Dolder et al. 2003;Hatano et al. ) Other effects of creatine
Cr induces differential expression of transcription factors and other genes (Hespel et al. 2001;Louis et al. 2004;Deldicque et al. 2008;Safdar et al. 2008) Cr reduces the appearance of inflammation markers during endurance exercise (Santos et al. 2004;Bassit et al. 2008) Cr activates cell signalling and enhances muscle cell differentiation (Ceddia and Sweeney 2004;Louis et al. 2004;Deldicque et al. 2007Deldicque et al. , 2008) ) Cr lowers homocysteine levels and lipid peroxidation (heart risk factors) (Deminice et al. 2009) Cr acts as an osmolyte, protecting cells against hypertonic stress (Alfieri et al. 2006) PCr binds to cell membranes and stabilizes and protects erythrocyte cell membranes (Saks et al. 1996;Tokarska-Schlattner et al. 2003, 2005a) interesting compensatory alterations give new insight into the kinds of problems that may have been generated in a given tissue by knocking out of either cytosolic and/or mitochondrial CK. A further interesting observation relates to the fact that CK exists as isoforms and that in a given cell usually a cytosolic CK isoform is co-expressed with a mitochondrial mtCK isoform, although the relative proportion may vary depending on cell type and organ (Wallimann and Hemmer 1994). After knocking out one CK isoform only the other CK isoform can at least partially compensate for the function of the other, e.g. cytosolic CK can partially compensate for mtCK (Watchko et al. 2000).
The most obvious CK knockout phenotypes in muscle relate (a) to force development and maintenance, as well as force-velocity relationship, and (b) to muscle relaxation and Ca 2? -handling, as well as to (c) CK-mediated membrane metabolic sensing. As to the first, transgenic mice that are completely deficient in muscle CK lack burst activity (van Deursen et al. 1993). The velocity and extent of muscle shortening, power and work after the initial series of stimuli are also significantly lower in the CK knockouts compared to wild type (Watchko et al. 2000). In transgenic mice with graded reduction of CK, muscle burst activity is reduced proportionally to the lowered levels of CK expression (van Deursen et al. 1994). These data are in line with findings that MM-CK is specifically localized at the M-band of sarcomeric muscle, where it regenerates in situ the ATP used for muscle contraction (Saks et al. 1984;Wallimann et al. 1984). In non-muscle cells, ablation of BB-CK leads to altered actin-based phagocytosis (Kuiper et al. 2008) and cell motility of cells cultured from CK knockout animals (Kuiper et al. 2009), indicating that CK is not only important for muscle contraction but also for phagocytosis and cell motility in general. As to the second, in CK knockout muscle, muscle relaxation time was longer, with changes also in intracellular Ca 2? -handling in transgenic muscle cells (Steeghs et al. 1997). In line with earlier findings that CK is crucially involved in local ATP regeneration in the vicinity of the SR Ca 2? -ATPase pump (Rossi et al. 1990), it was shown with CK knockout mice that the CK system is indeed essential for optimal refill of the SR Ca 2? store in skeletal muscle (de Groof et al. 2002). According to more recent data, CK, however, is not only important for Ca 2? cycling in muscle, but also in the brain, as shown with brain CK knockouts, where brain Ca 2? kinetics were affected (Streijger et al. 2010). Interestingly, Cr supplementation of myogenic cells from mdx dystrophic mice improves intracellular Ca 2? handling (Pulido et al. 1998) and Cr supplementation of athletes results in shortening of muscle relaxation times in vivo, presumably by improving SR-Ca 2? -pump function and intracellular Ca 2? handling (van Leemputte et al. 1999). The fact that elevating total Cr concentration in muscle by Cr supplementation of human subjects leads to an increase in muscle force and to faster muscle relaxation and recovery after exhaustive exercise, compared with non-supplemented subjects (see below), is fully in line with the described functions of the CK isoforms. In addition, these data indicate that by Cr supplementation, the efficiency of the respective subcellular CK micro-compartments can be improved via elevation of the PCr pool size. Finally, as to the third, by transgenic deletion of cytosolic MM-CK the observed integrative signalling through CK, where cellular energetics is coupled to membrane metabolic sensing, is lost (Abraham et al. 2002). This would corroborate the importance of CK for metabolic sensing and signalling at the plasma membrane.
Besides being expressed in all brain cells, real ''hot spots'' of CK expression and localization are Bergman glia and Purkinje cells in the cerebellum that are important for movement coordination and control, as well as neuronal cells in the hippocampus, where learning and memory functions reside, and finally epithelial cells in the choroid plexus that are rich in ATP-dependent pumps for homeostasis of ions and metabolites between the ventricular fluid/ brain interface (Wallimann and Hemmer 1994;Kaldis et al. 1996a). Accordingly, brain CK knockouts that present with permanently reduced body weight, as well as with altered brain morphology, display altered behaviour, e.g. low nestbuilding activity, less exploratory activity, less grooming, etc. and neurological difficulties in spatial learning and memory functions (Jost CR et al. 2002;in 't Zandt et al. 2004;Streijger et al. 2005).
Recent data also show that the same animals present with problems concerning thermoregulation, eventually succumbing to a sudden and severe drop in body temperature (Streijger et al. 2009). With respect to the involvement of CK in thermoregulation, it is interesting to observe that Cr supplementation in endurance athletes improved their performance during exhausting exercise under hot conditions. That is, in responders, whose muscle total Cr increased during supplementation, rectal temperature and heart rate lowered and peripheral key modulators and indices for the brain neurotransmitters, serotonin and dopamine, were influenced. The subjects in the Cr group reacted with reduced effort perception and completed the endurance task more easily compared to controls (Hadjicharalambous et al. 2008). At the same time, Cr reduced inflammatory and muscle soreness markers after a 30 km race (Santos et al. 2004).
Ablation of the genes for brain-type B-CK and ubiquitous mtCK in mice also leads to frequency-dependent hearing loss and problems with vestibular functions (Shin et al. 2007), which is in line with the very high concentrations of CK found in the respective cellular structures of the inner ear. Interestingly, Cr supplementation significantly attenuated noise-induced hearing loss (Minami et al. 2007). Finally, brain CK knockout has demonstrated the importance of CK for the energetics of bone metabolism and osteoclast function for bone resorption (Chang et al. 2008). This complements earlier results concerning the expression of CK in osteoblasts and the beneficial action of Cr on survival, differentiation and mineralization of osteoblasts in culture (Gerber et al. 2005), as well as with the stimulating effects of Cr on collagen type I synthesis and osteoprotegerin secretion of healthy and osteoporotic human osteoblasts (Gerber et al. 2008).
Thus, it seems obvious that CK takes over specific functions in almost every cell of the body, except for liver, where under normal healthy conditions no CK is expressed (Wallimann and Hemmer 1994). Surprisingly, after the first brain CK knockout transgenic mice became available, it took almost 10 years to figure out some of the most prominent phenotypes of this type of transgenic mice. Probably, it will take another decade still to discover the more subtle phenotypic changes, gone unnoticed, which are caused by brain-type CK ablation.
Pleiotropic effects of creatine Abrogation of CK enzymes in transgenic mice (see above) or depletion of the substrate Cr in Cr-analogue-fed (GPA) animals, respectively, both show muscle phenotypes with similar functional deficits (Mekhfi et al. 1990;Wyss and Wallimann 1994;O'Gorman et al. 1996O'Gorman et al. , 1997b;;Steeghs et al. 1997Steeghs et al. , 1998)). The affected functions, such as the development of muscle force and muscle relaxation, including intracellular Ca 2? handling, can be enhanced in wild-type animals, as well as in humans, by Cr supplementation (Kraemer and Volek 1999;van Leemputte et al. 1999). These mostly energy-related ergogenic effects of Cr in sports, based on the seminal work by Harris and Greenhaff in the early nineties (Harris et al. 1992;Greenhaff et al. 1993), are well known and in the meantime widely accepted (Kamber et al. 1999). The same holds true for the effects of Cr for rehabilitation (Hespel et al. 2001;Johnston et al. 2009) (Table 1). However, a number of potentially beneficial effects of Cr, which are not directly related to enhancement of cellular energetics, have emerged, for example the protective effects of Cr on mitochondrial permeability transition pore opening (O'Gorman et al. 1997a;Dolder et al. 2003), an early event in apoptosis, or the antioxidant effects of Cr, as well as the interference of Cr with cell signalling affecting the expression of muscle transcriptional factors (Hespel et al. 2001;Hespel and Derave 2007;Deldicque et al. 2008) or activating important signalling pathways such a p38 Akt/PKB (Hespel et al. 2001;Deldicque et al. 2007;Hespel and Derave 2007;Deldicque et al. 2008) or AMPK (Ceddia and Sweeney 2004) (Table 1).
Such beneficial effects of Cr may also alleviate toxic drug effects that target bioenergetics and mitochondria, such as the anti-cancer drug doxorubicin (Tokarska-Schlattner et al. 2002, 2006). Doxorubicin accumulates in mitochondria and affects their functions including inhibition of CK isoforms (Tokarska-Schlattner et al. 2002;Tokarska-Schlattner et al. 2005b, 2007). In an animal study, Cr supplementation in combination with vitamins was able to increase survival of doxorubicin-treated rats (Santos et al. 2007).
The pleiotropic effects of Cr on muscle growth and muscle performance have been documented in more than 400 publications to date. Cr has a scientifically unambiguously proven record of being a truly ergogenic nutritional supplement that reaches the target organs, elevates muscle total Cr and PCr pools, leads to an increase in muscle mass and elevates muscle performance in a number of sports (for reviews, see the position stands of the International Society of Sports Nutrition: Buford et al. 2007;Kerksick et al. 2008). The effects of Cr are most beneficial for high-intensity intermittent exercise (Kraemer and Volek 1999) but positive effects of Cr have also been noted for better fatigue resistance (Rawson et al. 2011) and for improved recovery after heavy exercise (Yquel et al. 2002). It has also been realized that Cr could alleviate or spare muscle damage and inflammation caused by excessive endurance performance experienced in an ironman competition (Bassit et al. 2008; see also Table 1). It is important to note that the mostly anecdotal side effects of Cr supplementation that are reported, e.g. via internet can be dismissed on the basis of solid scientific evidence, even if Cr is taken for extended periods of time (i.e. years), (Kreider et al. 2003;Francaux and Poortmans 2006;Persky and Rawson 2007;Bender et al. 2008a). However, the most relevant issue with respect to potential side effects of Cr, namely the chemical purity of the Cr used, is definitely an issue (Pischel and Gastner 2007). Many of the pleiotropic effects of Cr for sports, health and disease are discussed in detail in this volume. Here, we would like to point out some potentially important new applications of Cr supplementation and their potential socio-economic implications for humans and for the global ecosystem.
Creatine supplementation for normal healthy people?
The protective effects of Cr as an adjuvant therapeutic intervention in disease states, such as neuromuscular, neuro-degenerative diseases, as well as muscle-and neurorehabilitation, have been recently reviewed in a special volume of ''Subcellular Biochemistry'' (''Creatine and Creatine Kinase in Health and Disease'', edited by G. Salomons and M. Wyss). In particular, the neuroprotective role of Cr that is relevant to a number of neuromuscular and neuro-degenerative diseases is well documented (Matthews et al. 1998;Klivenyi et al. 1999;Brewer and Wallimann 2000;Wyss and Kaddurah-Daouk 2000;Baker and Tarnopolsky 2003;Andres et al. 2005a, b;Brosnan and Brosnan 2007;Rodriguez et al. 2007;Tarnopolsky 2007;Adhihetty and Beal 2008;Andres et al. 2008;Valastro et al. 2009;Gualano et al. 2010).
Little has been mentioned so far of the potential benefits of Cr supplementation for normal healthy people. In a placebo-controlled, randomized animal study, it was shown in fact that life-long Cr supplementation, even at very high daily dosage, is of significant benefit to life expectancy and most importantly also for life quality of normal healthy mice (Bender et al. 2008a). In a recent study with human subjects, glucose tolerance in healthy sedentary males undergoing aerobic training was improved by Cr supplementation. Thus, a change in life-style together with intake of Cr may prevent or delay the onset of health problems, such as type-2 diabetes, obesity and metabolic syndrome (Gualano et al. 2008a).
Cr supplementation, in conjunction with exercise, was shown to improve muscle performance in elderly men and postmenopausal women (Gotshalk et al. 2002(Gotshalk et al. , 2008)), as well as to increase bone mineral density in healthy elderly men (Chilibeck et al. 2005). This is in line with the findings that Cr increases survival, metabolic activity, as well as mineralization of cultured osteoblast cells in vitro (Gerber et al. 2005). Thus, Cr may not only be beneficial for muscle but also for bone health of normal healthy people. It is entirely conceivable that Cr supplementation could alleviate or prevent osteopenia and/or osteoporosis of postmenopausal women (Gerber et al. 2005(Gerber et al. , 2008)), as Cr has been shown to stimulate collagen type I synthesis and secretion of osteoprotegerin in human bone cells derived from osteopenic subjects (Gerber et al. 2008).
Positive effects of Cr supplementation on memory, learning and mental performance (Rae et al. 2003), as well as on cognitive performance, have been demonstrated (McMorris et al. 2007), and a reduction of mental fatigue by Cr was also shown (Watanabe et al. 2002).
Creatine for the elderly?
A simple and inexpensive intervention, a daily supplementation with 2-5 g of chemically pure Cr for healthy adults and most importantly for senior and elderly people (Gotshalk et al. 2002(Gotshalk et al. , 2008)), is likely to contribute as a preventive measure to muscle, bone and brain health, potentially saving billions of dollars otherwise spent for rehabilitation measures following accidents (Hespel and Derave 2007;Dalbo et al. 2009). Cr supplementation seems especially relevant for elderly, who often eat much less or no meat at all and thus likely have low tissue Cr levels, as limited data from vegans and vegetarians indicates (Burke et al. 2003;Watt et al. 2004). Recent nutritional recommendations by the US Society for Sarcopenia, Cachexia and Wasting Disease proposed Cr supplementation together with other measures for the management of sarcopenia (age-dependent progressive muscle loss) which is prevalent among the elderly (Morley et al. 2010). Interestingly, 2 weeks of 4 9 5 g of Cr daily improved cognitive performance in the elderly (McMorris et al. 2007). With respect to the possible beneficial effects of Cr supplementation that have been discussed for elderly (Dalbo et al. 2009), it is important to note that Cr acts as an osmolyte. Since Cr is taken up by the osmotically active sodium and chloride dependent CRT, concomitant import of NaCl into the target cells may lead to at least a temporary increase in the intracellular water content (Ziegenfuss et al. 1998). Hydration is an important physiological parameter in humans that gradually decreases with age (Aloia et al. 1998). However, in order to substantiate these preliminary results, many highly relevant to disease prevention and potentially with significant socio-economical health benefits, multi-centre epidemiological studies involving hundreds or thousands of subjects over prolonged periods of time would be necessary. Such studies would be expensive undertakings. However, as Cr promises only negligible financial returns to pharmaceutical companies, such studies would most likely have to be funded by government agencies.
Creatine as a prominent nutritional constituent for man since prehistoric times Concerning possible health benefits for healthy people, a legitimate question that may be asked is whether modern man, due to greatly changed eating habits, is justified in supplementing the diet with additional Cr particularly when this can be synthesised endogenously and obtained through a balanced meat and fish diet. To possibly answer this question we need to examine early hominid nutrition. A recent archaeological survey in Ethiopia brought to light stone-tool inflicted cut and percussion marks on ungulate bones that were dated to older than 3.39 million years. These, the oldest findings of this kind, after careful microscopic examination were identified as being the result of early hominid stone tools used for removing flesh from bones and for retrieving bone marrow (McPherron et al. 2010). These findings indicate that Australopithecus afarensis already practised butchery of large animals and consumed meat some 3.4 million years ago.
Humans have clearly evolved as carnivores/omnivores, ingesting large quantities of meat and fish, and thus necessarily also Cr, as a significant part of their diet (Broadhurst et al. 1998;Richards 2002). There is evidence that evolutionarily human brain development and growth were strongly dependent on the availability of high-quality food, such as meat and/or fish, representing nutritionally rich sources of protein, fatty acids, vitamins and minerals (Milton 2003) and incidentally also of Cr.
Evidence from isotopic analysis of skeletons of Neanderthals and modern Palaeolithic and Mesolithic humans highlights the importance of meat and fish in the hominid diet (Richards 2002). When successful at hunting or fishing, these hominids as true carnivores/omnivores, most likely devoured more than 1-2 kg of meat or fish per day during prolonged periods of time during the year, ingesting at least 5-10 g or more of Cr daily. The combination of high-quality diet and the higher proportion of maternal daily energy budget invested in the growing embryo during pregnancy of prehistoric women, would additionally have allowed for greater body weight as well as larger brain size (encephalization) relative to body weight of the infant at birth, compared with other primates (Ulijaszek 2002;Carlson and Kingston 2007). With respect to Cr, it is known that endogenous synthesis of Cr is energetically costly in terms of methyl-group equivalents. Carnivores/omnivores ingesting large amounts of Cr thus spare a significant proportion of the energy needed for acquisition of reactive methyl group equivalents in the form of S-adenosyl-L-methionine (AdoMet) that can be used for other anabolic synthetic pathways (Brosnan et al. 2007a, b).
It, therefore, seems that the evolutionary path of hominid development is tightly linked to food quality, e.g. to ingesting large amounts of meat and fish and concomitantly also of Cr. One might surmise from this that Cr supplementation should also belong to the nutritional requirements of modern man, depending on how much meat and/ or fish is actually ingested daily. Depending on the cultural and economic background, the present daily meat consumption varies from zero (vegans) to approximately 150 g (Switzerland) or 250-300 g (USA, Australia) of meat/ person/day (numbers include not only fresh but also processed meat that is known to contain much less Cr than fresh meat, or even none at all). These numbers correspond to a daily Cr consumption per person from zero to about 0.75-1.5 g and are clearly at the lower end of daily alimentary requirements for Cr that may be in the order of 2-4 g/person/day (see European Food Safety Authority web site: http://www.efsa.europa.eu/EFSA/efsa_locale- 1178620753812_1178620761727.htm).
Finally, the fact that in most people who ingest extra Cr the total Cr pool size (Cr ? PCr) in muscle is elevated by 5-20% indicates that in these the Cr pools are not saturated. This alone could be taken as an argument that even normal healthy people should supplement with Cr. The fact that meat consumption is recommended to be lowered globally for health (cf. high cholesterol, etc.) and ecological reasons lends additional support to the argument for Cr supplementation of the diet.
Using a special precocial mouse strain, the spiny mouse (Acomys cahirinus) with a longer pregnancy than that of normal mice, closely resembling human pregnancy, it was convincingly shown that Cr supplementation protects the brain of the mouse pups in vivo against hypoxia and thus significantly enhances survival of the offspring (Ireland et al. 2008). This corroborates earlier findings that Cr protects the brain of newborn rats against hypoxia (Adcock et al. 2002) or of adult mice in a model of stroke (Prass et al. 2007). Also, Cr displayed astonishingly positive effects in traumatic brain injury in animals (Sullivan et al. 2000;Hausmann et al. 2002), as well as in children and adolescent patients (Sakellaris et al. 2006). Maternal Cr supplementation from mid-pregnancy onwards has most recently been shown to protect the diaphragm of the newborn spiny mouse from intra-partum hypoxia-induced damage (Cannata et al. 2010).
What is new in the spiny mouse study is the fact that pregnant dams were fed Cr. This orally fed Cr is actively transported into the foetus via placental CRT (Ireland et al. 2009). Cr supplementation of the pregnant dam leads to an enhancement of total Cr levels in most organs, not only of the mother but also of the embryo, including the brain and thus protects the precocial mouse pups from episodes of hypoxia during a simulated hypoxic birth (Ireland et al. 2008;Cannata et al. 2010). These are important findings indicating that Cr supplementation during pregnancy may be a general protective measure to lower the incidence of brain damage and enhance survival also of human babies that go through periods of anoxia during birth or are at high-risk for an ischemic/anoxic birth to start with. Thus, it is entirely conceivable that the protective effects of Cr that are observed with experimental animals will also hold true for humans.
In line with this hypothesis is the fact that endogenous Cr synthesis in the spiny mouse foetus gradually develops, but only reaches a mature level some time after birth (Ireland et al. 2009). The placenta expresses relatively high amounts of Cr transporter (CRT) and the expression of CRT in the placenta is high during the entire pregnancy and increases even more before birth (Ireland et al. 2009). This indicates that a significant part of embryonic Cr is taken up via the placenta and endogenous Cr synthesis in the foetus is not yet fully established. Thus, it can be concluded that the spiny mouse foetus depends on Cr delivered by the mother via her placenta. Although Cr is basically free of significant side effects, the dosage of Cr, used in the experiments with spiny mice described above, was very high with 5%, (w/w) in the food. This amounts, depending on how much food a pregnant mouse consumes, to an equivalent of 20-50 g of Cr or more per day for an adult human. This, on the other hand, demonstrates that Cr is safe also at high dosages, even for the foetus, which is in line with a study using Cr supplementation for premature babies in a clinical set-up, where no serious side effects of Cr had been observed in preterm babies (Bohnhorst et al. 2004).
In addition to the obvious benefits of Cr for the baby, Cr supplementation of the mother during pregnancy could additionally be of benefit for the build-up of a strong uterus during the third trimester, when the uterus energetically matures by implementing the CK system (Dawson and Wray 1985;Clark et al. 1993;Clark 1994;Wallimann and Hemmer 1994). This could help to ease birthing, which largely depends on the energy charge of uterine smooth muscle and thus also of the PCr pool size (Kumar et al. 1962), to develop sufficient muscle force for the expulsion of the embryo. Therefore, it may be legitimate to propose that Cr supplementation should be a standard regimen during pregnancy, as well as after birth, both for the pregnant and lactating mother, as well as for the baby. This would favour healthy brain development of the embryo and baby, as it is obvious that Cr-deficient patients suffer from severe developmental delay with accompanying mental retardation (Schulze 2003).
Creatine as natural constituent of mother's milk Human colostrums and milk contain significant concentrations of Cr, in the range of 0.2 mM (Hulsemann et al. 1987;Peral et al. 2005). The same is true for cow and sow milk with approximately 0.8 mM Cr (Sheffy et al. 1952;Hulsemann et al. 1987;Peral et al. 2005). Interestingly, in a detailed study on sow colostrums and milk throughout lactation and weaning, both Cr and PCr were identified at up to 1.5 and 1.2 mM concentrations, respectively (Kennaugh et al. 1997). It seems that Cr values can vary significantly, depending on the analytical techniques used. In addition, the variation found in the absolute Cr values and the Cr/Crn ratio may indicate that not always fresh milk was analysed. Using non-destructive NMR methods and spectral peak assignments, however, a prominent proton NMR peak was identified as Cr in bovine milk (Hu et al. 2004), thus leaving no doubt that Cr is a genuine chemical constituent in fresh milk. It would be of importance to investigate in detail whether and how nutrition would influence the total Cr content in mother's milk.
A 3-to 4-month-old baby of 5 kg consumes approximately 800 ml milk per day from the mother at 0.2 mM Cr, or from cow milk, at 0.8 mM Cr (Hulsemann et al. 1987), amounting to a Cr ingestion of 4 or 16 mg Cr/kg body weight per day, respectively, which translates into approximately 0.3 or 1.2 g of Cr per day, respectively, for an adult person. Considering the fact that a significant amount of creatinine (25% in cow milk and more than 50% in human milk) was also found that may have arisen from Cr break-down during pasteurization of the milk, one can assume the above values to be lower-limit estimates, such that the actual Cr in fresh milk and, concomitantly, also the Cr intake by infants would be higher by a factor of 2-3 if really fresh milk were consumed.
Creatine supplementation of the mother during pregnancy, and of baby formulas and infant nutrition?
Cr is secreted by the mammary gland during lactation and has been shown to be absolutely required for normal brain development and brain function (see Wyss and Schulze 2002). According to recent data from rats, the rat mammary gland of dams does not synthesize the Cr to be secreted but is extracted from the circulation (Lamarre et al. 2010). Since there was no increase in endogenous Cr synthesis in the dams, this required an increase for the lactating mother of approximately 50% in alimentary Cr above the normal daily requirement. In rats, this is largely compensated by hyperphagia, as normal rat feed contains Cr that is introduced by dried meat and fish products (Lamarre et al. 2010). This suggests that pregnant, and even more so lactating mothers, who remain vegans or vegetarians during pregnancy and lactation, may suffer from an inadequate supply of Cr, which in turn may not be favourable for the development of the foetus and infant when nursing.
In the above work, it was further shown that while Cr is substantially accumulated in the growing pups, only 12% of this was obtained from the mother's milk (Lamarre et al. 2010). Thus, the relatively high need for Cr in the rat pups clearly places a metabolic burden on them, since, as mentioned earlier, endogenous Cr synthesis is energy costly and may use as much as 40% of the total SAM available for trans-methylation reactions (Brosnan and Brosnan 2007;Brosnan et al. 2007a, b). If this is also true for humans (Mudd et al. 2007), the legitimate question arises, as to whether supplementing lactating women with Cr and thus have them secrete more Cr into their mothers milk may relieve some of the metabolic burden of energycostly endogenous Cr synthesis by the baby. The same would hold true with respect to supplementing baby food and infant nutrition with Cr in non-nursing mothers or after weaning.
Except for soy-based baby milk, which is devoid of Cr, most of the baby formulas and follow-up baby nutrition preparations were found to contain Cr and Crn in a similar concentration range (0.3 and 0.1 mM, respectively) as mother's milk, although with significant variations (Hulsemann et al. 1987). Thus, pending more thorough investigations, it may turn out to be advisable to control and eventually supplement not only purely vegetarian soy-based baby food, but also all infant nutrition products, with Cr. Interesting in this context is the fact that it was already noted in 1913 that ''the increase in body weight of a baby after birth was roughly proportional to the Cr excreted in the urine by the respective mothers and that the excreted Cr/Crn ratio of the urine increased proportionally with mammary gland activity'' (Mellanby 1913). Thus Cr, be it delivered via mothers milk or supplemented externally, seems generally beneficial for growth and development of the infant.
Since Cr is definitely an important constituent in milk of mammals, including humans, it is hard to understand that there still exist baby formulas and early infant nutrition products, especially products based entirely on soy-bean, which do not contain any or only very little Cr. Changing this situation should be a high priority of International and European Child Nutrition Advisory Boards, and this should be an important focus of Cr research during the coming years. The socio-economic benefit of this and other Cr supplementation applications, e.g. for general skeletomuscular- (Tarnopolsky 2007) and neuro-rehabilitation (Sakellaris et al. 2006) seems obvious since every infant spared from irreversible brain damage caused for example by ischemia/anoxia during a difficult birth (Adcock et al. 2002;Ireland et al. 2008Ireland et al. , 2009;;Cannata et al. 2010) is beyond reckoning, not to speak of the relief from suffering. The objection that Cr could have side effects on human babies and infants has no scientific foundation. In a recent clinical study with premature infants, who were given 200 mg of Cr per kg body weight per day for 2 weeks, a dose corresponding to 14 g of Cr per day for an adult person of 70 kg, the treatment was well tolerated and no side effects were noted (Bohnhorst et al. 2004). However, there is an urgent need of clinically controlled studies in the field to determine a physiologically acceptable Cr dose for use with infants.
Children with a so-called Cr-deficiency syndrome, who have genetic defects either in one of the enzymes for endogenous Cr synthesis (AGAT or GAMT), or in the CRT, present with severe neurological symptoms, such as developmental and speech delay, mental retardation, autism and epilepsy (Schulze 2003), which is largely due to the complete absence of Cr in brain tissue. This emphasises the importance of Cr for normal brain development and function (Newmeyer et al. 2005) and, by implication, advocates the adoption of Cr supplementation for pregnant women, as well as for preterm babies and infants. It seems of utmost importance that Cr in infant food is recognized as an essential component of human nutrition and that its content in infant formulas should be mandatorily regulated and controlled.
Finally, the clinical tests (from urine and blood) for Cr-deficiency should be mandatory for all newborns. By this strategy, infants with treatable Cr-deficiencies (in AGAT or GAMT, but not in CRT) could be helped to lead a normal life.
Creatine supplementation of parenteral nutrition? Individuals, who are supported in intensive care units (ICU), as well as severely ill patients with a variety of disease states, e.g. cancer patients with cachexia, fully depend on parenteral food which ideally should include supplementary Cr if needs are to be met. Unfortunately, Cr has not yet been recognized as a nutrient to be included in parenteral food. In this way, atrophy of muscles, particularly intra-costal muscles and diaphragm, could be lessened in IUC patients that are supported by artificial ventilation. The longer artificial ventilation is implemented, the more difficulties such patients experience in resuming independent breathing. Since Cr has been shown to significantly alleviate muscle disuse atrophy (Johnston et al. 2009), it is to be expected that resumption of breathing and rehabilitation of ICU patients may also be positively influenced by Cr. The same, of course, holds true for severely ill patients presenting with cachexia, as often occurs for cancer patients. Parenteral supplementation with Cr would help to prevent the physiological sequelae of Cr depletion to be expected in such patients (Wyss and Wallimann 1994).
Creatine for dialysis patients?
As pointed out earlier, Cr is not toxic to the kidney but rather is important for kidney function itself, since the CK/ PCr system supports ion-pumps and metabolite transporters in this organ that are responsible for ion balance and resorption of metabolites from the urine. Patients with chronic renal failure (CRF) undergoing dialysis, who for obvious reasons are advised to restrict meat consumption, may present after a certain time with an altered skeletal and cardiac muscle energy metabolism, being low or deficient in Cr and showing a low PCr/ATP ratio (Pastoris et al. 1997;Tagami et al. 1998;Ogimoto et al. 2003). It is known that these patients lose muscle mass, become weaker and experience chronic fatigue. For these reasons, supplementation of dialysis patients with Cr might be viewed as an adjuvant therapy that should be implemented for counteracting some of the side effects of prolonged dialysis. An elegant method for Cr supplementation of this group of patients, undergoing either peritoneal or haemodialysis, would be to add appropriate amounts of Cr directly into the dialysis fluid. This would ease compliance and spare the patients to have to orally ingest yet another powder besides phosphate and calcium binders, etc. In the only study so with dialysis patients, involving the very limited number of five patients, oral Cr supplementation reduced spastic muscle cramps, a problem often observed in these patients (Chang et al. 2002).
Creatine as feed additive for animal nutrition, growing life stock and aquaculture? Endogenous synthesis of Cr is energy-costly and consumes some 40% of the SAM available for methylation reactions (Brosnan et al. 2007a(Brosnan et al. , b, 2009)). In addition, the three amino acids needed for Cr synthesis, arginine, glycine and methionine are valuable and some are 'physiologically' expensive. By including chemically synthesised Cr in a pure form in animal feeds, these amino acids would be spared for protein synthesis and growth. This holds true also for processed and dry pet food, which compared with fresh meat and offal contains very little Cr (Harris et al. 1997), such that supplementation with Cr of pet food would make sense.
Chickens fed during a growing period of 41 days with an entirely vegetable soy-based feed, to which 0.2% w/w of pure Cr was added, show a 4% greater body weight gain compared with chickens fed with normal meat-and fishmeal containing feed. In addition, feed consumption in the Cr group decreased by 2-3% and the weight gain was shown to be due to growth of lean muscle mass (Pfirter and Wallimann, unpublished data). Thus, Cr may have significant potential as an additive for animal feed, replacing millions of tons of meat-and fish-meal for animals, as well as for fish aquaculture. On a global scale, this could help with problems of world hunger and the prevention of overfishing of oceans for fish-meal production. Interestingly Cr supplementation of Drosophila melanogaster, known to express arginine kinase (AK) instead of CK (Wallimann and Eppenberger 1973), protects these flies from oxidative stress caused by exposure to rotenone, a potent mitochondrial toxin, and paraquat, a potent herbicidal redox cycler, that both generate ROS (Hosamani et al. 2010). Since these insects are not able to phosphorylate Cr into PCr, the protective effects of Cr observed cannot be due to improved cellular energetics, but are more likely related to anti-oxidant and/or anti-apoptotic effects of this guanidino compound. Thus Cr may be added to animal feed as a protectant also against environmental oxidative and toxic stress.
Acknowledgments Work from the authors cited in this review has been supported by the
Note added in proof After this review had been submitted some most recent clinically relevant publications have appeared which are fully supporting the views on the multiple beneficial actions of creatine supplementation expressed in the review presented here.
From the number of most recent publications in the field, plus those appearing in the present special volume of the Journal ''Amino
The conservation status of the members of the Honduran herpetofauna is discussed. Based on current and projected future human population growth, it is posited that the entire herpetofauna is endangered. The known herpetofauna of Honduras currently consists of 334 species, including 117 amphibians and 217 reptiles (including six marine reptiles, which are not discussed in this paper). The greatest number of species occur at low and moderate elevations in lowland and/or mesic forest formations, in the Northern and Southern Cordilleras of the Serranía, and the ecophysiographic areas of the Caribbean coastal plain and foothills. Slightly more than one-third of the herpetofauna consists of endemic species or those otherwise restricted to Nuclear Middle America. Honduras is an area severely affected by amphibian population decline, with close to one-half of the amphibian fauna threatened, endangered, or extinct. The principal threats to the survival of members of the herpetofauna are uncontrolled human population growth and its corollaries, habitat alteration and destruction, pollution, pest and predator control, overhunting, and overexploitation. No Honduran amphibians or reptiles are entirely free of human impact. A gauge is used to estimate environmental vulnerability of amphibian species, using measures of extent of geographic range, extent of ecological distribution, and degree of specialization of reproductive mode. A similar gauge is developed for reptiles, using the first two measures for amphibian vulnerability, and a third scale for the degree of human persecution. Based on these gauges, amphibians and reptiles show an actual range of Environmental Vulnerability Scores (EVS) almost as broad as the theoretical range. Based on the actual EVS, both amphibian and reptilian species are divided into three categories of low, medium, and high vulnerability. There are 24 low vulnerability amphibians and 47 reptiles, 43 medium vulnerability amphibians and 111 reptiles, and 50 high vulnerability amphibians and 53 reptiles. Theoretical EVS values are assessed against available information on current population status of endemic and Nuclear Middle American taxa. Almost half (48.8%) of the endemic species of Honduran amphibians are already extinct or have populations that are in decline. Populations of 40.0% of the Nuclear Middle American amphibian species are extirpated or in decline. A little less than a third (27.0%) of the endemic reptiles are thought to have declining populations. Almost six of every ten (54.5%) of the Nuclear Middle American reptilian species are thought to have declining populations. EVS values provide a useful indicator of potential for endangerment, illustrating that the species whose populations are currently in decline or are extinct or extirpated have relatively high EVS. All high EVS species need to be monitored closely for changes in population status. A set of recommendations are offered, assuming that biotic reserves in Honduras can be safeguarded, that it is hoped will lead to a system of robust, healthy, and economically self-sustaining protected areas for the country's herpetofauna. These recommendations will have to be enacted swiftly, however, due to unremitting pressure from human population growth and the resulting deforestation. Resumen.-Se discute el estatus de conservación de los miembros de la herpetofauna de Honduras. Basados en el crecimiento presente y proyectado de la población del ser humano, se propone que toda la fauna herpetológica de Honduras está en peligro de extinción. Lo que se conoce de la fauna herpetológica hondureña en el presente consiste de 334 especies, incluyendo 117 anfibios y 217 reptiles (incluyendo seis reptiles marinos, que no se discuten en este artículo). La mayoria de las especies se presentan en bajas y moderadas elevaciones en formaciones forestales de tierras bajas y/o húmedas, en las Cordilleras Septentrional y Meridional de la Serranía, y las áreas ecofisiográficas de la costa y las faldas de la montaña del Caribe. Un poco mas de un tercero de la fauna herpetológica consiste de especies endémicas o sino de esas especies restringidas al Mesoamérica Nuclear. Honduras es una área severemente afectada por la disminución de las poblaciones de anfibios, con cerca de la mitad de la fauna anfibia amenazada, en peligro, o extinta. Las principales amenazas a la sobreviviencia de los miembros de la fauna herpetológica son el crec-Correspondence.
The portion of the closing paragraph of E. O. Wilson's (1998) powerful book quoted above provides an extremely serious warning to our species, a warning that in continuing with our plan to place all the natural world in service to ourselves, we risk erasing any meaning for our continued existence. This concept is antipodal to the usual thinking that we encounter our raison d'être as we continue to subjugate Nature to our own designs. One of the central goals of conservation biology, then, is to attempt to bridge the gap between these antithetical worldviews in an effort to salvage and restore as much of the remaining global biodiversity as possible in the shortest time possible.
It is common knowledge among biologists that the greatest amount of biodiversity resides in the area between the Tropics of Cancer and Capricorn-the tropics. It is frequently stated that 40-80% of the diversity of life occurs in this region (Miller 2001;Raven and Berg 2001). Unfortunately, this region also is subject to the highest rates of human population growth. For example, in the Western Hemisphere, there are thirty-one countries that lie wholly within the tropics. The average natural increase for these thirty-one countries is 1.71% (data obtained from the 2000 World Population Data Sheet of the Population Reference Bureau, an insert in Raven and Berg 2001). This translates to an average doubling time of 40.9 years (using the formula DT = 70/natural increase).
The countries of Central America, however, are the fastest growing ones in the American tropics (data obtained from the 2000 World Population Data Sheet of the Population Reference Bureau, an insert in Raven and Berg 2001). Natural increase ranges from a low of 1.7 in Panama to a high of 3.0 in Nicaragua, with doubling times ranging from 23 years for Nicaragua to 41 years for Panama.
Growth rates, however, are significantly higher for the nations of northern Central America than are those for lower Central America. Costa Rica and Panama have growth rates of 1.8 and 1.7, respectively, whereas those for Belize, Guatemala, El Salvador, Honduras, and Nicaragua range from 2.4 to 3.0. For the latter five countries, these figures translate to doubling times ranging from 23 (Nicaragua) to 29 years (El Salvador). The natural increase of Honduras, at 2.8%, is the third highest in Central America, being exceeded only by those of Guatemala (2.9%) and Nicaragua (3.0%). Thus, its doubling time is the third fastest in the region, at 25 years.
The senior author has been working on the herpetofauna of Honduras since 1967. In the 35 years since then, the human population of the country has grown from about 2.4 million to a figure somewhat in excess of 6.7 million (the former figure imiento sin control de la población humana y sus vástagos, la alteración y destructión de habitación, polución, el control de pestes y predadores, el exceso de caza y explotación. Ningun anfibio o reptil hondureño está totalmente libre de el impacto humano. Se ha desarrollado una regla de medir para estimar la vulnerabilidad ambiental de las especies de anfibios, usando medidas de extensión del rango geografíco, amplitud de distribución ecológica, y estado de especialización del modo de reproducción. Se ha desarrollado una medida similar para los reptiles, usando las dos primeras medidas de vulnerabilidad usados con los anfibios, y una tercera medida para el grado de persecusión humana. Basados en estas medidas, los anfibios y reptiles muestran un rango actual de una marca de vulnerabilidad medioambiental (EVS) casi tan amplia como el rango teorético. Basados en la EVS, ambas especies de anfibios y reptiles están divididas en tres categorías, de baja, media, y alta vulnerabilidad. Hay 24 especies de anfibios y 47 de reptiles de baja vulnerabilidad, 43 especies de anfibios y 111 de reptiles de media vulnerabilidad, y 50 especies de anfibios y 53 de reptiles de alta vulnerabilidad. Teoréticamente, los valores de EVS son determinados de acuerdo de información disponible del estado presente de las taxas endémicas de is from Golenpaul, 1968, and the latter one is from data obtained from the 2001 World Population Data Sheet of the Population Reference Bureau, an insert in later copies of Raven and Berg 2001). In other words, in that 35-year period of time, the population of Honduras has doubled and increased by almost half again as much.
Habitat degradation and destruction are recognized as the major threats to biodiversity today (Raven and Berg 2001). Such degradation and destruction in Honduras is primarily fueled by deforestation (E. Wilson and Perlman 2000), occasioned by shifting agricultural practices, ranching, logging, and fuel gathering. The deforestation models in E. Wilson and Perlman (2000) indicate that the amount of forest remaining in 1995 amounted to 4.1 million hectares. Honduras, however, contains 43,277 sq. mi. or 11,208,935 hectares. Thus, in 1995 only about 37% of the original forested area of the country (i.e., once the entire country) remained. The E. Wilson and Perlman (2000) deforestation model for Honduras also indicates that the time to halve the remaining forest is 30.1 years. Thus, the 1995 figure of 4.1 million hectares will be down to 2.05 million hectares by about 2025. The deforestation rate indicated by E. Wilson and Perlman (2000) is -2.3% and will reduce the remaining forest in the country to 0.5 million hectares by the year 2085. It can be expected that, if these rates continue, no forest will remain in Honduras by the end of the present century.
Measured against this backdrop, it is abundantly clear that the Honduran herpetofauna, and indeed the entire biota, is endangered, in the best sense of the term. Equally clear, thus, is the rationale for an examination of the conservation status of the herpetofauna of the country. If we do not examine it now, we can only look to further deforestation, fueled by the uncontrolled growth of the human population, and increasing threats to the survival of the herpetofauna. We have no idea what the herpetofauna of Honduras looked like at the time of Columbus' arrival at Cabo de Honduras, opposite Trujillo, in 1502, but at least we do know that the known herpetofauna that existed when the senior author began to work in the country in 1967 is not the herpetofauna known today (see below).
It is the purpose of this paper to assess the conservation status of the known members of the Honduran herpetofauna and to construct a set of conservation and research priorities for the foreseeable future. It is hoped that the brutal honesty with which we have approached this work will act to spur the necessary steps to enable these priorities before this segment of the Honduran patrimony is lost for all time.
The modern history of the study of the amphibians and reptiles of Honduras began with the first trip to the country made by John R. Meyer in 1963. Meyer was "in country" for three months with a field crew from Texas A&M University led by the mammalogist Gerald V. Mankins. It was during this trip that Meyer began to formulate an idea for a dissertation topic dealing with a survey of the herpetofauna of Honduras. With his transfer to the University of Southern California under the mentorship of Jay M. Savage, the idea became a reality.
At about the same time, Larry D.
Wilson was also work-ing on his dissertation at Louisiana State University in Baton Rouge. Unaware of Meyer's dissertation work, Wilson began to survey various collections around the country to see what material from Honduras existed there. The word got around to Meyer, who then began to correspond with Wilson. In time, Meyer suggested that Wilson join him on a three-month field trip to the country during the summer of 1967. A second threemonth journey ensued in the summer of 1968. At this point, Meyer began to write his dissertation, which was completed in 1969 (Meyer 1969). The known herpetofauna as of that publication consisted of 196 species. Two years later, Meyer and Wilson (1971) provided a checklist of the amphibian fauna containing 52 species and in 1973 a checklist of the turtle, crocodilian, and lizard fauna listing 59 species (not 58, as stated in their abstract and introduction). Wilson and Meyer (1985) treated 95 species of snakes then known to occur in Honduras (Wilson and Meyer 1982, had treated 91 species of snakes in Honduras).
In 1976, Wilson began to work with James R. McCranie, and their first paper together (joined by Louis Porras) on Honduras appeared in 1978 (Wilson et al. 1978). These same three authors described in 1980 the first new species to result from the fieldwork up to that point (McCranie et al. 1980). In 1983, Wilson produced the first list of amphibians and reptiles for the country since the work of Meyer andWilson (1971, 1973) and Wilson and Meyer (1982). That list consisted of 208 species (56 amphibians and 152 reptiles). Wilson and McCranie (1994) produced a second update of the Honduran herpetofauna, listing a total of 277 species (89 amphibians and 188 reptiles).
The latest accounting of the species of amphibians is in McCranie and Wilson (2002). This book lists 117 species for Honduras, including two species of caecilians, 25 species of salamanders, and 90 species of anurans (one of which is reported in an addendum). The most recent list of the reptiles is in Wilson and McCranie (2002), in which are included 217 species (14 turtles, two crocodilians, 88 lizards, and 113 snakes). The total known herpetofauna, thus, as of these two publications, consists of 334 species (including six marine reptiles).
McCranie and Wilson (2002) hypothesized that seven additional species of amphibians probably reside in Honduras. A similar work in progress on the reptiles of Honduras (McCranie and Wilson, in preparation) lists 13 species of probable occurrence. At the present time, then, we know the herpetofauna consists of 334 species, and we think it may contain as many as 20 more species, apart from any new taxa that may be discovered. The above summarizes our current understanding of the composition of the Honduran herpetofauna.
Our understanding of the geographic and ecological distribution of the members of the herpetofauna of Honduras is summarized in McCranie and Wilson (2002) for the amphibians and, to a lesser extent, in Wilson et al. (2001). The latter situation is the case because Wilson et al. (2001) spent over five years in press and could not be consistently updated to the point it appeared in print. For example, Wilson et al. (2001) considered 276 species of amphibians and reptiles, but did not include five species of marine turtles, one species of marine snake, and six reptile species restricted in Honduras to the Swan Islands and the Miskito Keys. Inclusion of these 12 species would have raised their tally to 288 species, which is 46 species fewer than the number now known to occur in the country. Thus, the information presented below is somewhat more accurate for the amphibians than it is for the reptiles, although the major distributional patterns discussed are not affected much by the relative lack of currency of the information for the reptiles, nor will it have much affect on the conclusions reached in the remainder of this paper.
Both Wilson et al. (2001) and McCranie and Wilson (2002) discussed ecological distribution of Honduran amphibians and reptiles with respect to ecological formations, physiographic regions, elevation, and ecophysiographic areas. They also discussed the broad patterns of geographic distribution of these animals.
With regard to distribution in ecological formations (modified from those of Holdridge 1967), Wilson et al. (2001) indicated that the greatest number of species occur in lowland formations (Lowland Moist Forest, Lowland Dry Forest, and Lowland Arid Forest formations) and mesic formations (Lowland Moist Forest, Premontane Wet Forest, Lower Montane Wet Forest, and Lower Montane Moist Forest formations). For the amphibians alone, however, the greatest numbers of species are found in only three of the four mesic formations (Premontane Wet Forest, Lowland Moist Forest, and Lower Montane Wet Forest formations).
With reference to distribution in physiographic regions, Wilson et al. (2001) noted that the greatest numbers of species are found in the Northern Cordillera and the Southern Cordillera, these two areas comprising the Serranía of Honduras. The same pattern was discovered for the amphibians when considered alone (McCranie and Wilson 2002).
Analysis of distribution with respect to elevation indicates that the greatest number of amphibians and reptiles occur at low elevations (0-600 m), although moderate elevations (601-1500 m) harbor almost as many (Wilson et al. 2001). When amphibians are considered alone, however, there is a significantly greater number of species known from moderate elevations (88 species) than from low elevations (65 species). In addition, a sizable number of species (56) also occurs at intermediate elevations (1501-2700 m).
Combining ecological formations and physiographic regions gives rise to ecophysiographic areas (see Wilson et al. 2001 for a discussion). Thirty-eight such areas were recognized by Wilson et al. (2001), of which 28 were subjected to analysis. McCranie and Wilson (2002), however, presented data on amphibian distribution in 32 of the 38 areas (see McCranie and Wilson 2002 for a map showing the distribution of these areas). Wilson et al. (2001) showed that the highest numbers of species occurred (in decreasing order) in the Eastern Caribbean Lowlands, the West-central Caribbean Lowlands, the Sula Valley, and the Central Caribbean Slope, all of which are Caribbean lowland regions or the foothills above such areas. When the amphibians are considered alone, however, a slightly different pattern emerges. The highest numbers of species of amphibians are found in the Eastern Caribbean Lowlands, the Eastern Caribbean Slope, the Central Caribbean Slope, and the Western Caribbean Slope. The prevalence of foothill regions in this list is reflective of the sizable presence of amphibians at moderate elevations in the country (see above).
Analysis of the broad patterns of geographic distribution by Wilson et al. (2001) showed that the largest numbers of species are endemic to the country or otherwise restricted to Nuclear Middle America (about a third of the herpetofauna therein considered). Slightly more than 90 percent of the herpetofauna were distributed in the area from Mexico to South America. The amphibians, when considered alone (McCranie and Wilson 2002), show the same pattern, with 56.9% either endemic to Honduras or to Nuclear Middle America and 94.0% distributed in the area from Mexico to South America.
The overall outcome of the research on the Honduran herpetofauna that has taken place since 1967 is the description of a large number of new taxa, the discovery of a sizable number of species new to the herpetofauna, and a few resurrections of formerly synonymized taxa. More recently, however, we have entered a new era in our studies in Honduras, as detailed by McCranie and Wilson (in press) for the amphibians. As noted above, McCranie and Wilson (2002) treated 116 species of amphibians (and another one in an addendum). The majority of these 116 amphibian species are either endemic to Honduras (41 species) or otherwise endemic to Nuclear Middle America (25 species). Thus, 56.9% of the amphibian fauna falls into these two distributional categories, as noted above. The analysis presented by McCranie and Wilson (in press) indicates that of the 41 endemics, six apparently have already disappeared. The populations of an additional 14 are in apparent decline (field work in 2001 indicated that one of the 14 species thought to be in decline by McCranie and Wilson, in press, has also disappeared) and there are four species for which we do not currently know the population status. Thus, only 17 of 41 species (41.5%) appear to have stable populations at the present time. Of the 25 species otherwise restricted to Nuclear Middle America, the populations of nine species appear to be in decline and those of one species appears to have been extirpated in Honduras. We have no data on the populations of an additional four species. Thus, only 11 of 25 species (44.0%) appear to have populations that are stable at this time. Of the 50 remaining amphibian species not discussed above, McCranie and Wilson (in press) determined that 25 (50.0%) of them require relatively undisturbed forest regions to survive, and, thus, have lost much of their habitat in recent years. In summary, the populations of only 53 of 116 species of Honduran amphibians (45.7%) appear to be stable or nearly so. Thus, close to half the known amphibian fauna of Honduras is threatened, endangered, or now extinct. This sad picture is being repeated throughout much of Latin America (Young et al. 2001).
In a following section, we attempt to establish a set of conservation priorities for all the members of the Honduran herpetofauna, using revised environmental vulnerability scores, first developed and used by Wilson and McCranie (1992).
Threats to the survival of amphibians and reptiles of Honduras Wilson et al. (2001:109) opined that, "The most serious of the plethora of environmental problems impacting the planet currently, perhaps, is biodiversity decline, for this is the only one that is irreversible. As species of organisms are pushed to extinction, the information stored in their genomes is irretrievably lost. What importance such creatures have in maintaining the planet's life support systems and what more immediate or direct value that information content may have for humanity is most often extremely imperfectly known to completely unknown. Upon the extinction of the organisms, such enlightenment becomes permanently unattainable." This opinion is based on a cascade of modern research concerning the nature and extent of environmental problems, most specifically about the above-discussed problem of biodiversity decline (see, for example : Ehrlich andEhrlich 1981, 1996;E. Wilson 1984E. Wilson , 1988 [ed.] [ed.], 1992; E. Wilson and Perlman 2000;Miller 2001;Raven and Berg 2001).
The anthropogenic threats to the Earth's biota are fairly clearly identified. E. Wilson and Perlman (2000), for example, identify the following threats as most important:
• Habitat loss and fragmentation • Exotic species • Overhunting • Degradation of air, water, and soil • Synergistic pressures Raven and Berg (2001) listed the following factors as most important for U.S. plants and animals:
• Habitat loss and degradation
McCranie and Wilson (2002) identified habitat alteration and destruction, pollution, and pest and predator control as the threats of greatest importance to Honduran amphibians. When one considers the reptile segment of the herpetofauna, then overhunting and overexploitation must be added to the list. However, it may be shown that the synergistic interactions of these various threats will represent the ultimate threat (E. Wilson and Perlman 2000), pushing the existing natural systems in Honduras beyond any hope of recovery. Given the rate at which habitat alteration and destruction is proceeding, as especially measured by the rate of deforestation (see the Introduction), it may be hypothesized that the collapse of most to all of the populations of the country's amphibians and reptiles will be complete at or before the end of the present century. In the same period of time, based on Honduras's human population doubling time of 25 years (data obtained from the 2000 World Population Data Sheet of the Population Reference Bureau, an insert in Raven and Berg 2001), its population will increase theoretically by a factor of 16 times! One of the most basic questions facing the populace of Honduras is what the country will be doing with its 107.2 million people it is scheduled to have by the year 2101.
In recent years, additional threats have been manifested. One such threat comes in the form of a chytrid fungus that has been implicated as a proximate cause of mortality for anurans in Australia, Costa Rica, and Panama (see Berger et al. 1998, Lips 1999). This effect is especially startling, inasmuch as it has been occurring "… in pristine areas at moderate to intermediate elevations" (McCranie and Wilson 2002, p. 539). Many tadpoles of several Honduran species of montane hylids of the genus Plectrohyla, as well as a species of Ptychohyla, have been found to have deformed keratinized mouthparts, likely a symptom of infection by a chytrid fungus (McCranie and Wilson 2002; also see Fellers et al. 2001). Another threat may be connected to "documented climatic changes associated with recent warming" (McCranie and Wilson 2002, p. 527-528), strongly implicated by Pounds et al. (1999) to be responsible for amphibian population crashes in a Costa Rican montane habitat. We suspect "these same climatic changes are also likely taking place in montane habitats within Honduras" ( McCranie and Wilson 2002 What is especially frightening about these recent developments involving pathogens and climatic change is that they produce unanticipated changes that make it difficult to impossible to predict their effects. As such, it becomes difficult to impossible to plan for these effects. They appear to have the potential to become an environmental "super-problem," in the sense of Bright (2000). Bright (2000) uses this term to describe environmental synergisms resulting from the interaction of two or more environmental problems, so that their combined effect is greater than the sum of their individual effects. These problems represent an environmental worstcase scenario-the point when environmental problems become so serious that they produce unanticipated results, the successful resolution of which threaten to slip forever from the grasp of humanity. It is against this terrifying backdrop that we proceed with the effort to assign conservation priorities for the members of the herpetofauna of Honduras. It may be stated without fear of contradiction that there are no populations of Honduran amphibians and reptiles that are entirely free of anthropogenic impact (Wilson et al. 2001, McCranie andWilson 2002, McCranie andWilson, in press).
Prior attempts have been made by us to assess the effectiveness of the current system of biotic reserves in Honduras in protecting the country's herpetofauna (Wilson et al. 2001), to determine the status of amphibian populations (McCranie and Wilson, in press), and to anticipate the future of the amphibian faunal component (McCranie and Wilson 2002). Each of these efforts has pointed to significant threats to the integrity of herpetofaunal populations. In a very real sense, this is all we have been able to do-to point to these threats. Addressing these threats in any meaningful way is the responsibility of the people of Honduras-through their government, information media, educational systems, and environmental organizations. We have written this paper in the hope that looking at these problems in a different way than has been done heretofore may act to focus sufficient attention before it is too late-if it is not too late already. An overriding problem is that there is little consensus in the literature concerning the number and individual sizes of the protected areas in the country (see Table 15 in Wilson et al. 2001;Anonymous 2001).
Many others share these concerns, of course. In fact, Honduras is one of the countries in the Western Hemisphere that figures into the Mesoamerican Biological Corridor Project ("Paseo Pantera"), as described by Illueca (1997). While expansive and desirable in concept, there are serious problems in its design and prospects in Honduras. The map of the components of this project in Mesoamerica includes a number of "protected areas" (incidentally, one of these "protected areas," the Mayan ruins of Copán, Honduras, is mismapped; what is shown apparently is the Parque Nacional Montecristo-Trifinio) and "desired green connections." We have previously discussed the pressures existing in the "protected areas" (here and in Wilson et al. 2001). Even more significantly, however, are the problems associated with attempting to turn the "desired green connections" into anything actually "green" (i.e., ecologically restored). For example, one of these connections traverses the area between the Maya Mountains Biosphere Reserve in Belize, the Copán Maya Ruins in the department of Copán in extreme western Honduras, and the Río Plátano Biosphere Reserve in northeastern Honduras. The intervening area encompasses about the western two-thirds of Honduras, in which area lives the large majority of the human population of the country. This is also the area that has suffered greatly at the hands of agriculturists for centuries, to the point that Hondurans, especially the landless poor, are moving in significant numbers to the less heavily exploited Mosquitia in eastern Honduras. Creating a "green connection" through this area of the country appears to us to be an impossibly large task.
Several years ago (Wilson and McCranie 1992), we developed an environmental vulnerability gauge for use with amphibian populations. We then (McCranie and Wilson 2002) updated it for use with the 116 species of amphibians treated in The Amphibians of Honduras. For this paper, we have developed a similar gauge for the reptiles. The gauge for amphibians and that for reptiles resemble one another in using scales for extent of geographic range and ecological distribution. The two gauges differ from one another in that susceptibility of reproductive mode to anthropogenic pressure is used for amphibians and extent of human persecution is used for reptiles (see below).
We use these gauges to establish a set of conservation priorities for the remaining species of the Honduran herpetofauna. This is an approach different from the one we adopted in Wilson et al. (2001), which attempted to evaluate the effectiveness of the existing system of biotic reserves to protect all members of the herpetofauna known at the time, and to make suggestions about where additional reserves needed to be established. In essence, we have been forced to adopt a different approach, given the mute testimony provided in recent years by disappearing Honduran amphibians.
As noted above, this environmental vulnerability gauge for both amphibians and reptiles has three components, which are described below. The first component of the gauge, applicable to both groups, deals with the extent of the geographic range using the following scale:
1 = widespread in and outside of Honduras 2 = distribution peripheral to Honduras, but widespread elsewhere 3 = distribution restricted to Nuclear Middle America (exclusive of Honduran endemics) 4 = distribution restricted to Honduras 5 = known only from the vicinity of the type locality As is evident, in a rough sense, the degree of restriction of geographic range increases as the scale number increases.
The second gauge component, also applicable to both groups, indicates the extent of ecological distribution, based on a modified version of the forest formations of Holdridge (1967), using the following scale (omitting consideration of the Montane Rainforest formation, the herpetofauna of which is almost completely unknown):
1 = occurs in eight formations 2 = occurs in seven formations 3 = occurs in six formations 4 = occurs in five formations 5 = occurs in four formations 6 = occurs in three formations 7 = occurs in two formations 8 = occurs in one formation
The degree of restriction of ecological range increases as the scale number increases, similar to that of geographic range in the previous component.
In gauging the degree of specialization of reproductive mode in amphibians, as it relates to the effect of environmental modification, especially deforestation, we use the following scale: 1 = both eggs and tadpoles in large or small bodies of lentic or lotic water 2 = eggs in foam nests, tadpoles in small bodies of lentic or lotic water 3 = tadpoles occur in small bodies of lentic or lotic water, eggs elsewhere 4 = eggs laid in moist situations on land or moist arboreal situations, direct development 5 = eggs and tadpoles in water-retaining arboreal bromeliads or water-filled tree cavities Again, increase in number signifies probable increase in reproductive vulnerability to the effects of habitat degradation.
In light of the fact that reptiles are amniote vertebrates and, thus, do not possess the biphasic life cycle or the range of reproductive modes typical of amphibians, it is necessary to develop another gauge of human pressure on the populations of these animals. In addition, reptiles, being vertebrates fully adapted to life on land, are often more noticeable to humans and more frequently encountered than are amphibians, especially larval amphibians. Moreover, many, if not most, reptiles are the subjects of superstition, ignorance, fear, and, as a consequence, outright killing upon sight. Finally, given that all Honduran reptiles are scaled vertebrates and some are large enough to be of commercial interest for their hides, meat, and/or eggs, these species are hunted (i.e., actively sought) for these products. Taking these biological and sociological features into consideration, we developed the following scale to indicate the degree of human persecution: 1 = fossorial, usually escape human notice 2 = semifossorial, or nocturnal arboreal or aquatic, nonvenomous and usually nonmimicking, sometimes escape human notice 3 = terrestrial and/or arboreal or aquatic, generally ignored by humans 4 = terrestrial and/or arboreal or aquatic, thought to be harmful, may be killed on sight 5 = venomous species or mimics thereof, killed on sight 6 = commercially or noncommercially exploited for hides and/or meat and/or eggs
As with the previously discussed components, the degree of threat from human beings roughly increases as the scale number increases.
In order to obtain this rough idea of environmental vulnerability, thus, each of the three applicable scores has been determined for each Honduran amphibian and reptilian species. Then the numbers associated with the three scales have been added to obtain a composite score. These composite scores can range theoretically from a low of three to a high of 18 for amphibians and from a low of three to a high of 19 for reptiles.
The composite environmental vulnerability scores (EVS; used either in singular or plural form, as determined by context) for amphibians (Table 1) actually range from a low of three to a high of 17, almost the entire gamut. The numbers of species attaining the various EVS are as follows:
EVS 3-1 species EVS 11-12 species EVS 4-1 species EVS 12-13 species EVS 5-5 species EVS 13-13 species EVS 6-7 species EVS 14-15 species EVS 7-2 species EVS 15-17 species EVS 8-2 species EVS 16-10 species EVS 9-6 species EVS 17-8 species EVS 10-5 species Using this measure, the least vulnerable amphibian species are Bufo marinus, B. valliceps, Hyla microcephala, Phrynohyas venulosa, Scinax staufferi, Smilisca baudinii, and Rana berlandieri. They are all 1-1-1, 1-2-1, or 1-3-1 species (species widespread geographically in and outside of Honduras, of broad ecological occurrence, and having the least derived reproductive mode). The most vulnerable species are Bolitoglossa carri, B. decora, B. longissima, Nototriton lignicola, Eleutherodactylus chrysozetetes, E. coffeus, E. cruzi, and E. merendonensis. They are all 5-8-4 species (species known only from the vicinity of the type locality, in one forest formation, with eggs laid in moist situations on land or moist arboreal situations). In addition, three of the four species of Eleutherodactylus (save for E. coffeus for which there are no data available) appear to have already disappeared or are in decline (McCranie and Wilson, in press).
We have used the same method in this paper as McCranie and Wilson (2002). Thus, we have divided the species of Honduran amphibians into three categories of environmental vulnerability, i.e., low vulnerability, of medium vulnerability, and high vulnerability. This categorization provides an initial rough means of gauging the degree of attention that ought to be focused on the various taxa. Thus, the species that can be expected to have the best chance to survive in the face of continued environmental degradation are those in the first category. These 24 species make up only 20.5% of the Honduran amphibian fauna. A larger group of 43 species, making up 36.8% belongs to the medium category; nonetheless, this is a heterogeneous grouping, created due to a lack of weighting of the three categories used to compute the EVS, in which relatively widespread species, such as Agalychnis callidryas, are grouped with highly restricted ones, such as Plectrohyla chrysopleura. A larger group of 50 high vulnerability species, making up 42.7%, can be expected to have the poorest chance for survival. Almost all of these species are endemic to Honduras or are otherwise restricted to Nuclear Middle America. Additionally, recent declines or disappearances in amphibian populations from moderate to intermediate elevation, pristine habitats were not considered in this analysis. The importance of these declines and disappearances, however, is discussed in the following section.
The composite environmental vulnerability scores (EVS) for reptiles ( The least vulnerable reptilian species, by this measure, are Norops tropidonotus, Enulius flavitorques, Imantodes cenchoa, and Ninia sebae. They are 1-1-2, 1-1-3, or 1-3-2 species (widespread geographically, occurring in six or eight forest formations, and semifossorial or terrestrial/arboreal, sometimes escaping human notice). The most vulnerable reptile is Ctenosaura bakeri, 5-8-6 species (known only from the vicinity of the type locality, in one forest formation, and used for its meat and eggs locally). The next most vulnerable is Ctenosaura oedirhina, a 4-8-6 species (a Honduran endemic, occurring in one forest formation, and used for its meat and eggs locally).
As for the amphibians, we have divided the species of Honduran reptiles into three categories of environmental vulnerability, as indicated in Table 2. As above, this categorization is intended as a coarse gauge as to the degree of attention that should be brought to bear on the various species. There are 47 low vulnerability species, making up only 22.3% of the Honduran reptilian fauna. A slightly larger group of 53 species, making up 25.1% of the taxa, comprises the high vulnerability category. Many of these species (35) are endemic to Honduras. The largest group of 111 species, as with the amphibians, is composed of taxa of intermediate vulnerability (52.6% of total). Most of these species (93) are geographically widespread, although in many cases occurring peripherally to Honduras, and many (66) are known from only one or two forest formations.
Table 1. Environmental vulnerability scores (EVS) for the 117 species of amphibians of Honduras. Numbers for each gauge explained in text.
The table is broken into three parts: low vulnerability species (EVS of 3-9; 24 species; 20.5%); medium vulnerability species (EVS of 10-13; 43 species; 36.8%); and high vulnerability species (EVS of 14-17; 50 species; 42.7%). Updated from Table 33 in McCranie and Wilson (2002). Categorization of EVS provides a means to assign conservation priorities, with high vulnerability species given highest priority, medium vulnerability species intermediate priority, and low vulnerability species lowest priority. The highest priority taxa include 50 amphibians and 53 reptiles (total of 103 species or 31.4% of 328 total species); the intermediate priority taxa consist of 43 amphibians and 111 reptiles (total of 154 species or 47.0%); and the low priority taxa comprise 24 amphibians and 47 reptiles (total of 71 species or 21.6%).
The above discussion attempts to assign conservation priorities to the members of the Honduran herpetofauna on a largely theoretical basis, with the assumption that there are features of distribution (geographic and ecological), life history (reproductive mode), and human persecution that can act as a rough gauge of vulnerability to anthropogenic environmental pressures, in a similar manner as has been done for threatened and endangered species in general (see Raven and Berg 2001 for a discussion of such features).
As noted in a previous section, however, there are factors at work in Honduras, as elsewhere in the world, the effect of which were not predicted by the typical models of species endangerment. The unanticipated factors apparently of greatest importance are chytridiomycosis (Berger et al. 1998) and climatic warming (Pounds et al. 1999), although neither has been conclusively demonstrated to be in effect in Honduras.
Whatever the causative factors that may be involved, it is apparent that populations of many members of the Honduran herpetofauna are in decline or have disappeared since the early years of the 1990s (Wilson andMcCranie 1998, McCranie andWilson 2002, in press). The declines have been substantiated best among amphibian populations. Unfortunately, these declines have involved the two most important groups of amphibians, those endemic to Honduras and those otherwise restricted to Nuclear Middle America (Table 3). As noted by McCranie and Wilson (in press), of the 41 species of endemic amphibians, six are feared extinct and 14 appear to have declining populations (field work in 2001 indicated that one of the 14 species, Eleutherodactylus stadelmani, thought to be in decline by McCranie and Wilson, in press, has also disappeared). In addition, we have no data for four species. Only 17 species appear to have stable populations. Thus, 20 of the 41 endemic species of Honduran amphibians (48.8%), or almost half, are already gone or are in decline.
The seven endemic amphibian species feared extinct have EVS ranging between 14 and 17 (mean 15.6). The 13 species whose populations are in decline have EVS from 12 to 17 (mean 14.4). Of considerable interest is the fact that the EVS for the 17 endemics thought to have stable populations range from 11 to 17, with a mean value of 15.0. The implication of these data are that there is an urgent need to monitor populations of these supposed "stable" species, because 14 of the 17 have scores indicative of high vulnerability to environmental pressures.
McCranie and Wilson (in press) also discussed the population status of 25 amphibian species not endemic to Honduras, but restricted in distribution to Nuclear Middle America. They considered nine species to be in decline and one to probably have been extirpated. The EVS of the nine in decline range from nine to 16 (mean 12.1). The one species thought extirpated (Bolitoglossa occidentalis) has an EVS of 14. These data indicate that EVS of 13 and above are indicative of species that need to be monitored, but that scores below that level do not insulate a species from anthropogenic pressure. As we have noted above, there is no species of Honduran amphibian safe from human depredation, although there are clearly some species capable of persisting as commensals of human beings.
The picture for Honduran reptiles is somewhat less clear. This is due to the fully terrestrial life cycle of most reptiles, which allows for habitation of niches removed from water, in turn increasing the potential breadth of occurrence. Nonetheless, it is possible to comment on the current popula- tion status of reptiles endemic to Honduras or otherwise restricted to Nuclear Central America. Thirty-seven species of reptiles are endemic to Honduras (Table 4). Of these 37 species, only 19 species (51.4%) are thought to have stable populations. Ten (27.0%) are considered to have declining populations, primarily on the basis of destruction of habitat within their ranges. Finally, eight species (21.6%) are poorly known enough so that we are uncertain of their status. The ten endemic reptile species considered to have declining populations have EVS ranging between 12 and 16 (mean 14.9). The EVS for the 19 endemics thought to have stable populations range from 14 to 19 (mean 15.4), which is higher than the mean for those species thought to have declining populations. It is interesting that the reptilian endemics thought to have stable populations also have a higher mean EVS than those thought to have declining populations. The implication of these data is same as that for the analogous data for amphibians. The populations of these endemics need to be monitored carefully, inasmuch as all have scores indicating high vulnerability to environmental pressures.
We also determined the population status for those reptile species not endemic to Honduras but restricted in distribution to Nuclear Middle America. Of these 22 species, only eight (36.4%) are considered to have stable populations, at least somewhere in their known ranges in Honduras. Twelve species (54.5%) are thought to have declining populations. Finally, two species (9.1%) are too poorly known to judge their current population status.
The 12 Nuclear Middle American reptile species that appear to have declining populations have EVS ranging between ten and 15 (mean 12.8). Following the same pattern as indicated above, the EVS for the eight species appearing to have stable populations range from 12 to 14 (mean 12.9), which is slightly higher than the mean for the declining population Nuclear Middle American species. The populations of these species also need to be closely monitored.
In general, it should be understood that the population status of amphibian and reptile species in Honduras potentially can change relatively rapidly. As habitats are degraded, the fabric of community structure unravels. The community inhabitants depend on the integrity of this structure in order to obtain the materials and energy necessary to support their life processes. Thus, they are links in biogeochemical cycles and food webs, through which these materials and energy move,
X Bothriechis thalassinus X respectively. Thus, for example, given that amphibian populations are undergoing apparent increasing decline, this can be expected to adversely affect the populations of amphibian-eating snakes. In turn, decline of these snake populations should affect the populations of ophiophagous snakes, and so on. Thus does the straight edge of much human thinking cut deeply.
Plates 2-14 show some of the primary forest left in Honduras, plus some of the extensive deforestation taking place in the country. Plates 15-18 show some Honduran endemic species now feared extinct. Plates 19-38 show some of the Honduran endemic species in which all known populations are believed to be declining. Finally, plates 39-50 show some of the Nuclear Middle America-restricted species (exclusive of the Honduran endemics) in which all known populations are believed to be declining.
Biodiversity decline is one of the most serious environmental problems, if not the most serious (Wilson et al. 2001). Since it is a problem, it cries out for solutions. Unfortunately, one of the tenets of the problem solving critical thinking strategy (see Chaffee 1994 for a description of the strategy) is that a problem cannot be solved by simply treating its symptoms. Biodiversity decline is a symptom of habitat loss and degradation, in turn a symptom of runaway human population growth. Uncontrolled population growth is, in turn, a symptom of the mismanaged human mind, to use a phrase coined by E. O. Wilson (1988). The "cascade of deeper problems arising within the human psyche" (Wilson et al. 2001, p. 109) referred to by E. O. Wilson (1988) has been explored at length by L. D. Wilson in a series of papers (1997 a, b, 1998, 1999, 2000, 2001). L. D. Wilson (2001) concluded, after a lengthy argument presented in this series, that the sustainable society described by the better environmental science texts (see for example Miller 2001, andRaven andBerg 2001) will only come about (if it ever does) by a fundamental reform of the educational process, so as to enable us to use education as a kind of species-wide psychotherapy. This view, then, treats the "mismanagement of the human mind" (E. O. Wilson 1988) as a pervasive psychological illness in need of broadbased therapy.
Until and unless the "mismanaged human mind" is treated successfully, then we argue that none of the problems that cascade from it, which are, after all, the persistent problems of humankind, will ever encounter workable and lasting solutions. Having said this, then it must be understood that the recommendations we outline below will only work if the geometrically advancing problems of uncontrolled human population growth and its corollaries, habitat loss and degradation, are solved. If not, then the exercise below is merely a monument to futility.
Given the above, we have to assume that it is possible to guard the integrity of established biotic reserves in Honduras. Based on our decades-long field experience, this is only happening in a limited way. It is still the case that most biotic reserves in the country exist only on paper, without the appropriate resources dedicated to establish boundaries, hire personnel to police them, build facilities for housing administrative, scientific, and security personnel, and fund the scientific studies necessary to make such reserves sustainable. This situation will have to change and change rapidly, for the pressure of a 25-year doubling time will brook no idleness.
It is also evident that we have been idle too long, and that the study of the Honduran herpetofauna has turned a corner into a torturous maze from which there is no easy exit. It is already clear, as is discussed above, that a new era has been breached-one in which advances in our cataloguing of the herpetodiversity of Honduras is being offset by documented losses of that same diversity over the last decade or so. We are, thus, fighting an uphill battle on very slippery slopes.
In full light of the provisos identified in this section above, the following recommendations concerning the protection of the members of the Honduran herpetofauna are made:
• The system of biotic reserves should be expanded to include areas for protection of species not currently known to reside in any legally established reserve. The locations of such areas are discussed by Wilson et al. (2001) and McCranie and Wilson (2002). Of the Honduran endemics, there are 14 such species. For the Nuclear Middle American species, seven species are involved.
• The entire system should be evaluated to ascertain the health of the populations of amphibians and reptiles resident within the various reserves. At least an initial effort can be accomplished by use of Rapid Ecological Assessment Program methodology (see Parker and Bailey 1991).
• Following this evaluation, the system of reserves should be adjusted to the extent possible to provide maximal protection of the remaining populations of resident amphibians and reptiles. Undoubtedly, this step also would involve establishment of additional reserves. Wilson et al. (2001) and McCranie and Wilson (2002) provide some guidance for such decisions.
• Steps then should be taken to clearly identify the limits of the reserves, build facilities to house personnel, involve local people in planning and decision making, make employment available to local people, and put the resulting revenues into local communities for future improvements. Meyer and Meerman (2001) discussed this type of "participatory" management strategy, which they advocate to replace the traditional "exclusionary" management strategy maintained by them to be ineffective over the long term. These steps, which need to occur as rapidly as possible, will obviously require appropriate allocation of governmental funds. The administration of the new Honduran president, Ricardo Maduro Joest, is just beginning. It remains to be seen what priority is established by the new government to address these issues.
• Once facilities are available for housing personnel, then the longer-term scientific survey work and other sorts of scientific studies can begin, with the goal of establishing the biological worth of the various reserves. Opportunities for cooperation in such studies between resident and foreign scientists should be explored. We continue to explore such collaborations with various Honduran biologists.
• With completion of facilities and scientific studies can come educational and ecotourist programs, with the goal of making the reserves economically self supporting. Again, cooperative undertakings should be encouraged. Such steps would involve reaching out to various Honduran and foreign governmental and nongovernmental organizations.
• Our strongest recommendation is that the steps outlined above be taken with all dispatch possible. We have demonstrated that populations of a highly significant number of species of Honduran amphibians and reptiles are already in decline or have disappeared, especially of the most important segment containing the endemic species and those whose distribution is otherwise restricted to Nuclear Middle America. In addition, deforestation has been demonstrated to be increasing at an exponential rate, commensurate with the increase in human population. Deforestation is the principal type of habitat destruction in Honduras, which is, in turn, the major threat to the highly distinctive and important Honduran herpetofauna. There is, in the final analysis, no time to dawdle.
"We must learn to use our intelligence to live more lightly on the land, so that we do not degrade the only home we have-and the only one we can leave to our children."
Acknowledgments.-Over the decades that we have worked on the herpetofauna of Honduras, we have incurred indebtedness to uncounted individuals and organizations. First and foremost, this work would not have been possible without the support of the personnel of Recursos Naturales Renovables and Corporación Hondureña de Desarrollo Forestal (COHDEFOR), who made the necessary collecting and export permits available to us over the years. In addition, our good friend Mario R. Espinal has acted as our agent in Honduras to assure that the gaining of these permits went as smoothly as possible, that vehicles were ready when we needed them, and that other myriad details upon which the success of our field work depended were appropriately managed. Finally, it has been our great and continuing pleasure to work with uncounted Hondurans who have made we gringos feel welcome in what is, in a very real sense, our second home. Without the unstinting efforts of these Honduran friends, we would never have unraveled the secrets of the animals to which we have devoted our scientific lives.
We also wish to thank Nidia Romer, who kindly translated the Abstract into Spanish for use as the Resumen. Finally, we owe a sincere debt of gratitude to Louis Porras, Jay M. Savage, and Jack W. Sites, the reviewers of this paper, whose suggestions materially improved our presentation.
Continued on page 20.
DOI: 10.1514/journal.arc.0000012.t001
Amphib. Reptile Conserv. | http://www.herpetofauna.org
Based on specimens without precise locality data and one sight record in the Middle Choluteca Valley.
However, this species is extirpated on the Swan Islands, the only place where this species is known in Honduras. DOI: 10.1514/journal.arc.0000012.t002
DOI: 10.1514/journal.arc.0000012.t004
Table 3. Current status of populations of Honduran amphibian endemics and species otherwise restricted to Nuclear Middle America. Stable = at least some populations stable; Declining = all populations believed to be declining. Extinct category applies to Honduran endemics; extirpated category applies to Nuclear Middle American endemics (excluding those endemic to Honduras).
The conservation status of the herpetofauna of Honduras
Volume 3 | Number 1 | Page 29 Amphib. Reptile Conserv. | http://www.herpetofauna.org DOI: 10.1514/journal.arc.0000012.t003
Deposition of immunoglobulin light chains is a result of clonal proliferation of monoclonal plasma cells that secrete free immunoglobulin light chains, also called Bence Jones proteins (BJP). These BJP are present in circulation in large amounts and excreted in urine in various light chain diseases such as light chain amyloidosis (AL), light chain deposition disease (LCDD) and multiple myeloma (MM). BJP from patients with AL, LCDD and MM were purified from their urine and studies were performed to determine their secondary structure, thermodynamic stability and aggregate formation kinetics. Our results show that LCDD and MM proteins have the lowest free energy of folding while all proteins show similar melting temperatures. Incubation of the BJP at their melting temperature produced morphologically different aggregates: amyloid fibrils from the AL proteins, amorphous aggregates from the LCDD proteins and large spherical species from the MM proteins. The aggregates formed under in vitro conditions suggested that the various proteins derived from patients with different light chain diseases might follow different aggregation pathways.
Light chain amyloidosis (AL), light chain deposition disease (LCDD) and multiple myeloma (MM) are immunoproliferative disorders characterized by excessive proliferation of monoclonal plasma cells, excretion of high amounts of Bence Jones proteins (BJP) and deposition of immunoglobulin light chains. BJP are soluble free monoclonal immunoglobulin light chains that are secreted into circulation and excreted in urine [1]. The symptoms and tissue damage vary between these three light chain diseases. AL is characterized by the deposition of immunoglobulin light chains as amyloid fibrils in vital organs leading to organ failure. The organs most commonly affected include the kidney, heart and liver [2]. AL can also affect tissues such as peripheral nerve, gastrointestinal tract and pulmonary [3]. LCDD is characterized by granular amorphous aggregates in the basement membrane of the kidney. The kidney is the most frequently affected organ, but LCDD can also occur in other organs such as the heart and liver, where it is usually asymptomatic [4]. MM is characterized by bone lesions, hypercalcemia, renal failure and anemia. MM can form casts in the kidney that can lead to renal complications [5]. Lambda light chains are more prevalent in AL, while kappa light chains are more prevalent in LCDD and MM [6,7].
The ability of BJP to aggregate could be caused by numerous factors, in particular the combination of thermodynamic instability due to somatic mutations, germline context and/or proteolysis [8]. According to Hurle et al., domain stability of proteins plays an important role in fibrillogenesis. The less stable domains tend to form aggregates and any mutation that decreases domain stability could lead to AL [9].
There is a bias for certain germline sequences towards amyloidogenicity. AL has a preference for certain germline sequences including VlIII 3r, VlVI 6a, VkI O18/O8 and VkIV B3 [6]. In LCDD, there is a preference for VkIV B3 [10]. Some possible explanations for this bias include the fact that certain germline gene families are more available, leading to recombination events and clonal expansion [6]. It is also possible that some germline genes are selected due to their antigen diversity.
The majority of AL amyloid fibrils are made up of N-terminal fragments corresponding to the variable domain, suggesting the possibility of a proteolytic event. Partial proteolysis could destabilize proteins and predispose them to form aggregates [8].
In this study, we compared the protein structure, protein stability and aggregation properties of light chain proteins from AL, LCDD and MM patients in order to gain an understanding about the similarities and differences between BJP involved in these different light chain deposition diseases.
Bone marrow cells were collected following the guidelines from the Institutional Review Board at the Mayo Clinic. RNA extracted from patients' bone marrow cells was used in a reverse transcriptase polymerase chain reaction (RT-PCR) to produce cDNA. The cDNA was amplified with degenerate primers encoding the N-terminus of the germline sequences by PCR (six different primers for lambda proteins and three different primers for kappa proteins) as previously reported [6]. Light chain sequences were analyzed using Vbase (www.mrc-cpe.cam.ac.uk) to determine the germline donor sequence for each DNA variable domain sequence. Specific germline primers corresponding to the variable light chain gene were used to amplify the gene corresponding to the light chain variable domain by PCR and then cloned into pCR 1 II TOPO 1 vector following the protocol from the TOPO TA cloning 1 kit (Invitrogen, Carlsbad, CA, USA). The DNA Synthesis and Sequencing Core Facility at the Mayo Clinic synthesized primers and sequenced plasmids. DNA sequences were verified in Vbase and by BLAST 2 sequence alignment (www.ncbi.nlm.nih. gov/blast/bl2seq/wblast2.cgi) and the sequences were translated into their amino acid sequence using ExPasy (www.us.expasy.org/tools/dna.html) and mutations in the patient sequence were noted in comparison to the germline. The sequences have been deposited in GenBank with the following accession numbers DQ240234 (AL-02), DQ240235 (AL-03), DQ240236 (MM-01), DQ240237 (MM-02), DQ240238 (LCDD-01), and DQ240239 (LCDD-02, predominant VkI O12/O2 sequence).
Patients' urine samples were collected following the guidelines from the Institutional Review Board at the Mayo Clinic. Patients' urine samples were dialyzed in nanopure water overnight at 48C. The dialyzed urine was filtered with 0.22 mm disposable filter and 0.02% sodium azide was added. BJP were purified by size exclusion chromatography using a HiLoad 16/60 Superdex 75 prep grade column on an AKTA FPLC (GE Healthcare, Piscataway, NJ, USA) in 10 mM Tris-HCl pH 7.4 buffer. Pure fractions were checked by SDS polyacrylamide gel electrophoresis (SDS-PAGE) and stained with Coomassie blue. The extinction coefficient for each protein was calculated using the biopolymer calculator (http://paris.chem. yale.edu/cgi-bin/extinct.pl) which was used with its absorbance at 280 nm to determine the protein concentrations. Pure fractions were combined, concentrated and stored at 48C.
The molecular weight of the purified protein was determined on a Superdex 75 10/30 size exclusion column on an AKTA FPLC (GE Healthcare) with 10 mM Tris-HCl pH 7.4 and 150 mM NaCl buffer. Molecular weight standards of albumin, carbonic anhydrase, cytochrome C and aprotinin (Sigma, St Louis, MO, USA) were prepared in 10 mM Tris-HCl pH 7.4 and 150 mM NaCl buffer with a concentration of 3 mg/ml for aprotinin and carbonic anhydrase and a concentration of 2 mg/ml for albumin and cytochrome C. The column was calibrated with each standard at a flow rate of 0.3 ml/min. The elution volume for each standard peak was determined. The purified BJP were injected and eluted through the column at a flow rate of 0.3 ml/min. Blue dextran, concentration of 1 mg/ml, was injected onto the column to determine its void volume, which is used to calculate the ratio of elution to void volume (Ve/Vo) for each molecular weight standard. A calibration curve was produced by plotting the logarithms of each standard as a function of their Ve/Vo and determining the line of best fit. After determining the Ve/Vo of each BJP, the molecular weight was calculated by using the equation for the line of best fit.
Protein secondary structure was measured by Far UV-CD spectra (260-200 nm) on an AVIV 215 circular dichroism (CD) spectrometer (AVIV Biomedicals Inc., Summerset, NJ, USA) in the continuous mode by taking measurements every 1 nm with an averaging time of 5 s at 48C in a 0.2 cm path-length cuvette. Protein concentrations varied between 5.5-25.6 mM. Thermal denaturations were carried out following the ellipticity at 218 nm (maximum b-sheet signal) for each protein. The ellipticity was monitored at 28C intervals from 4 to 908C with an equilibration time of 1 min and an averaging time of 30 s in a 0.2 cm path-length cuvette. A refolding curve was monitored immediately after the unfolding curve acquisition from 90 to 48C using the same parameters as previously stated. Thermal denaturation data were processed following a two-state transition model. Folded and unfolded baselines were linear extrapolated from the data with a minimum of 10 points for the majority of the proteins. In some cases (uMM-02), we used fewer points for the unfolded baseline. The fraction folded (FF) at each temperature was calculated by using the following equation: FF ¼ (ellipticity observedÀellipticity of the unfolded)= (ellipticity of foldedÀellipticity of unfolded).
Melting temperature (T m ) was calculated at the midpoint transition for each protein. DG was calculated according to the equation
from Pace where DH is derived from the van't Hoff equation:
ln Keq ¼ ðÀDH/RÞ Ã ð1=TÞ þ ðDS/RÞ and DCp % 172 þ 17:6
where N is the number of amino acid residues and SS is the number of disulfide cross-links in the protein [11]. An alternative method tested used the points in the FF transition to determine the free energy of folding. The equilibrium constant (Keq) was calculated from these points. DG folding was calculated using the following equation:
where R ¼ gas constant (1.98 cal mol 71 K 71 ), T ¼ temperature in Kelvin. DG(48C) was calculated by extrapolating free energy of folding versus the temperature to determine the line of best fit for the free energy of folding at 48C/277 K. Both methods gave rise to very similar DG folding values. Chemical denaturations were done at 48C for all proteins except for uAL-03, which was done at 228C in a 0.2 cm path-length cuvette. A Far UV-CD spectra of the folded (absence of urea) and unfolded (presence of urea) samples were collected to determine the wavelength with the maximum ellipticity difference between them. Ellipticity at that wavelength (218 nm) was monitored in 1 s intervals during a 60 s scan, with an equilibration time of 10 min prior to the scan. The data were averaged for each urea concentration. Denaturation curves were obtained by mixing equal volumes of two protein stock solutions, in the presence and absence of urea, to change the urea concentration while keeping the protein concentration constant. Both the initial and final urea concentrations were confirmed by refractometry and calculated using the equation [urea] ¼ 117:66 Ã ðDNÞ þ 29:753 Ã ðDNÞ 2 þ 185:56 Ã ðDNÞ 3 ;
where DN is the difference in refraction between the sample and buffer [11]. The denaturation curves were processed the same way as the two-state transition model used in the thermal denaturation to calculate fraction folded curves. The melting concentration of denaturant (C m ) was determined at the midpoint transition of the unfolding curve. DG folding was calculated using the same equation as used with the thermal denaturation data. Free energy of folding versus the concentration of urea was extrapolated to determine the line of best fit for the free energy of folding in the absence of denaturant (DG(H 2 O)).
A stock solution of ANS (8-anilino-1-napthalenesulfonic acid) was prepared, filtered and its concentration determined by checking its absorbance at 350 nm and using an extinction coefficient of 5000 (M Ã cm) 71 [12]. A 5 mM protein sample was prepared with 2.3 mM ANS in 10 mM Tris-HCl pH 7.4 buffer in the presence and absence of 6 M urea. Samples were placed in a quartz cuvette with a pathlength of 1 cm in a temperature-controlled Model QM-2001 fluorometer (Photon Technology International, Lawrenceville, NJ, USA). ANS fluorescence was monitored at an excitation wavelength of 370 nm and by an emission wavelength scan from 400-620 nm with slit widths of 7 and 8 nm at a temperature of 268C.
Urine samples, pure BJP and known positive controls were run on a SDS-PAGE gel and transferred to an Immobilon P (PVDF) membrane (Millipore, Billerica, MA, USA) using a semidry transfer apparatus for 1 h. The membrane was blocked for 1 h in PBS with 5% non-fat dry milk rocking at room temperature. The membrane was rocked overnight at 48C with 1:500 dilution of the primary antibody (anti-human kappa (AFF) or anti-human lambda (AFF); The Binding Site, San Diego CA, USA) in blocking buffer. The membrane was washed with two washes of PBS, three washes of PBS þ 0.1% Tween 20 followed by one wash with PBS. It was rocked for 1 h with a 1:8000 dilution of secondary antibody conjugated to horseradish peroxidase (HRP; rabbit polyclonal to sheep IgG H þ L (HRP); Novus Biologicals, Littleton, Co, USA) in blocking buffer at room temperature. The membrane was washed again with PBS and PBS þ 0.1% Tween 20 following the previous wash method listed above. Antibody bound to the protein was detected using an ECL chemiluminescence reagent (GE Healthcare) and exposed to film.
Aggregate formation assays incubated at the T m of each protein were set up with 5 mM protein in 10 mM Tris-HCl pH 7.4 buffer with 15 mM thioflavine T (ThT) with a final volume of 1.5 ml. Samples were placed in a quartz cuvette with a stir bar and incubated at their T m (578C for uAL-02, uAL-03, uMM-01 and uMM-02, 608C for uLCDD-01 and uLCDD-02) in a temperature-controlled Model QM-2001 fluorometer (Photon Technology International). Tryptophan fluorescence was monitored using an excitation wavelength of 294 nm and emission wavelength of 350 nm. ThT fluorescence was monitored using an excitation wavelength of 450 nm and emission wavelength of 480 nm. The two different probes were followed simultaneously for 72 h with continuous stirring at 300 rpm. Shutters were open only during the acquisition. All reactions were measured in a 1 cm path-length cuvette with slit widths of 3 nm, taking an acquisition every 25 s with an integration time of 0.5 s.
Aggregate formation assays incubated at 378C were set up with 5 mM protein in 10 mM Tris-HCl pH 7.4 with 5 mM ThT in triplicate in a 384-well, flat-bottomed, high-binding, polystyrene Costar plate (Corning Incorporated, Corning, NY, USA) in a GENios FL plate reader (TECAN US, Durham, NC, USA). The plate was incubated at 378C with a reading occurring every 15 min with 3 min of shaking prior to each reading.
Endpoint samples from the aggregation experiments were centrifuged and resuspended in 100 ml of buffer. Three microliters of the concentrated aggregate was placed on a 300-mesh copper formvar/carbon grid and negatively stained with either 4% uranyl acetate or 1% phosphotungstic acid and examined on a JEOL 1200 EX transmission electron microscope.
Ten micrograms of each protein were sent to Scientific Research Consortium, Inc. in St. Paul, Minnesota where acid hydrolysis was performed.
Twenty-five micrograms of each protein were taken to the Mayo Clinic Proteomics Research Center to carry out trypsin digestion/tandem mass spectrometry and intact mass analysis. The proteins to be analyzed by nanoLC-ESI-tandem mass spectrometry were initially reduced and alkylated with DTT and iodoacetamide, digested with trypsin and the peptides run on a ThermoFinnigan LTQ Orbitrap. The MS/MS spectra were searched with Sequest using both Swiss-Prot database and the expected peptides from the sample protein sequences. Intact mass measurements were made by LC-MS using a C18 reverse-phase HPLC column eluting into an Agilent LC/MSD TOF mass spectrometer.
We studied two BJP for each type of light chain disease. These BJP were called uAL-02, uAL-03, uLCDD-01, uLCDD-02, uMM-01 and uMM-02. The 'u' denotes these proteins were derived from a patient's urine sample and the name without the 'u' denotes the corresponding DNA sequence name. Table I contains information for each protein studied such as the germline sequence, disease, organ involvement, year of diagnosis and confirmation of diagnosis for each protein. Each protein sequence was translated from the DNA sequence and compared with its corresponding germline to note any mutations which are highlighted in gray in Figure A day 100 post-stem cell transplant bone marrow sample was used in the cloning process for LCDD-02. We found cDNA sequences that matched two germline sequences, VkI O12/O2 and VkI L1. The sequence corresponding to the VkI O12/O2 germline was the most dominant (eight of 21 sequences) and was used as the reference sequence for this study (Figure 1). The nucleotide sequence comparisons for VkI L1 samples yielded all different sequences. It is interesting to note there are a difference of seven amino acids between the VkI O12/O2 and L1 germline sequences and in comparing their nucleotide sequences, there are 16 nucleotide changes. Germline nucleotide sequences were found in Vbase (www.mrc-cpe.cam.ac.uk). Since the bone marrow sample used to generate the LCDD-02 sequence was taken poststem cell transplant, we think it was presumably polyclonal. Unfortunately, no pretreatment bone marrow sample was available for LCDD-02.
In Figure 1, the protein sequence for the variable domain of LCDD-02 is identical to MM-02 including the location and number of mutations in both of them. Both of these protein sequences correspond to the same germline, VkI O12/O2. However, when comparing the nucleotide sequences for MM-02 and LCDD-02, there are significant differences between them. We analyzed 12 LCDD-02 sequences and found nucleotide changes in different positions and at different frequencies confirming the polyclonal nature of the sample. Taking these differences into account, the sequences are different at a nucleotide level even though the amino acid sequences appear identical.
The constant domains for the AL and MM proteins were not sequenced. Most of the constant domain for MM-01 was determined by the Mayo Clinic Proteomics Research Center upon complete sequencing analysis with the kappa constant domain from the IMGT website (http://imgt.cines.fr/). However, based on the small region sequenced corresponding to the constant domain as part of the light chain variable domain (VL) sequencing, the constant domain of the lambda proteins was determined. There are six possible constant domains for lambda light chains. uAL-02 has a constant domain of IGLC1 and uAL-03 has a constant domain of IGLC2. For the MM proteins, the unique kappa constant domain was assigned. The sequences for all constant domains were found on the IMGT website (http://imgt.cines.fr/).
The proteins were purified by size exclusion chromatography with 10 mM Tris-HCl pH 7.4 buffer. A 25-26 kDa protein band is visible in all samples except the purified LCDD proteins which has a 12 kDa protein band (Figure 2). A Western blot confirmed the purified proteins were immunoglobulin light chain proteins (data not shown). The molecular weight of the proteins was verified by analytical size exclusion chromatography (data not shown) with uAL-02 and uAL-03 being dimers, and uLCDD-01, uLCDD-02 and uMM-01 being monomers. uMM-02 migrated as a pentamer according to the analytical size exclusion column results, which could be contributed to the long-term storage of pure protein at 48C before injecting onto the column (248 days).
Amino acid analysis results show 98% or more identity to the predicted sequence from translation of Table I. Protein comparison table. Protein Isotype Germline Disease Diagnosis year Diagnosis confirmed Organ(s)/tissue involved Molecular weight (kDa) uAL-02 Lambda lI-1c AL 2003 Endomyocardial biopsy Heart, liver 25 uAL-03 Lambda lIII-3r AL 2003 Renal biopsy Renal 25 uLCDD-01 Kappa kIV-B3 LCDD 2003 Renal biopsy Renal 12 uLCDD-02 Kappa kI-012/02 LCDD 2004 Renal biopsy Renal 12 uMM-01 Kappa kII-A19 MM 1986 Lytic lesions and a metastic bone survey Bone marrow 25 uMM-02 Kappa kI-012/02 MM 2003 M spike and a negative bone x-ray Bone marrow 25 AL, light chain amyloidosis; LCDD, light chain deposition disease; MM, multiple myeloma.
DNA for all proteins. Mass spectrometry of tryptic fragments had an overall good coverage of all proteins confirming the sequences from the DNA translation. The intact mass measurement of uAL-02, uAL-03 and uMM-01 were done by reversephase LC-electrospray-TOF mass spectrometry. The other three proteins did not give reliable results. The reduced samples of uAL-02, uAL-03 and uMM-01 had an observed mass close to the theoretical mass calculated for each of them. The secondary structure of the six purified BJP was determined by Far UV-CD spectra (260-200 nm). We observed the expected b-sheet structure with a minimum around 218 nm for each protein (Figure 3). No notable differences were found in any of the spectra for these samples and all three groups of proteins look very similar. uLCDD-02 has a less intense mean residue ellipticity (MRE) value around 218 nm than the rest of the proteins and it has a shifted maximum around 206 nm. All proteins followed a two-state unfolding transition with similar T m values (Figure 4 and Table II). All of the proteins were able to refold reversibly except for uAL-03. uLCDD-01 and uMM-01 have the lowest values/most favorable free energy of folding from the thermal denaturation data (DG(48C)) while both AL proteins show the highest/less favorable values (Table II). The DG(H 2 O) calculated from chemical denaturation data follow the same trend but show different values. uAL-03 and uMM-02 show the highest C m while uMM-01 and uMM-02 show the lowest DG folding (Table II). Urea denaturations for uLCDD-01 and uLCDD-02 did not follow a twostate unfolding transition.
We conducted aggregation assays by incubating each protein (5 mM) with 15 mM ThT in 10 mM Tris-HCl pH 7.4 at their T m (578C for uAL-02, uAL-03, uMM-01, uMM-02, 608C for uLCDD-01, uLCDD-02) (Table II) as was previously reported [13,14]. The kinetics of aggregate formation were followed simultaneously by tryptophan and thioflavine T (ThT) emission wavelengths of 350 nm and 480 nm, respectively. The changes on tryptophan fluorescence for the two AL proteins studied have different early aggregation events but eventually converge into the same process (Figure 5, panel A). The changes in tryptophan fluorescence followed the same trend for the two LCDD proteins but LCDD-01 had higher fluorescence intensity throughout the whole assay when compared to LCDD-02 (Figure 5, panel B). MM proteins had a slight variation in the beginning of the aggregation process but the fluorescence signals eventually converge (Figure 5, panel C). The initial changes seen in Figure 4, panel A and B (AL and LCDD graphs) are due to the initial equilibration of the protein samples to their T m . The most significant changes in the BJP happened within the first 2 h of the reaction. In conclusion, no pattern of tryptophan fluorescence can be matched to the different aggregation processes by the different BJP proteins.
The intensity of the initial and endpoint ThT fluorescence for each BJP is summarized in Figure 6 along with the fold change for each protein. The largest fold change is found in both of the AL proteins. The formation of fibrils or aggregates was confirmed by EM as seen in Figure 7. Both of the AL proteins formed fibrils with a diameter ranging from 14-50 nm. LCDD proteins formed amorphous aggregates with diameters between 50-300 nm. MM proteins formed spherical species with diameters of 100-150 nm. MM and LCDD proteins are considered non-pathologic due to the lack of the formation of fibrils. There is no significant increase in ThT fluorescence after 8 days following the aggregation kinetics of the samples at 378C. The fold change for the proteins at 378C is lower than the fold change at their T m . This would suggest the ability to form aggregates under these conditions is occurring slower than samples incubated at their T m with continuous agitation of 300 rpm (Figure 6).
The predicted amino acid sequences from the cloned MM-02 and LCDD-02 samples are identical for the variable domain. However, it is important to note a few differences between these two samples. The nucleotide sequences are different between these two samples. The MM-02 protein also contains the constant domain which is missing in the LCDD-02 protein. It is possible that the dominant clone found for LCDD-02 is not the pathologic clone but it may be the most abundant sequence in the polyclonal bone marrow specimen taken post-stem cell transplant. MM-02 is a clonal sequence in that almost all of the nucleotide sequences are identical. This is not true for the LCDD-02 nucleotide sequences.
Our results indicate that BJP derived from patients afflicted with these three different light chain diseases present very similar T m values. Both thermal and chemical denaturation derived DG folding values indicated that MM and LCDD proteins are more stable than AL proteins, which is in agreement with previous reports [12,[15][16][17].
Due to the absence of two-state unfolding transition for LCDD proteins using urea denaturation, we wanted to know if they were sampling unfolded states and exposing hydrophobic patches in the absence of denaturant at room temperature. For that purpose, the proteins were incubated with 2.3 mM ANS (8anilino-1-naphthalene-sulfonic acid) in the presence and absence of 6 M urea to see if enhanced fluorescence with ANS could be detected. ANS is a hydrophobic dye used to detect exposed hydrophobic surfaces as well as partially folded intermediates [18][19][20]. Fluorescence enhancement of ANS in the presence of both proteins suggested that LCDD proteins are sampling partially unfolded states (data not shown). The presence of 6 M urea further enhanced ANS fluorescence, which suggested that the proteins were able to expose more hydrophobic surfaces and therefore might be partially unfolded in
Table II. Thermodynamic parameters of Bence Jones proteins. Protein T m (8C) DG(48C) (kcal/mol) C m (M) DG(H 2 0) (kcal/mol) uAL-02 57.2 + 0.5 76.8 3.1 + 0.5 73.1 uAL-03 57.3 + 0.4 76.6 4.4 + 0.5 72.2 uLCDD-01 59.6 + 0.4 717.4 NA* NA* uLCDD-02 56.0 + 1.4 712.6 NA* NA* uMM-01 56.6 + 0.4 717.1 4.0 + 0.8 73.6 uMM-02 57.8 + 0.6 715.1 6.2 + 0.5 74.0 *Unable to determine these values due to lack of two-state transition. AL, light chain amyloidosis; LCDD, light chain deposition disease; MM, multiple myeloma. the absence of denaturant (data not shown). Since a two-state unfolding transition was observed when a thermal denaturation was performed, it is possible that the unfolded state populated in urea is different from the one populated at high temperatures. The LCDD proteins in this study lack the constant domain, which may help explain their lack of a two-state unfolding transition when followed by urea.
Striking differences between the different BJP were found in their aggregation properties. The different BJP were incubated at their T m in 10 mM Tris-HCl pH 7.4 buffer for 72 h. These solution conditions have been reported to maximize aggregation [13,14,21]. Our results show the largest fold change in the ThT fluorescence intensity between the initial and endpoints for the AL proteins.
According to our results, the tryptophan emission fluorescence decreases in intensity over time during the in vitro aggregation assay for all the BJP in this study. The changes in the signal can be in response to conformational transitions, denaturation or changes in the environment of the protein [22].
The proteins studied were modeled using 1BRE.pdb (k) and 1CD0.pdb (l) to determine the location of each tryptophan residue in the protein to get an idea of which ones were contributing to the fluorescence which may help us understand the conformational changes occurring during aggregation. The six proteins studied have a tryptophan residue at position 35. Both of the AL proteins have four tryptophan residues which are in the same position in each of these sequences even though they belong to two different variable and constant domain germline sequences. The two tryptophan residues in the MM proteins are in the same position even though they belong to different germline sequences. LCDD-01 has two tryptophan residues one at position 35 protein and tryptophan 50. LCDD-02 only has one tryptophan in the variable region at position 35.
AL proteins formed amyloid fibrils with diameters of 14-50 nm. These diameters are slightly larger than the range of 7-12 nm that has been previously reported for the variable domain amyloid fibrils [23]. The difference in diameter could be due to the presence of the constant domain as part of the fibril. MM proteins consistently formed large spherical species with diameters of 100-150 nm, which is contrary to previous reports of variable domain MM proteins that were able to form amyloid fibrils [16,[24][25][26]. LCDD proteins formed predominantly amorphous aggregates that had diameters of 50-300 nm. MM and LCDD proteins did not form fibrils and in turn are non-pathogenic.
What causes a protein to misfold and form either fibrils or aggregates? Clues to the type of aggregation formed by a protein could be found in the pathway the protein follows in becoming an aggregate. The aggregation pathway could be either an on-or offfolding pathway, going through a possible intermediate before reaching its form of deposition. Khurana et al. has reported that partially unfolded intermediates could lead to fibrils or amorphous aggregates [12]. According to Vidal et al., offpathway aggregation does not necessarily form amyloid fibrils. For example, in LCDD, the formation of fibrils may be avoided due to the formation of intermediates that lead to aggregate formation in an off-pathway fashion [27]. Figure 8 shows a model of the various misfolding pathways a protein could follow on its way to becoming an aggregate. It is possible the spheres formed by MM proteins and LCDD amorphous aggregates are a trapped intermediate state (most likely off-pathway) after 72 h of incubation and are unable to proceed to fibrils. It is not known if over time the protein will be able to leave this trapped state and go on to form fibrils.
The mutations for the various BJP in this study can be located throughout the protein, but they tend to cluster in certain areas of the variable domain [28]. The variable domain has an immunoglobulin fold consisting of two antiparallel b-sheets packed together and joined by a disulfide bond [29]. AL proteins have the largest number of mutations out of the six proteins studied and most of their mutations are located in the b-sheet that is part of the heavy chain/light chain dimer interface. uLCDD-01 has mutations in the top or bottom of the beta barrel. uMM-01, uMM-02, and uLCDD-02 have their mutations located in the N and C terminus strands. In conclusion, we have determined thermodynamic parameters for human derived BJP, we have characterized the aggregation properties and described their differences. This study has shed some light into the differences among BJP involved in different light chain diseases.
We would like to thank the
This is a consensus document produced by expert members of a Working Party established by the Association of Anaesthetists of Great Britain and Ireland (AAGBI). It has been seen and approved by the Council of the AAGBI.
The AAGBI acknowledges the contribution received from the British Committee for Standards in Haematology Transfusion Task Force and the Appropriate Use of Blood Group. Summary 1. Hospitals must have a major haemorrhage protocol in place and this should include clinical, laboratory and logistic responses. 2. Immediate control of obvious bleeding is of paramount importance (pressure, tourniquet, haemostatic dressings). 3. The major haemorrhage protocol must be mobilised immediately when a massive haemorrhage situation is declared. 4. A fibrinogen < 1 g.l )1 or a prothrombin time (PT) and activated partial thromboplastin time (aPTT) of > 1.5 times normal represents established haemostatic failure and is predictive of microvascular bleeding.
Early infusion of fresh frozen plasma (FFP; 15 ml.kg )1 ) should be used to prevent this occurring if a senior clinician anticipates a massive haemorrhage. 5. Established coagulopathy will require more than 15 ml.kg )1 of FFP to correct. The most effective way to achieve fibrinogen replacement rapidly is by giving fibrinogen concentrate or cryoprecipitate if fibrinogen is unavailable.
6. 1:1:1 red cell:FFP:platelet regimens, as used by the military, are reserved for the most severely traumatised patients. 7. A minimum target platelet count of 75 • 10 9 .l )1 is appropriate in this clinical situation. 8. Group-specific blood can be issued without performing an antibody screen because patients will have minimal circulating antibodies. O negative blood should only be used if blood is needed immediately. 9. In hospitals where the need to treat massive haemorrhage is frequent, the use of locally developed shock packs may be helpful. 10. Standard venous thromboprophylaxis should be commenced as soon as possible after haemostasis has been secured as patients develop a prothrombotic state following massive haemorrhage.
Re-use of this article is permitted in accordance with the Creative Commons Deed, Attribution 2.5, which does not permit commercial exploitation.
There are an increasing number of severely injured patients who present to hospital each year. Trauma is the leading cause of death in all ages from 1 to 44 years. Haemorrhagic shock accounts for 80% of deaths in the operating theatre and up to 50% of deaths in the first 24 h after injury. Only 16% of major emergency departments in the UK use a massive haemorrhage guideline [1].
The management of massive haemorrhage is usually only one component of the management of a critically unwell patient. These guidelines are intended to supplement current resuscitation guidelines and are specifically directed at improving management of massive haemorrhage [2]. The guidance is intended to provide a better understanding of the priorities in specific situations. Effective teamwork and communication are an essential part of this process.
Definitions of massive haemorrhage vary and have limited value. The Working Party suggests that the nature of the injury will usually alert the anaesthetist to the probability of massive haemorrhage and can be arbitrarily considered as a situation where 1-1.5 blood volumes may need to be infused either acutely or within a 24-h period.
The formulation of guidance in the style of previous AAGBI guidelines has been difficult in such a rapidly changing area. The grade of evidence has not been mentioned within the text, but the editing of the final draft and Working Party membership have been cross-checked with other recently published documents. The Working Party believes that at the current time, its advice is consistent with recently published European guidelines and the availability of current evidence [3][4][5].
It is envisaged that the website version of this document will be updated at least annually and earlier if an addendum or correction is deemed urgent.
Hospitals must have a major haemorrhage protocol in place and this should include clinical, laboratory and logistic responses. Protocols should be adapted to specific clinical areas. It is essential to develop an effective method of triggering the appropriate major haemorrhage protocol.
The team leader is the person who declares a massive haemorrhage situation; this is usually the consultant or the most senior doctor at the scene. Their role is to direct and co-ordinate the management of the patient with massive haemorrhage.
The team leader should appoint a member of the team as communications lead whose sole role is to communicate with the laboratories and other departments.
A member of the team should be allocated to convey blood samples, blood and blood components between the laboratory and the clinical area. This role is usually taken by a porter or healthcare support worker who should ideally be in constant radio communication with the team. In their absence, a nurse or doctor should be identified to take on this role.
A member of the team should be identified whose role is to secure intravenous access, either peripherally or centrally. Large-bore 8-Fr. central access is the ideal in adults; in the event of failure, intra-osseous or surgical venous access may be required.
The switchboard must alert certain key clinical and support people when a massive haemorrhage situation is declared. These include
There are both clinical and logistic issues to consider. These include clinical management of the patient, setting processes in place to deliver blood and blood components to the patient, and organisation of emergency interventions to stop the bleeding (surgical or radiological). There are two common scenarios:
A massively bleeding ⁄ injured ⁄ ill patient en route With warning, resources and personnel can be mobilised to be in position to receive the patient. A brief history can alert the team to the risk of massive bleeding: Immediate actions in dealing with a patient with massive haemorrhage • Control obvious bleeding points (pressure, tourniquet, haemostatic dressings) • Administer high F I O 2 • IV access -largest bore possible including central access • If patient is conscious and talking and a peripheral pulse is present, the blood pressure is adequate.
• Baseline bloods -full blood count (FBC), prothrombin time (PT), activated partial thomboplastin time (aPTT), Clauss fibrinogen* and cross-match.
• If available, undertake near-patient testing e.g.
• Fluid resuscitation -in the massive haemorrhage patient, this means warmed blood and blood components. In terms of time of availability, blood group O is the quickest, followed by group specific, then crossmatched blood.
• Actively warm the patient and all transfused fluids.
• Next steps: rapid access to imaging (ultrasound, radiography, CT), appropriate use of focused assessment with sonography for trauma scanning and ⁄ or early whole body CT if the patient is sufficiently stable, or surgery and further component therapy.
• Alert theatre team about the need for cell salvage autotransfusion. *A derived fibrinogen is likely to be misleading and should not be used.
• Look at injury patterns • Look for obvious blood loss (on clothes, on the floor, in drains) • Look for indications of internal blood loss • Assess physiology (skin colour, heart rate, blood pressure, capillary refill, conscious level) Some patients compensate well despite significant blood loss. A rapid clinical assessment will give very strong indications of those at risk. It is important to restore organ perfusion, but it is not necessary to achieve a normal blood pressure at this stage [6][7][8].
Once control of bleeding is achieved, aggressive attempts should be made to normalise blood pressure, acid-base status and temperature, but vasopressors should be avoided. Active warming is required. Coagulopathy should be anticipated and, if possible, prevented. If present, it should be treated aggressively (see Dealing with coagulation problems).
Surgery must be considered early. However, surgery may have to be interrupted and limited to 'damage control'. Once bleeding has been controlled, abnormal physiology can be corrected [9][10][11].
Following treatment for massive haemorrhage, the patient should be admitted to a critical care area for monitoring and observation, and monitoring of coagulation, haemoglobin and blood gases, together with wound drain assessment to identify overt or covert bleeding.
Standard venous thromboprophylaxis should be commenced as soon as possible after bleeding has been controlled, as patients rapidly develop a prothrombotic state. Temporary inferior vena cava filtration may be necessary.
The haemostatic defect in massive haemorrhage will vary, depending on the amount and cause of bleeding and underlying patient-related factors. It is likely to evolve rapidly. Patient management should be guided by laboratory results and near-patient testing, but led by the clinical scenario.
All patients being treated for massive haemorrhage are at risk of dilutional coagulopathy leading to reduced platelets, fibrinogen and other coagulation factors. This occurs if volume replacement is with red cells, crystalloid and plasma expanders, and insufficient infusion of fresh frozen plasma (FFP) and platelets. Dilutional coagulopathy should be prevented by early infusion of FFP.
Some patients with massive haemorrhage are also at risk of a consumptive coagulopathy and are liable to develop haemostatic failure without significant dilution. Consumption is commonly seen in obstetric haemorrhage, particularly associated with placental abruption and amniotic fluid embolus, in patients on cardiopulmonary bypass (CPB), following massive trauma especially involving head injury, and in the context of sepsis.
Activation of anticoagulant pathways is associated with massive trauma and patients may have haemostatic compromise without abnormal coagulation tests [12].
Platelet dysfunction is associated with CPB, renal disease and anti-platelet medication.
Hyperfibrinolysis is particularly associated with obstetric haemorrhage, CPB and liver surgery.
In the context of massive haemorrhage, warfarin should be reversed with a prothrombin complex concentrate (PCC) and intravenous vitamin K (5-10 mg). The dose is dependent on the international normalised ratio (INR) (see Table 1).
Unfractionated heparin can be reversed with protamine (1 mg protamine reverses 100 u heparin). Excess protamine induces a coagulopathy. Usual reversal is by infusing either 25 or 50 mg of intravenous protamine.
Low molecular weight heparin can be partially reversed with protamine.
Direct thrombin and factor Xa inhibitors e.g. fondaparinux, dabigatran and rivaroxaban cannot be reversed.
Patients taking aspirin have a low risk of increased bleeding, whilst those on P2Y12 antagonists have a higher risk. The anti-platelet effect of aspirin can be reversed by platelet transfusion, but the effect of the P2Y12 antagonist, clopidogrel, is only partially reversed by platelets.
It is very likely that patients with an inherited bleeding disorder will be registered with a haemophilia centre and urgent advice should be sought if they present with massive haemorrhage.
Liver disease is associated with decreased production of coagulation factors, natural anticoagulants and the production of dysfunctional fibrinogen (dysfibrinogenaemia). It should be anticipated that these patients will develop a clinically significant dilutional coagulopathy and haemostatic failure with bleeds less than one blood volume.
Interpretation of laboratory tests A fibrinogen < 1 g.l )1 or a PT and aPTT > 1.5 times normal represents an established haemostatic failure and is predictive of microvascular bleeding. Early infusion of FFP should be used to prevent this occurring if a senior clinician anticipates a massive haemorrhage.
Clauss fibrinogen is an easily available test and should be specifically requested if not part of the routine coagulation screen. The fibrinogen level is more sensitive than the PT and aPTT to a developing dilutional or consumptive coagulopathy. Levels below 1 g.l )1 , in the context of massive haemorrhage, are usually insufficient, and emerging evidence suggests that a level above 1.5 g.l )1 is required. Higher levels are likely to improve haemostasis further.
A platelet count below 50 • 10 9 .l )1 is strongly associated with haemostatic compromise and microvascular bleeding in a patient being treated for massive haemorrhage. A minimum target platelet count of 75 • 10 9 .l )1 is appropriate in this clinical situation.
The PT is an insensitive test for haemostatic compromise and a relatively normal result should not necessarily reassure the clinician. It is common practice to correct to PT to within 1.5 of normal; however, this may not be an appropriate target in many situations [13].
An INR is not an appropriate test in massive haemorrhage because it is standardised for warfarin control, and results may be misleading in the context of dilutional and consumptive coagulopathies and liver disease.
The aPTT is commonly used to guide blood product replacement but, as with the PT, correcting to 1.5 times normal is not necessarily an appropriate strategy because haemostatic failure may already be significant at this level. The aPTT should be maintained below 1.5 times normal as the minimum target.
If whole blood point of care testing is used, a protocol for blood product usage based on thromboelastogram (TEG ⁄ ROTEM) results should be agreed in advance.
Haemostatic tests and FBC should be repeated at least every hour if bleeding is ongoing, so that trends may be observed and adequacy of replacement therapy documented. Widespread microvascular oozing is a clinical marker of haemostatic failure irrespective of blood tests and should be treated aggressively.
The coagulopathy during massive haemorrhage is likely to evolve rapidly and regular clinical review and blood tests are required. It is important to anticipate and prevent haemostatic failure, but if haemostatic failure has occurred, standard regimens (e.g. FFP 15 ml.kg )1 ) can be predicted to be inadequate and larger volumes of FFP are likely to be required.
Emerging evidence supports the early use of FFP to prevent dilutional coagulopathy. If an experienced clinician anticipates a blood loss of one blood volume, FFP should be infused to prevent coagulopathy. While FFP 15 ml.kg )1 is appropriate for uncomplicated cases, increased volumes of FFP will be needed if a consumptive Blood transfusion and the anaesthetist: management of massive haemorrhage coagulopathy is likely or the patient has underlying liver disease.
A minimum target platelet count of 75 • 10 9 .l )1 is appropriate in this clinical situation.
1:1:1 red cell:FFP:platelet regimens, as used by the military, are reserved for the most severely traumatised patient and are not routinely recommended [14,15].
In the context of massive haemorrhage, patients with widespread microvascular oozing or with coagulation tests that demonstrate inadequate haemostasis (fibrinogen < 1 g.l )1 or PT ⁄ aPTT > 1.5 above normal), should be given FFP in doses likely to correct the coagulation factor deficiencies. This will require more than 15 ml.kg )1 , and at least 30 ml.kg )1 would be a reasonable first-line response [16,17].
Platelets should be maintained at at least 75 • 10 9 .l )1 [2,18].
Although it is often recommended that hypofibrinogenaemia unresponsive to FFP be treated with cryoprecipitate, treatment may be associated with delays because of thawing and transportation.
Fibrinogen replacement can be achieved much more rapidly and predictably with fibrinogen concentrate (no requirement for thawing as for cryoprecipitate) given at a dose of 30-60 mg.kg )1 . This product is not currently licensed in the UK and must be given on a named patient basis. Fibrinogen concentrate is licensed in many European countries to treat both congenital and acquired hypofibrinogenaemia [19].
Intravenous tranexamic acid should be used in clinical situations where increased fibrinolysis can be anticipated. Support for its use has strengthened recently with the positive report of its use in traumatic haemorrhage [20] (see Other interventions on pharmacological management of massive haemorrhage).
Hypocalcaemia and hypomagnesaemia are often associated with massively transfused patients and will need monitoring and correction. rFVIIa This drug has been used for treatment of massive haemorrhage unresponsive to conventional therapy. Recent review of data has highlighted the risk of arterial thrombotic complications and the specification of product characteristics now states: 'Safety and efficacy of NovoSeven (rFVIIa) have not been established outside the approved indications and therefore NovoSeven should not be used'. Where centres decide to use this therapy, local protocols must be agreed in advance.
rFVIIa is usually given with tranexamic acid and is not as efficacious if the patient has a low fibrinogen.
Some centres use PCC (concentrated factors II, VII, IX and X) in certain clinical situations such as liver disease and post-CPB; local protocols must be agreed in advance.
Positive patient identification is essential at all stages of the blood transfusion process and a patient should have two identification bands in situ. The healthcare professional administering the blood component must perform the final administrative check for every component given. All persons involved in the administration of blood must be trained and certificated in accordance with national standards.
Pre-transfusion procedures are designed to determine the patient's ABO and Rhesus D (RhD) status, to detect red cell antibodies that could haemolyse transfused cells and confirm compatibility with each of the units of red cells to be transfused. Red cell selection may be based on a serological cross-match or electronic issue. Standard issue of red cells may take approximately 45 min.
Group O RhD negative is the blood group of choice for transfusion of red cells in an emergency where the clinical need is immediate. However, overdependence on group O RhD negative red cells may have an adverse impact on local and national blood stock management and it is considered acceptable to give O RhD positive red cells to male patients. Hospitals should avoid the need for elective transfusion of group O RhD negative red cells to non-O RhD negative recipients. Clinical staff should endeavour to provide immediate blood samples for grouping in order to allow the use of group specific blood.
In the emergency situation, blood can be issued following identification of group without knowing the result of an antibody screen -'group specific blood'. Grouping can be performed in about 10 min, not including transfer time, and group specific blood can be issued. This of course is a higher risk strategy and depends on the urgency for blood. In massive bleeding, patients will have minimal circulating antibodies, so will usually accept group specific blood without reaction. If the patient survives, antibodies may develop at a later stage.
Women who are RhD negative and of childbearing age, who are resuscitated with Rh D positive blood or platelets, can develop immune anti-D, which can cause haemolytic disease of the newborn in subsequent pregnancies. To prevent this, a combination of exchange transfusion and anti-D can be administered, on the advice of a haematologist, within 72 h of the transfusion.
Cold chain requirements are essential under European Law. Blood should be transfused within 4 h of leaving a controlled environment. Blood issued cannot be returned to stock if out of a controlled temperature and monitored fridge for longer than 30 min. If blood is issued within a correctly packed and validated transport box, the blood should be placed back in a blood fridge normally within 2 h, providing the box is unopened. Blood transfusion laboratory staff will then assess the acceptability of blood for return to stock. Only in exceptional circumstances should blood components be transferred with a patient between hospital trusts or health boards.
It is a statutory requirement that the fate of all blood components must be accounted for. These records must be held for 30 years. Staff must be familiar with local protocols for recording blood use in clinical notes and for informing the hospital transfusion laboratory.
The hospital transfusion committee is the ideal forum to allow cross-specialty discussions about protocols and organisation for dealing with massive haemorrhage. Audit of previous instances allows a refinement of response to ensure efficient and timely treatment. This level of organisation can only be arranged at a local level.
Stock management of labile components when there is unpredictable demand is a challenge. Large stock-holding is associated with wastage, whereas insufficient stocks may lead to clinical disaster.
Most hospitals rely on rapid re-supply of platelets from the Blood Service rather than holding stocks. Anaesthetists need to be aware of local arrangements and the normal time interval for obtaining platelets in an emergency.
National demand for blood components may exceed supply. National blood shortage plans will be activated in the event of red cell and platelet shortages. Guidance is given for the prioritisation of patient groups during shortage. The transfusion support during massive haemorrhage is a priority. However, it is expected that all efforts be made to stop the bleeding and reduce the need for donor blood. The use of cell salvage is encouraged in all cases of massive haemorrhage (Further information is available in Blood Transfusion and the Anaesthetist -Intraoperative Cell Salvage. AAGBI: http://www.aagbi. org/publications/guidelines/docs/cell%20_salvage_2009_ amended.pdf).
This section advises on the appropriate use of blood components during massive haemorrhage (see Dealing with coagulation problems). This advice is required because red cell concentrates do not contain coagulation factors or platelets.
Patients with massive haemorrhage may require all blood components. Blood may be required not just at the time of resuscitation, but also during initial and repeat surgery. The benefits of timely and appropriate transfusion support in this situation outweigh the potential risks of transfusion and may reduce total exposure to blood components.
(Further information is available in Blood Transfusion and the Anaesthetist -Blood Component Therapy. AAGBI: http://www.aagbi.org/publications/guidelines/docs/ bloodtransfusion06.pdf).
A comprehensive guideline for neonatal and paediatric transfusion together with a recent update statement is available at http://www.bcshguidelines.com/. Useful principles are: minimise and stop blood loss; minimise donor exposure; and use paediatric components where readily available (See Table 2).
All blood components should be administered using a blood component administration set, which incorporates a 170-200 lm filter.
There is no current need to use any sort of additional filter in massive haemorrhage when using allogeneic product, as pre-storage leucodepletion has rendered this process unnecessary. If red cell salvage is being used, a 40-lm filter may still be indicated e.g. if small bone fragments contaminate the surgical field.
Although a special platelet giving set is ideally used for platelet transfusion, it is unnecessary in massive haemorrhage. The important issue is to administer platelets via a clean 170-200 lm giving set (see below), as one that has previously been used for red cells may cause the platelets to stick to the red cells and therefore reduce the effective transfused platelet dose.
The use of an adequate warming device is recommended in massively bleeding patients and this equipment needs to be available in all emergency rooms and theatre suites, allowing adequate warming of administered blood at high infusion rates.
• Only blood component administration sets that are compatible with the infusion device should be used (check manufacturers' recommendations). Infusion devices should be regularly maintained and any adverse outcome as a result of using an infusion device to transfuse red cells should be reported to the Medical Devices section of the Medicines and Healthcare products Regulatory Agency (MHRA).
• Administration sets used with infusion devices should incorporate an integral mesh filter (170-200 lm).
• The pre-administration checking procedure should include a check of the device and device settings.
• Either gravity or electronic infusion devices may be used for the administration of blood and blood components. Infusion devices allow a precise infusion rate to be specified.
• Rapid infusion devices may be used when large volumes have to be infused quickly, as in massive haemorrhage. These typically have a range of 6-30 l.h )1 and usually incorporate a blood-warming device.
• Infusion devices should only be used if the manufacturer verifies them as safe for this purpose and they are CE-marked.
• The volume delivered should be monitored regularly throughout the infusion to ensure that the expected volume is delivered at the required rate.
Pressure devices • External pressure devices make it possible to administer a unit of red cells within a few minutes. They should only be used in an emergency situation together with a large-gauge venous access cannula or device.
• External pressure devices should:
• Exert pressure evenly over the entire bag;
• Have a gauge to measure the pressure;
• Not exceed 300 mmHg of pressure;
• Be monitored at all times when in use.
• In all adults undergoing elective or emergency surgery (including surgery for trauma) under general or regional anaesthesia, 'intravenous fluids (500 ml or more) and blood components should be warmed to 37 °C' [21,22]. The greatest benefit is from the controlled warming of red cells (stored at 4 °C) rather than platelets (stored at 22 ± 2 °C) or FFP ⁄ cryoprecipitate (thawed to 37 °C) [23]. Of note, there is no evidence to suggest that infusion of platelets or FFP through a blood warmer is harmful.
• In most other clinical situations where there is concern, it is sufficient to allow blood to rise to ambient temperature before transfusion. Special consideration should be given when rapidly transfusing large volumes to neonates, children, elderly patients and patients susceptible to cardiac dysfunction.
• Blood should only be warmed using approved, specifically designed and regularly maintained blood warming equipment with a visible thermometer and audible warning. Settings should be monitored regularly throughout the transfusion.
• Blood components should never be warmed using improvisations, such as putting the pack in warm water, in a microwave or on a radiator.
Fibrinolysis is the process whereby established fibrin clot is broken down. This can occur in an accelerated fashion, destabilising effective coagulation in many clinical situations associated with massive haemorrhage, including multiple trauma, obstetric haemorrhage and major organ surgery (e.g. cardiothoracic, liver) including transplantation surgery. Accelerated fibrinolysis can be identified by laboratory assay of d-dimers or fibrin degradation products, or by use of coagulation monitors such as TEG or ROTEM. It is accepted that not all hospitals can provide either the hardware or expertise to interpret TEG and ROTEM.
Antifibrinolytic drugs, such as tranexamic acid, have been used to reverse established fibrinolysis in the setting of massive blood transfusion.
Whilst systematic reviews fail to demonstrate evidence from randomised controlled trials to support the routine use of antifibrinolytic agents in managing massive haemorrhage, they are considered effective if accelerated fibrinolysis is identified.
Tranexamic acid inhibits plasminogen activation, and at high concentration inhibits plasmin. The recent CRASH-2 trial supports its use at a loading dose of 1 g over 10 min followed by 1 g over 8 h [20]. There are few adverse events or side effects associated with tranexamic acid use in the setting of massive haemorrhage [20]. Repeat doses should be used with caution in patients with renal impairment, as the drug is predominantly excreted unchanged by the kidneys. It is contraindicated in patients with subarachnoid haemorrhage, as anecdotal experience suggests that cerebral oedema and cerebral infarction may occur.
Aprotinin is a serine protease inhibitor, inhibiting trypsin, chymotrypsin, plasmin and kallikrein. It has been used to reduce blood loss associated with accelerated fibrinolysis in major surgery (e.g. cardiothoracic surgery, liver transplantation). Recently, there have been concerns about the safety of aprotinin. Anaphylaxis occurs at a rate of 1:200 in first-time use. A study performed in cardiac surgery patients reported in 2006 showed that there was indeed a risk of acute renal failure, myocardial infarction and heart failure, as well as stroke and encephalopathy [24]. As a result of this, and other follow-up work, the MHRA recommends that aprotinin should only be used when the likely benefits outweigh any risks to individual patients. As a result, use of aprotinin is now limited to highly specialised surgical situations, e.g. cardiac and liver transplantation, and is used on a named patient basis only.
Coagulation factor concentrates may be required for patients with inherited bleeding disorders such as haemophilia or von Willebrand disease. They should only be used under the guidance of a haemophilia centre. (For recombinant factor VIIa, prothrombin complex concentrate and fibrinogen concentrate, see Dealing with coagulation problems).
Radiologically aided arterial embolisation These techniques are becoming more widespread and successful cessation of bleeding can be achieved with embolisation of bleeding arteries following angiographic imaging. The suitability of such manoeuvres needs to be assessed in each individual case and will also depend on availability of an interventional radiologist. The technique can be remarkably effective and may eliminate the need for surgical intervention, particularly in major obstetric haemorrhage.
The use of intra-operative cell salvage can be very effective at both reducing demand on allogeneic supplies and providing a readily available red cell supply in massive haemorrhage. National Institute of Health and Clinical Excellence (NICE) guidelines have also supported its use where large blood loss is experienced in obstetric haemorrhage and complex urological surgery such as radical prostatectomy. The indications for cell salvage are detailed in Blood Transfusion and the Anaesthetist -Intra-operative Cell Salvage (http://www.aagbi.org/ publications/guidelines/docs/cell%20_salvage_2009_ amended.pdf).
Ó 2010 The Authors Anaesthesia Ó 2010 The Association of Anaesthetists of Great Britain and Ireland
Ó 2010 The Authors
Anaesthesia Ó 2010 The Association of Anaesthetists of Great Britain and Ireland
Soluble epoxide hydrolase (sEH) is a promising therapeutic target for the treatment of hypertension, pain, and inflammation-related diseases. In order to enable the development of sEH inhibitors (sEHIs), assays are needed for determination of their potency. Therefore, we developed a new method utilizing an epoxide of arachidonic acid (14 (15)-EpETrE) as substrate. Incubation samples were directly injected without purification into an online solid phase extraction (SPE) liquid chromatography electrospray ionization tandem mass spectrometry (LC-ESI-MS-MS) setup allowing a total run time of only 108 s for a full gradient separation. Analytes were extracted from the matrix within 30 s by turbulent flow chromatography. Subsequently, a full gradient separation was carried out on a 50X2.1 mm RP-18 column filled with 1.7 μm core-shell particles. The analytes were detected with high sensitivity by ESI-MS-MS in SRM mode. The substrate 14(15)-EpETrE eluted at a stable retention time of 96±1 s and its sEH hydrolysis product 14,15-DiHETrE at 63±1 s with narrow peak width (full width at half maximum height: 1.5±0.1 s). The analytical performance of the method was excellent, with a limit of detection of 2 fmol on column, a linear range of over three orders of magnitude, and a negligible carry-over of 0.1% for 14,15-DiHETrE. The enzyme assay was carried out in a 96-well plate format, and near perfect sigmoidal dose-response curves were obtained for 12 concentrations of each inhibitor in only 22 min, enabling precise determination of IC 50 values. In contrast with other approaches, this method enables quantitative evaluation of potent sEHIs with picomolar potencies because only 33 pmol L -1 sEH were used in the reaction vessel. This was demonstrated by ranking ten compounds by their activity; in the fluorescence method all yielded IC 50 ≤ 1 nmol L -1 . Comparison of 13 inhibitors with IC 50 values >1 nmol L -1 showed a good correlation with the fluorescence method (linear correlation coefficient 0.9, slope 0.95, Spearman's rho 0.9). For individual compounds, however, up to eightfold differences in potencies between this and the fluorescence method were obtained. Therefore, enzyme assays using natural substrate, as described here, are indispensable for reliable determination of structure-activity relationships for sEH inhibition.
Soluble epoxide hydrolase (sEH) inhibitors are a promising new class of potential drugs for treatment of a variety of diseases, for example inflammation, hypertension, and pain [1,2]. In order to develop new sEH inhibitors (sEHI) analytical techniques are needed to identify active compounds and quantitatively measure their potencies. Several
The online version of this article (doi:10.1007/s00216-011-4861-2) contains supplementary material, which is available to authorized users.
in-vitro assays have been described utilizing surrogate substrates [3], for example cyano(6-methoxynaphthalen-2yl)methyl trans-[(3-phenyloxiran-2-yl)methyl] carbonate (CMNPC) [4,5] or tritium-labeled trans-diphenylpropene oxide (t-DPPO) [6]. However, because of the different recognition of dissimilar substrates by the enzyme, the measured potencies of sEHIs may differ among these methods. In order to obtain results predictive for in-vivo potency inhibition, assays utilizing the natural substrates are advantageous. Modern mass spectrometry (MS) enables parallel measurement of many natural enzyme substrates and products and is, thus, an excellent tool for measurement of enzyme activity and inhibition [7][8][9][10][11]. For the sEH, known natural substrates are epoxy fatty acids, which are metabolized to their corresponding fatty acid diols [12,13]. Among the epoxy fatty acids, arachidonic acid epoxides (EpETrEs) are best characterized. These have several biological effects, for example vasodilatory, antiinflammatory, and analgesic activity [1,2,[14][15][16][17]. EpETrEs and their corresponding diols (DiHETrEs) can be sensitively detected by liquid chromatography electrospray (LC-ESI) MS [18,19]. Consequentially, LC-ESI-MS has already been used to monitor conversion of 14(15)-EpETrE to 14,15-DiHETrE [3]. However, no LC-MSbased approach using natural a substrate has been described for the rapid determination of the potency of sEHI. For maximum sEH activity in cell-free in-vitro assays, volatile salts and stabilizing protein BSA are usually present in high concentrations [3]. Therefore, direct injection of these samples on conventional LC columns may lead to an irreversible absorption of proteins on the stationary phase, resulting in loss of chromatographic efficiency [20]. Moreover ESI-MS detection is significantly affected by this matrix, because of signal suppression or enhancement [21]. Matrix effects can still occur even when most of the proteins have been precipitated by organic solvent and removed by centrifugation [22]. Thus, a sample preparation step is needed before LC-ESI-MS analysis to ensure sensitive and reliable determination of small amounts of product formed in a difficult matrix. One fully automatable strategy is application of online solid-phase extraction (SPE), which enables direct injection of crude samples [23][24][25]. One of the most promising techniques for online SPE of protein-containing samples is the application of short, narrow columns filled with large particles (50-60 μm) [23][24][25]. At high flow rates, turbulent flow results, enhancing mass transfer between the mobile and stationary phases. This enables the separation of the small analyte molecules from the matrix, because of the larger diffusion coefficient of proteins [23][24][25]. Turbulent-flow chromatography (TFC) significantly reduces matrix effects by proteins in LC-ESI-MS quantification [26], and has found broad application in bioanalytical research, particularly for the analysis of drugs in biological samples [23][24][25]27]. TFC was recently introduced as sample preparation for ESI-MS based enzyme inhibition assays by Vogel and coworkers [11]. In this work, we developed one of the fastest online SPE-LC-ESI-MS-MS methods described, by combining TFC with a separation on a sub-2 μm core-shell particlefilled RP-18 separation column. Together with a streamlined sEH inhibition assay in plate format, this method enables ranking of sEHIs with picomolar potencies by utilizing the endogenous substrate 14(15)EpETrE.
14(15)-Epoxy-eicosatrienoic acid (14(15)-EpETrE) and 14,15-dihydroxyeicosatrienoic acid (14,15-DiHETrE) (Fig. 1) were purchased from Cayman Chemicals (Ann Arbor, MI, USA). The internal standards, 10(11)-epoxydecaheptanoic acid (10(11)-EpHep) and 10,11-dihydroxydecaheptanoic acid (10,11-DiHHep) were synthesized as described elsewhere [18]. Urea derivatives previously synthesized in our laboratory were used as epoxide hydrolase inhibitors (sEHI) [28][29][30][31][32]. The chemical structures of four of the inhibitors are shown in Figs. 2 and 3. All other chemicals were obtained from Fischer Scientific (Pittsburgh, PA, USA) and were of the highest quality available. Baculovirus-expressed human soluble epoxide hydrolase (sEH) was purified by affinity column chroma- tography and its 100 μmol L -1 stock solutions in sodium phosphate buffer (100 mmol L -1 pH 7.4) was kept at -80 °C until use [33].
Preparation of stock and standard solutions Stock solutions of 14(15)-EpETrE, 14,15-DiHETrE, 10 (11)-EpHep, 10,11-DiHHep, and sEHIs were prepared in DMSO and kept at 4 °C. All solutions for the assay were prepared on the day of analysis in 0.1 mol L -1 sodium phosphate buffer (pH 7.4) containing 0.1 gL -1 bovine serum albumin (BSA) and kept on ice until incubation. For the assay a 1 μmol L -1 enzyme solution of sEH was prepared from the stock solution in buffer; this was stable at 4 °C for seven days. Before the incubation, this solution was further diluted with buffer to concentrations of 0.1, 0.33, and 1 nmol L -1 . The substrate solution was freshly prepared by 1:100 dilution from the 1.5 mmol L -1 DMSO stock solution to 14,(15)-EpETrE in buffer (15 μmol L -1 , 1% DMSO).
Inhibitor solutions of 3, 10, 30, and 100 μmol L -1 were prepared by diluting stock solutions in buffer (final DMSO concentration: 1%) within an hour before analysis and subsequently further diluted as described in the section "sEH inhibition assay". The reaction quench solution consisted of 20 μmol L -1 sEHI (1) and 200 nmol L -1 of the I.S. 10(11)-EpHep and 10,11-DiHHep in ACN-water 50:50 (v/v). For calibration a 0.5 mmol L -1 DMSO solution of 14(15)-EpETrE and 14,15-EpETrE was sequentially diluted in a 96-well plate in 50:50 (v/v) ACN-water. Subsequently, the dilutions were mixed 1:1 with quench solution (final concentration 0.04 nmol L -1 -2.5 μmol L -1 ) and analyzed in the same way as samples.
Online SPE-LC was performed in back-flush mode utilizing a setup similar to that recently described [27]. In brief, the analytes (injection volume 20 μL) were extracted using a (50×0.5 mm, 50 μm particle size) Cyclone RP-18 column (Thermo Fisher Scientific, Waltham, MA, USA) at a flow rate of 1500 μL min -1 0.1% acetic acid. After 30 s the sixport valve was switched, and the analytes were separated by a 78-s binary gradient on a (50×2.1 mm) Kinetex coreshell reversed-phase column (Phenomenex, Torrance, CA, USA) at a flow rate of 500 μL min -1 . Mass spectrometric detection was carried out in selected reaction monitoring mode (SRM) on an ABI 4000 TRAP tandem mass 0.01 0.1 1 10 100 1000 0 20 40 60 80 100 concentration sEHi (nM) direct analysis storage for 7 days (-30°C) 0.01 0.1 1 10 100 1000 0 20 40 60 80 100 human sEH inhibition % concentration sEHi (nM) day 3 day 1 day 2 B A O F F F H N O NH N S O O CH 3 Fig. 2 Reproducibility and robustness of the method. Shown are the dose-response curves for sEHI 1. a Means of three independent determinations on three different days, using two different batches of substrate and enzyme. b Comparison of a direct measurement after incubation and after 7 days storage at -30 °C. The structure of the inhibitor is shown in panel A
spectrometer after negative electrospray ionization (ESI). Details of the instruments used, the gradients applied, and the mass spectrometric conditions are presented in the Electronic Supplementary Material.
All sEH incubations were carried out in polypropylene 96well plates (Fisher Scientific) in a heated (30 °C) shaker. For optimizing sEH concentration and incubation time, all wells were filled with 50 μL buffer. Enzyme solutions of 1, 0.33, and 0.1 nmol L -1 and buffer as control (50 μL) were added to two rows for each concentration, by use of a twelve-channel pipette. Following pre-incubation for 5 min the conversion was started by adding 50 μL substrate solution. After 0, 5, 10, 20, and 40 min, 150 μL of quench solution was added to two rows at each time point. After incubation the plate was kept at 4 °C until online LC-MS-MS analysis.
For the sEHI potency assay, wells in columns 2-12 were filled with 50 μL buffer. Column 1 wells were filled with 75 μL inhibitor solution (2-3 wells for each inhibitor). By use of an eight-channel pipette, 25 μL from each cell in row 1 was transferred to the cells in row 2 and mixed three times with the pipette. This procedure was repeated for rows 2-12, and, finally, 25 μL from row 12 cells was removed and discarded. No inhibitor was added to wells G12 and H12, which served as positive and negative control, respectively. Thereafter, 50 μL enzyme solution was added to all wells (except the negative control H12, to which 50 μL additional buffer was added instead) and the plate was preincubated for 5 min. After addition of 50 μL substrate solution (0.1 μmol L -1 ) to all wells with a twelve channel pipette, the plate was incubated for 30 min. The reaction was stopped by adding 150 μL quench solution to all wells and directly analyzed by online LC-MS-MS, or kept at 4 °C till analysis. The measured effect on sEH activity for each well was calculated as percentage of sEH inhibition based on the area ratio, R, of the product 14,15-DiHETrE and its internal standard, 10,11-DiHHep by use of Eq.( 1).
The IC 50 for the sEHI were calculated by fitting the dose-response curves of percentage inhibition values vs. log concentration with Origin 7.0 (OriginLab, MA, USA).
The potency of the sEHI was compared with the value from the commonly used fluorescence assay utilizing cyano (6-methoxynaphthalen-2-yl)methyl trans-[(3-phenyloxiran-2-yl)methyl] carbonate (CMNPC) as substrate as described elsewhere [3,4].
In the development of the new sEH inhibition assay a major emphasis was set on a rapid method of detection of the substrate 14(15)-EpETrE and its hydrolysis product 14,15-DiHETrE (Fig. 1). The odd-chain fatty acid epoxide 10, (11)-EpHep and its corresponding diol, 10,11-DiHHep (Fig. 1) were used as internal standards. These were chosen because they have similar physicochemical properties to the arachidonate oxylipins, and they do not occur biologically in relevant amounts. The crude samples arising from enzymatic incubations were mixed with I.S. solution and directly analyzed by online-SPE-LC-MS-MS (Electronic Supplementary Material Fig. S1). The analytes were completely extracted by the online-SPE column and no analyte was detected in the flow through up to an elution volume of 6 mL (Electronic Supplementary Material Fig. S2). Salts and protein were directed to waste, and a minimum extraction time of 30 s corresponding to elution of 20 void volumes of the SPE column was found to be sufficient. Thereafter, the six-port valve was switched and the analytes were eluted in back-flush mode from the SPE column by the more hydrophobic flow (57% ACN) delivered by pump 2 (Electronic Supplementary Material Fig. S1). The separation was carried out on a short 50× 2.1 mm, 1.7 μm not fully porous "core-shell" particle filled, RP-18 column at a flow rate of 500 μL min -1 . Efficient mass transfer between stationary and mobile phases results from the short diffusion path in the shell type particles [34,35]. In addition to the advantages of sub-2-μm particle size, this column type leads to very high chromatographic resolution [36]. The mobile phase gradient was optimized to fully separate analytes from the void volume (150 μL) of the analytical column where polar matrix compounds coextracted by SPE elute (Fig. 1). The dihydroxy fatty acid I. S. 10,11 DiHHep eluted first at a retention time of 59±1 s followed by 14,15-DiHETrE at 63± 1 s. Despite the proximity of the retention times, these compounds were separated almost to the baseline, because of the narrow peak width, with full width at half maximum height (FWHM) of 1.4 ± 0.1 s and 1.5 ± 0.1 s, respectively (Fig. 1). After elution of the diols, the gradient was increased to 95% ACN over a period of 6 s to elute the epoxides. The 14,(15)-EpETrE and its I.S. 10,(11)-EpHept co-eluted at 96±1 s as narrow peaks (Fig. 1). Although it is not ideal that the substrate and its I.S. co-elute, quantification for the assay was carried out on the basis of product formation, and the product and its I.S. are ideally baseline separated. MS detection of the fatty acid derivatives was carried out after negative ESI in selected reaction monitoring mode (SRM), using the same transitions as previously described [18,19].
A disadvantage inherent in the application of online SPE is the risk of carry over from the previous sample [25]. To investigate carry over, the highest concentration sample (2.5 μmol L -1 ) was injected and the analyte area obtained was compared with that for a subsequent blank injection. We found carry-over was 0.1% for 14,15-DiHETrE and 0.4% for 14(15)-EpETrE. Therefore it is expected there will be negligible interference from carry-over in the assay.
The method was calibrated using a series of standard solutions of 14,15-DiHETrE and 14(15)-EpETrE, which were treated in the same manner as incubation samples. The limit of detection (LOD, S/N=3) for 14,15-DiHETrE and 14(15)-EpETrE was 0.1 nmol L -1 (2 fmol on column) and method provided for both analytes a broad linear detection range over three orders of magnitude (R 2 ≥0.999) from the limit of quantification (LOQ, S/N=9) up to 800 nmol L -1 . To investigate matrix effects on the quantification, we spiked the protein-containing reaction buffer with 10-800 nmol L -1 of the analytes. These concentration ranges correspond to the product concentration at a conversion rate of 0.2-30% of the enzymatic assay (section "sEH inhibition assay"). As shown in Electronic Supplementary Material Fig. S3, recovery of the analytes was within ±20% for all the concentrations tested, so direct injection of quenched incubation samples did not compromise analytical performance. because of the ultra-rapid online-SPE-LC-MS-MS (total analysis time 1.8 min) more than 30 samples can be analyzed in one hour, rendering this approach ideal for enzyme inhibition assays with a large sample sets.
The enzyme inhibition assay was developed in a robust 96well-plate format for rapid investigation of sEHI libraries. In a final volume of 150 μL the assay was carried out by mixing equal volumes (50 μL) of inhibitor (buffer for controls), substrate, and enzyme solution followed by incubation. All pipetting steps were performed with a volume of at least 25 μL, enabling reproducible use of multichannel pipettes and making adaptation for use with fully-automated pipetting robots easily possible. The enzyme assay was carried in 0.1 mol L -1 sodium phosphate buffer (pH 7.4) containing 0.1 gL -1 bovine serum albumin (BSA) at 30 °C, as described previously [3]. The substrate concentration in the assay was set to the K M value of 5 μmol L -1 for the conversion of 14(15)-EpETrE by human sEH [37]. Incubation time and enzyme concentration were optimized to keep the enzyme concentration as low as possible to enable measurement of IC 50 values for potent inhibitors. It was found that an enzyme concentration of only 33.3 pmol L -1 was sufficient. The conversion of 14 (15)-EpETrE was linear over the incubation time. Within 30 min, 298±18 nmol L -1 14,15-DiHETrE was formed. This corresponds to substrate conversion of 5.9±0.4% and thus indicates that the initial velocity of the enzyme reaction was constant over the whole incubation time. Because of the high sensitivity of the online LC-ESI-MS-MS method, low product concentrations can be reliably quantified (section "Online SPE-LC-ESI-MS-MS"). Because no significant non-enzymatic chemical hydrolysis of the substrate occurred (0.13±0.06%), the approach enables quantitative measurement of sEH inhibition with a good signal-to-noise ratio based on product formation. However, the substrate concentration in the samples exceeds the dynamic range of the detection method, and thus could not be monitored simultaneously, although it would be technically feasible to do so. After the incubation, the reaction was stopped by adding an equal volume of quenching solution (150 μL). The ACN content of this solution was adjusted to be as high as possible to denature the enzyme and increase the solubility of the low polar oxylipins. However, concentrations above 50% ACN in water in the quenching solution caused precipitation of proteins and salts, making direct injection into the online SPE-LC-MS-MS system impossible. When 50:50 (v/v) ACN-water was used as the quenching solution, it was also found that sEH was not fully inactivated by the organic solvent and the 14,15-DiHETrE concentration increased in the quenched samples over time. Therefore sEHI (1) was added to the quenching solvent at a high concentration of 20 μmol L -1 to fully inactivate the enzyme. With this optimized quenching solution incubation samples were stable over a storage time of at least 7 days at -20 °C or 4 °C (Fig. 2).
Measurement of sEHI potency was carried out by investigating 14,15-DiHETrE formation in the presence of twelve concentrations of each inhibitor compared with control samples. In a generic scheme, 50 μL of 3 μmol L -1 solutions of sEHI in buffer (1% DMSO) were added to the wells of the first row of a 96-well plate and were sequentially diluted threefold per step. The resulting series of dilutions (final concentration in assay 6 pmol L -1 -1 μmol L -1 ) enabled the direct determination of sEHI potency within a wide dynamic range. To ensure the accuracy of the measurements, at least three concentrations above and three concentrations below the IC 50 should be investigated [38]. The dilution series therefore enabled the evaluation of sEHI with an IC 50 range of 0.1 nmol L -1 to 100 nmol L -1 , without any modification of this procedure. However, for investigation of sEHI of low potency (high IC 50 value), the initial concentration of the inhibitor solution was increased to 3.3, 10, and 33.3 μmol L -1 . The resulting dose-response curves for sEHI inhibitors had a nearly theoretical sigmodial shape as shown in Fig. 3 for three inhibitors of significantly different potency. Together with the precise online SPE-LC-ESI-MS-MS measurement (intra sample variation <5%) this yielded consistent data sets for each inhibitor, which could be accurately fitted to obtain IC 50 values (Fig. 3). The precision and reproducibility of the approach was demonstrated by repeated investigation of sEHI (1) (Fig. 2). The variation in the IC 50 values of three replicates on a single plate (intra-plate variation) was consistently low (15±6%, n=9). The intraday variation of three independent investigations on three different plates on a single day (n=3) was also good (9± 6%). The inter-day stability was calculated on the basis of the potencies determined on three different days using different batches of enzyme and substrate solution. The IC 50 values for sEHI (1) were 2.1±0.1, 1.9±0.1, and 3.1± 0.5 nmol L -1 . The resulting intraday variation of only 25% emphasizes the robustness of the approach, rendering it ideal for the determination of sEHI potency with high precision.
To demonstrate the performance of this approach a library of 13 competitive inhibitors, sEHI 11-23, was investigated for their potency, and the results obtained were compared with those from the commonly used fluorescence assay with CMNPC as surrogate substrate. As shown in Table 1 the IC 50 of the investigated sEHI varied over three orders of magnitude from 2 nmol L -1 to 1300 nmol L -1 based on the fluorescence assay. It was found that IC 50 values obtained with the LC-MS based assay using the natural substrate and the fluorescence assay correlated well, with a linear correlation coefficient between the methods of R 2 =0.9 (Fig. 4). Because the slope of the linear correlation (potency fluorescence vs. natural substrate assay) was 0.95 overall the same potencies were obtained by both methods. Moreover, both methods sort the potency of tested sEHI in the same rank order, with a Spearman's rank correlation coefficient of 0.9. When comparing individual sEHI, the difference between both methods was generally lower than a factor of 2 (Table 1, Fig. 4). The two exceptions to the excellent correlation of IC 50 values between the assays were sEHI (14) and sEHI (21). The IC 50 value obtained by the new natural substrate assay was eightfold higher IC 50 for sEHI (14) and a factor of 6 lower for sEHI (21) than those obtained from the fluorescence assay. However, these differences in IC 50 values also occur between sEH assays utilizing different surrogate substrates. Tsai et al. reported recently reported strong variances between the potencies observed by the fluorescence assay using CMNPC as substrate and the radiometric assay utilizing t-DPPO as substrate [32]. For example the IC 50 value for sEH 16 (Fig. 3) was twentyfold higher for the radiometric assay compared with the fluorescence assay. Interestingly, the potency of this sEH determined with the natural substrate assay was in general agreement with those from the fluorescence assay (Table 1, Fig. 4). All tested sEHi were urea derivatives, for which it is known that they only act as competitive inhibitors [29][30][31][32]. Thus it is unlikely that an allosteric or irreversible inhibitory mode of action of the inhibitors causes the differences in the measured potencies. On the basis of on these findings, it is concluded that the IC 50 value obtained for sEHI is not only affected by assay conditions, but is also dependent on the substrate used. The major mode of action of sEHI as a potential pharmaceutical is thought to be the stabilization of biologically active fatty acid epoxides, for example 14(15)-EpETrE [1,39]. Thus, the values obtained from the natural substrate assay described in here should be more predictive of the efficacy of the sEHI in vivo.
The sEH concentration used in this assay is only 33 pmol L -1 . This concentration is significantly lower than that used in other methods (1 nmol L -1 ) [3][4][5][6]. Given that the lowest IC 50 value which can be reliable determined with an enzyme activity assay is equivalent to the enzyme concentration, this method is capable of quantitatively determining the potency of compounds with thirtyfold higher potency than all other methods up to an IC 50 value as low as 0.03 nmol L -1 . For the first time, the potency of the most powerful sEHI can be quantitatively investigated. This unique feature of the new method was demonstrated by ranking the competitive sEHI (1-10) by their potency. All these inhibitors have potency≤1 nmol L -1 in the fluorescence assay (Fig. 4; Table 1). However, none of these compounds had a potency significantly higher than 1 nmol L -1 in the natural substrate assay. Moreover, many of the potent sEHI had IC 50 values of 5 nmol L -1 and higher, as predicted by the fluorescence method. With potency of approximately 0.7±0.3 nmol L -1 , sEHI (2) (Fig. 3) is the most potent of the inhibitors in the compound library investigated. This is comparable with the potencies of best sEHI inhibitors described so far [31] and no picomolar sEHIs were identified in this study. Nevertheless, the new approach described herein would enable characterization of picomolar IC 50 s, and thus might lead to the development of still more potent sEHI.
A new, LC-MS-based natural substrate approach for potency measurement of sEHI has been developed. In combination with a generic assay scheme in 96-well plate format, this new method enables reliable measurement of the potency of sEHI with high precision. By application of one of the fastest online SPE-LC-MS-MS systems described, just less than 22 min were needed for determination of the potency of a single compound. With an analysis time of 2.8 h per 96-well plate, this assay cannot compete with spectral or fluorescence-based highthroughput assays. However, investigation of a small sEHI library indicates that, despite good overall correlation with data from surrogate assays, the determined potency for individual sEHI is substrate-dependent. The mode of action of sEHI in vivo is thought to be the stabilization of fatty acid epoxides. Thus, the utilization of an assay using the endogenous substrates is indispensable for reliable structure-activity relationship (SAR) analysis of sEHIs. With the method described herein, a fast highly automated technique is now available, which enables further development of highly potent sEHI.
Acknowledgements This study was supported by the
We have developed a method for intact mass analysis of detergent-solubilized and purified integral membrane proteins using liquid chromatography-mass spectrometry (LC-MS) with methanol as the organic mobile phase. Membrane proteins and detergents are separated chromatographically during the isocratic stage of the gradient profile from a 150-mm C3 reversedphase column. The mass accuracy is comparable to standard methods employed for soluble proteins; the sensitivity is 10-fold lower, requiring 0.2-5 μg of protein. The method is also compatible with our standard LC-MS method used for intact mass analysis of soluble proteins and may therefore be applied on a multiuser instrument or in a high-throughput environment.
Integral membrane proteins represent approximately 25% of the human proteome and are of crucial biological importance in regulating the composition of the cell and the extracellular environment, membrane potential, metabolism, cell structure, and signaling pathways. As a result of their position as the gateway to cells, they are also the targets for many drugs, in fact more than 50% of commercially available small molecule drugs target membrane proteins such as G-protein-coupled receptors, ion channels, transporters, other receptors, and enzymes [1]. Despite the importance of membrane proteins in biology and pharmacology, to date only 246 crystal structures exist for membrane proteins [2] in comparison to over 50,000 for soluble proteins, and only 14 of those structures are of human integral membrane proteins. The difficulties in expression, extraction, and isolation of homogeneous membrane proteins of high quality and quantity for crystallization trials have been well documented [3,4]. Analysis of the proteins produced is also difficult. In particular, SDS-PAGE 5 provides very poor estimates of the molecular mass of membrane proteins.
Intact mass analysis by MS has proved to be an invaluable method for routine protein identification and quality assessment in the high-throughput production of soluble proteins for crystallography [5,6]. Published LC-MS techniques exist for intact mass analysis of integral membrane proteins using nonstandard mobile and stationary phases (for example, Refs. [7][8][9][10]). These methods are not easily adapted to routine mass spectrometric analyses and have not been widely adopted. We have developed a simple and robust LC-MS method for routinely determining accurate mass to within 1.0 Da of the calculated mass for integral membrane proteins, with a sensitivity ∼10-fold lower than our existing methods for soluble proteins. This fast and convenient method does not require pretreatment or protein precipitation for removal of detergents. Moreover, the method may be used interchangeably with soluble protein analysis on the same LC-MS instrument.
LC-MS grade reagents were purchased from Fluka (Sigma-Aldrich, UK). Detergents were purchased from Anatrace (Maumee, OH, USA). All other reagents were analytical grade purchased from Sigma-Aldrich.
Proteins from a variety of membrane protein families, including channels, transporters, and enzymes, were expressed in bacterial, yeast, or baculovirus-infected insect cell cultures using a variety of expression vectors and cell lines. All proteins examined were expressed as fusions to oligohistidine or FLAG tags. The proteins were extracted from whole cells or from membrane preparations using detergents, and were purified by immobilized metal affinity chromatography (IMAC) or M2 anti-FLAG agarose affinity purification as appropriate. The protein was further separated from contaminants using gel filtration chromatography, using a gel filtration buffer (typically 10-20 mM HEPES or Tris, pH 7.5, 150 mM NaCl or KCl, 0-10% glycerol, ±DTT) in the presence of appropriate detergent and in one instance (SCD; stearoyl-CoA desaturase) the addition of a known lipid mix. Where noted, the purification tag was cleaved using TEV protease. In some cases, the purified proteins were finally concentrated using either 50-or 100-kDa molecular weight cutoff Centricon (Millipore, Watford, UK) concentrator to between 0.1 and 6 mg/ml.
Reversed-phase chromatography was performed in-line prior to mass spectrometry using an Agilent 1100 HPLC system. Protein was diluted (between 5-fold and 20-fold in different experiments) with a solution of 1% formic acid to an injection volume of 10 μl and loaded on to a 4.6 mm internal diameter Zorbax 300SB-C3 column with column length of 50, 150, or 200 mm. Effective column length was extended by serial attachment of 50-and 150-mm columns. The solvent system used consisted of 0.1% formic acid in ultrahigh purity water (Millipore) (solvent A) and 0.1% formic acid in methanol (solvent B). The details of chromatography varied in different experiments as indicated in parentheses. Proteins were injected at initial conditions of 95% A and 5% B and a flow rate of 0.5 ml/min. After 1 min at 5% B, an initial linear gradient from 5% B to 95% B was applied, for 7 min (2-13 min). Elution then proceeded isocratically at 95% B for 10 min (3-15 min) before applying a second linear gradient from 95% B to 5% B over 2 min. Equilibration at 5% B for 2-4 min returned the system to the initial conditions (the times were varied in different experiments, as indicated in the figures). Chromatography was performed in a column oven set to 40 °C. The total amount of protein loaded varied from 0.2 to 6.0 μg.
Protein intact mass was determined using an MSD-ToF electrospray ionization orthogonal time-of-flight mass spectrometer (Agilent). The instrument was configured with the standard ESI source and operated in positive-ion mode. The ion source was operated with the capillary voltage at 4000 V, nebulizer pressure at 50 psig, drying gas at 350 °C, and drying gas flow rate at 10 L/min. The instrument ion optic voltages were as follows: fragmentor 250 V, skimmer 60 V, and octopole RF 250 V. These parameters are identical to our standard methods for intact mass measurement of soluble proteins.
Data analysis was performed using the MassHunter Qualitative Analysis Version B.01.03 Build 1.3.157.0 (Agilent) software. Individual LC scans were inspected and selected manually based on signal to noise and on the presence of a characteristic protein ionization series with at least five charge states. Varying numbers of scans were combined depending on the width of the chromatographic peak. Deconvolution was performed between 200 m/z and 3500 m/z, using peaks with a ratio of signal to noise greater than 30:1. The mass range for deconvolution was 10,000-100,000 Da and the step mass was 1 Da. Average mass was calculated at 90% of the peak height using a minimum of five consecutive charge states and a minimum "protein fit" score of 8. The technique used was to "walk" through the total ion chromatogram (TIC) one scan at a time, searching each individual m/z spectrum for a characteristic protein-like multiple ionization envelope. The initial walk through was used to determine the chromatographic characteristics of the detergent and to eliminate those regions where detergent is presumed to have caused detector saturation. Those spectra found to include a protein ionization envelope were combined and deconvoluted. Typically membrane proteins elute directly prior to or directly after the detergent, and are often coelute with some associated detergent. The ionization of the detergent species may then mask the protein ionization. When this was noted, axis zoom was used to clearly visualize the protein ionization spectrum, and only an area between the detergent species was used for deconvolution. Deconvolution of the m/z spectrum to neutral charge state was achieved using the set parameters described. The deconvoluted mass was compared with the theoretical mass derived from the DNA sequence of the expression construct for each protein. Where significant discrepancies between observed and theoretical mass occurred, these differences in mass were compared with expected mass changes for known modifications in the Unimod database (http://www.unimod.org).
The identity of proteins in solution or in gel bands excised following SDS-PAGE was confirmed by tryptic digestion and tandem MS (LC-MS/MS). SDS-PAGE gel bands were excised as 1 × 4-mm slices using a gel cutting tip (GeneCatcher, Web Scientific) and stored in 10% MeOH at 4 °C. Prior to digestion, the methanol solution was removed and replaced with 100% acetonitrile for 2-5 min. The solution was then removed and replaced with 100 μl of 100 mM NH 4 HCO 3 (pH 8.0). For digestion of proteins in solution, an aliquot (<30 μl) of the protein was added directly to 100 μl of 100 mM NH 4 HCO 3 . For phosphopeptide analysis, 30 μl of protein (0.5 mg/ml) in solution was diluted to 100 μl with 100 mM NH 4 HCO 3 (pH 8.0). In all cases, 1 μl of 1 M dithiothreitol was added and incubated at 56 °C for 40 min. Four microliters of 1 M iodoacetamide was then added and the reaction incubated at ambient temperature in the dark for 20 min. A further 1 μl of 1 M dithiothrietol, 200 μl of 100 mM NH 4 HCO 3 , and 1 μl of trypsin solution (sequencing grade, Sigma Cat. No. T 6567, 1 mg/ml in 0.01 M HCl) was then added. Tryptic digestion proceeded at 37 °C for 16 h and was terminated by the addition of 3 μl of formic acid. LC-MS/MS was performed using a Dionex U3000 nano HPLC coupled to a Bruker Esquire HCT ion trap mass spectrometer. The amount of 1-5 μl of tryptic digest was loaded on to a 200 μm i.d. × 5 cm PS-DVB monolith column (Pepswift, Dionex Corp.) A linear gradient of 0% B to 15% B was developed over 5 min, followed by a second linear gradient from 15% B to 40% B over 2 min. The column was washed at 90% for 2 min and then equilibrated at 0% B for a further 6 min. Solvent A was 2% (v/v) acetonitrile, 0.1% formic acid in water, solvent B was 80% acetonitrile, 0.1% formic acid. The flow rate was 2.5 μl/min. The mass spectrometer was operated in positive ion, standard enhanced mode with a scan rate of 8100 m/z/s and a scan range of 250-1800 m/z. The trap accumulation time was 200 ms and the accumulation target was 200,000 counts. Data-dependent peptide fragmentation was performed in Auto MSMS mode.
Phosphorylation mapping of protein KCNJ12 was performed as follows: phosphopeptide enrichment and tandem mass spectrometry were performed using the same instrumentation as for general peptide MS/MS. A total of 75 μl of KCNJ12 tryptic digest was loaded sequentially on to a 300 μm i.d. × 2-mm TiO 2 precolumn (5-μm particle size, Titanshphere, GL-Sciences, Japan) followed by a linear gradient of 0% B, 10% C to 90% B, 10% C over 14 min to elute nonphosphorylated peptides. Solvent A was 2% (v/v) acetonitrile, 0.1% formic acid in water, solvent B was 80% acetonitrile, 0.1% formic acid, and solvent C was 100% formic acid, with a flow rate of 3 μl/min. Phosphopeptides were eluted from the precolumn by injection of 5 μl of ammonium hydroxide solution (28% NH 3 ) from a sealed vial, and eluted isocratically in 100% solvent (A). Eluted phosphopeptides were passed directly to the electrospray source without further chromatography. Mass spectrometer parameters were as described above. Compound extraction and peptide deconvolution were performed using the DA data analysis program (Bruker Daltonik). Database searching was performed using the Mascot 2.2.04 search algorithm (Matrix Science) with the following search parameters: charge states +2, +3; MS tolerance 1.5 Da; MSMS tolerance 0.5 Da; UniProt/SwissProt database without taxonomic restrictions.
Initial attempts to obtain intact mass for membrane proteins by LC-MS using our standard protocols for soluble proteins were unsuccessful. We found, however, that by replacing the organic acetonitrile phase with methanol we could record an accurate intact mass for the potassium channel KCNJ12 [11] (Fig. 1). A complex total ion current pattern appeared late in the chromatogram; a careful analysis of individual spectra was used to resolve specific features within this region of the chromatogram. The first major component appeared to be TEV protease (chromatogram region 1 in Fig. 1A, elution time between 8.7 and 9.0 min); the presence of TEV protease was verified by deconvolution (data not shown). There followed a short peak of the detergent Cymal-6 (the major component of regions 2 and 3, elution time between 9.106 and 9.160 min), and then a dip in the TIC which appears to be detector saturation between 9.232 and 9.341. Between 10.409 and 10.698 min (region 4) another protein envelope was observed (Fig. 1B and C). The deconvoluted mass spectrum shows one peak at 38738.6 Da, which corresponds to the accurate mass of the KCNJ12 protein after TEV cleavage (Fig. 1D). However, the most prominent peak indicates a mass of 38817.4 Da. The mass difference of +78.8 Da suggests a phosphorylation. This was later confirmed by tryptic digest and phosphopeptide analysis, mapping the phosphorylation site to either Thr-353 or Ser-354 (data not shown). Phosphorylation of KCNJ12 at Thr-353 was suggested previously [12] based on site-directed mutagenesis.
These results indicated that accurate mass analysis of membrane proteins was feasible. The key seemed to be the use of methanol as the mobile phase, and the careful search of multiple individual spectra for a characteristic protein ionization envelope. However, the analyses were not easy to reproduce with KCNJ12 or with other proteins (the initial success may have been associated with some unusual properties of a particular protein batch, e.g., purity or detergent:protein ratio). Furthermore, the elution of protein peaks after the isocratic segment of the chromatogram indicated that there was scope for optimization of the chromatographic procedure.
Examination of the total ion chromatograms from the analyses of KCNJ12 (Fig. 1) and a second membrane protein HVCN1 (data not shown) clearly showed that elution of both proteins and detergent occurs toward the end or after the 3 min 95% methanol isocratic section. We attempted three modifications of the chromatographic protocol: shortening the initial gradient, extending the length of the isocratic section at 95% B, and using longer columns.
The elution protocol was modified by shortening the initial gradient to 2 min (1.2 CV) rather than 6 min and extending the 95% methanol isocratic section to 7 min (4.2 CV). This method was applied to KCNJ12 (Fig. 2A) and SCD (Fig. 2B). This new LC-MS method was shown to be more reliable and able to separate membrane proteins from detergent. This separation occurred entirely over the isocratic section. Elution of salts was observed during the initial gradient phase. Step elution in 95% methanol plus 0.1% formic acid was attempted but proved unsuccessful (data not shown) as the detergent and protein eluted over a narrow section of the TIC and detector saturation was still observed.
As another route to improve the chromatographic separation of membrane proteins and detergents we tested longer C-3 columns (150 and 200 mm versus 50 mm). The initial methanol gradient was developed over 7 min (1.4 CV and 1.05 CV for the 150-and 200-mm columns, respectively) and the 95% methanol isocratic section was extended to 10 min (2 CV and 1.5 CV, respectively). Overall, this improved the chromatographic separation as shown for SCD (Fig. 2C and D). Two other proteins, HVCN1 and ZMPSTE24, also showed improved separation on the longer columns (data not shown). The different species of detergent were also separable using the longer column and the extended isocratic stage (Table 1). Further extension to 200 mm also proved beneficial and increased the separation of the free detergent from the protein without detector saturation. Furthermore, we observed the separation of excess free lipids from protein isolated in the presence of detergent and lipid species (Fig. 4).
The details of the chromatograms vary between samples. Although we generally find that free detergents elute early, followed by proteins, and then other contaminants, the details of elution times, peak shape, and size vary between different samples. It is crucial, therefore, to carefully scan the m/z spectra throughout the chromatogram to identify the region containing data on protein m/z.
Using methanol gradient elution and one of the three chromatographic protocols outlined above we obtained LC-MS data and intact mass values for seven integral membrane proteins from five distinct families, purified in the presence of six different detergents. The choice of detergent was dependent on the protein purification, rather than as a parameter for LC-MS. The observed masses compared with theoretical mass for each protein, using data from 28 separate measurements, are shown in Table 2. In all cases where the mass difference (ΔM) was ⩾2 Da, a plausible interpretation could be provided: most often, the removal of the initiator methionine [13], with an occasional acetylation or carboxylation. These enzymatic posttranslational modifications are commonly observed for the expression systems used (Refs. [14][15][16] and unpublished observations). Taking these presumed posttranslational modifications into account, the mean mass accuracy for proteins ranging in size from 34 to 67 kDa was ±2 Da and comparable to that observed for soluble proteins using standard methods. Since the predicted mass and identity of each protein was known in advance, in all cases the mass accuracy was deemed sufficient to confirm the identity and integrity of the analyzed protein, as well as providing a tentative identification of modification states.
In addition to accurate mass spectra for integral membrane proteins, it was possible to obtain spectra for free detergent species in the solution and to identify detergent species which appear to be noncovalently associated with the membrane protein. Fig. 3 shows an example of protein-associated detergent molecules. The deconvoluted mass spectrum of the ATP binding cassette protein ABCB10 shows a mass of 67117.37 Da, which we interpret as the predicted mass after methionine removal and acetylation. Three additional peaks match the protein mass plus 1, 2, or 3 molecules of dodecyl maltoside, the detergent used in purification. In addition to protein-associated detergents, free detergents and lipids are clearly seen in m/z spectra: for example, the peaks of DDM in Fig. 3B, of Cymal-6 in Fig. 1B, and of lipids (DOPC, dioleoylphosphatidylcholine) in Fig. 4B. Note that the detergent peaks often dominate the m/z spectrum (as in Fig. 3B), making the protein spectrum hard to detect in a casual inspection. Table 1 lists the ions we observed in spectra of detergentcontaining samples. These data can be useful in monitoring the efficiency of detergent exchange; it is very important to verify that the LC column is completely washed of detergent residues, as these tend to persist and appear in subsequent samples.
The largest protein for which we obtained mass spectra was of 67 kDa; attempts to analyze significantly larger proteins (140-250 kDa) failed. This is commonly observed with soluble proteins, although the size limitation is unlikely to be absolute. Similarly, we have obtained experimental masses of proteins purified in some detergents (DM, DDM, Cymal-6, FC14, FC16, OG ± cholesterol hemisuccinate), but not with other detergents, including LDAO and C18E8.
The concentrations of the protein and the detergent are also important for success. Protein samples as low as 0.27 μg could be analyzed, although more typically 1-5 μg was required (protein concentrations above 0.2 mg/ml have been successful). Many detergents are concentrated along with the protein in standard centrifugal concentrators, even when 100-kDa cutoff membranes are used. The high detergent concentrations, often >0.5%, suppress the protein signal. One possible way to alleviate this problem is to purify the protein by gel filtration in a buffer containing no more than three times the critical micelle concentration of the protein, then take a sample of the highest concentration fraction from the gel filtration for LC-MS analysis, avoiding the accumulation of detergents in a spin concentrator.
Finally, the purity of the protein is an important factor; because of the complexity of the protein-detergent spectra it is difficult to deconvolute mixed or heterogeneously degraded proteins. However, in cases where the spectra are exceptionally clean, we have been able to resolve separate proteins in a mixed ion population (Fig. 5).
As an alternative method of protein identification we have applied tryptic digest followed by tandem MS on an ion trap. The procedures were the same as those used for soluble proteins. This analysis could be applied to purified proteins as well as isolated gel bands from impure preparations. A summary of MS/MS data obtained for the proteins described in this work is shown in Table 3. The peptide coverage is low, typically limited to the hydrophilic regions of an integral membrane protein. MS/MS results are often sufficient for unequivocal identification of the protein, and are particularly useful in identification of gel bands from intermediate stages of purification. However, peptide data very rarely confirm the integrity of the protein and its posttranslational modification status. Intact mass and tryptic digest MS/MS are best viewed as complementary analytical techniques.
Accurate determination of the intact mass of proteins is an important tool in protein analysis. Whether purifying naturally expressed or recombinant proteins, knowledge of the intact mass can establish the identity of the protein as well as its integrity, and provide data on protein modifications. To achieve this, the analysis should be reproducible, of high mass accuracy, and (ideally) using standardized methods that allow routine application in a busy MS facility. Such capabilities have long existed for soluble proteins, and have been used as a major tool for identification and quality assurance in a high-throughput protein production [5,6].
There is a general perception that membrane proteins are incompatible with intact mass analysis by standard LC-MS. We have attempted to apply our standard LC-MS methods to integral membrane proteins, using acetonitrile as the mobile phase, but failed to produce protein mass spectra. There are several processes which may be contributing to the difficulties in membrane protein MS:
Proteins may precipitate during the chromatography phase.
Proteins may bind tightly to the HPLC stationary phase and thus fail to elute.
(iii) Detergents may undergo preferential ionization and thereby suppress ionization of the associated membrane protein.
The proteins may remain associated with heterogeneous amounts of detergents.
Detergents are known to cause signal suppression and instrument contamination in LC-MS analysis and their use is often avoided. Previous reports have accomplished mass determination of membrane proteins by a number of approaches. Treatment of the samples prior to chromatography to remove detergents followed by direct infusion in formic acid or using a specialized LC setup enabled accurate mass analysis of bacteriorhodopsin and spinach thylakoid membrane proteins D1 and D2 [9,10,17]. Cadene and Chait [7] and Takayama et al. [8] demonstrated MALDI-MS analysis of integral membrane proteins with useful though rather lower mass accuracy. When the objective is to analyze multisubunit or protein-detergent complexes, direct infusion MS has been the method of choice [18][19][20][21]. In particular Barrera et al. [19] have shown that gas-phase ionization occurs for proteindetergent complexes in the native state, and that dissociation of these complexes can be initiated by manipulation of the collision cell voltages. The mass accuracy has been sufficient to determine the gas-phase interactions taking place, but unit mass accuracy, necessary for confirmation of sequence, has not been shown. The specially modified m/z range instruments used in this work appear to be limited to a resolution of only 3000 [22].
To our knowledge, neither LC-electrospray-ionization nor MALDI methods have been widely adopted. When faced with the challenge of integral membrane protein analysis most laboratories have chosen a "bottom-up" approach by LC-MSMS analysis of tryptic peptides (reviewed in Ref. [23]).
The initial observation of this study was prompted by a global shortage of acetonitrile, which led us to test the use of methanol as the organic phase for LC-MS of soluble proteins. Slightly longer retention times and peak broadening were observed for soluble proteins, but there was no loss in sensitivity or mass accuracy. In addition we found that there was considerably less carryover of proteins between sequential samples, which substantially improved our throughput for soluble proteins. We have therefore adopted methanol as the organic phase of choice for all our LC-MS analyses.
Application of methanol-based elution enabled us initially to measure the intact mass of the integral membrane protein KCNJ12. The difficulty of reproducing the initial result, and the appearance of the protein spectrum at the end of the programmed gradient indicated that the chromatographic method could be optimized. Because of practical considerations (the availability of limited quantities of a variety of proteins), we did not perform an exhaustive optimization for every protein, but the large number of analyses and the consistent results for a diverse set of membrane proteins indicate that the methods are robust and widely applicable. We can summarize the conclusions and resulting guidelines as follows:
(1) Chromatography using a short gradient from 5 to 95% methanol (in 0.1% formic acid), followed by a prolonged isocratic segment in 95% methanol/0.1% formic acid, provides effective separation of membrane proteins from detergents and other buffer components.
(2) Prolonging the isocratic segment, using longer columns, or a combination of both can improve the separation.
(3) To fully utilize the chromatographic separation, m/z spectra should be scanned individually to identify the segments with the best protein ionization signal. Even then, the protein spectrum may be dwarfed by the more abundant free detergent ions, so the spectra should be scrutinized carefully for weak signals.
Results of experiments not discussed in detail indicate that optimization of the chromatographic separation has a greater impact than varying the MS parameters on our instrument.
The errors in protein mass determination seem to be less than 2 Da, after taking into account plausible interpretations of larger mass differences. This accuracy allows us to make precise predictions as to the source of observed differences, which can then be tested by other means (for example, we have performed phosphopeptide mapping in one instance, but other methods such as N-terminal sequencing may be applied where appropriate).
The ratio of protein to detergent and the amount of protein are important determinants of success; samples for analysis should be taken from steps in the purification protocol that generate high protein-detergent ratios. In particular, concentrating protein-detergent mixes using centrifugal ultrafiltration often leads to excessive levels of detergent. The nature of the protein is also important: in addition to an apparent size limitation (>67 kDa), there are marked differences in the quality and signal to noise of spectra of different proteins (e.g., Figs. 1B, 3B and 4C). It is not easy at this stage to define the precise characteristics of protein samples that correlate with successful mass spectra; such generalizations may become possible with the accumulation of a larger data set.
The nature and homogeneity of the detergent are important. Maltosides and glucosides (such as DM, DDM), FC14, FC16, and Cymal-6 work well, whereas C12E8 C12E10, LDAO, Tween 20, and Triton X100 do not.
The ionized proteins are almost completely stripped of bound detergent, although a variable fraction of ion complexes may contain 1-5 bound detergent molecules.
Mass spectra from different regions of the chromatogram also provide valuable information on detergents and lipids, both free and protein-associated. Special care should be taken to avoid the effects of sample carryover, which are particularly pronounced with hydrophobic analytes.
In summary, we present a robust and widely applicable method for LC-MS analysis of the intact mass of integral membrane proteins. Unit mass accuracy allows identification of these proteins and their probable modifications. The method can be easily integrated in routine operation of an instrument used for both soluble and membrane proteins. Although we have not tested this, it is conceivable that the sensitivity may be improved by using narrow-bore or smaller columns. We expect that future extension of the method to analysis of more problematic proteins can be based on further optimization of the liquid chromatography step, including the use of alternative solvent systems or columns. Berridge et al. Page 11 Published as: Anal Biochem. 2011 March 15; 410(2): 272-280. Sponsored Document Sponsored Document Sponsored Document Berridge et al. Page 12 Published as: Anal Biochem. 2011 March 15; 410(2): 272-280. Berridge et al. Page 13 Published as: Anal Biochem. 2011 March 15; 410(2): 272-280. Sponsored Document Sponsored Document Sponsored Document Intact mass analysis of a mixed population of proteins. A mixture of the membrane protein KCNJ12, the soluble TEV protease used to cleave the fusion tag, and the membrane protein HVCN1, which occurred as a contaminant from an earlier analysis on the same LC column. Table 2
Summary of MS results for seven proteins analyzed. The results shown present the best data obtained for each protein.
Published as: Anal Biochem. 2011 March 15; 410(2): 272-280.
Berridge et al. Page 15 Published as: Anal Biochem. 2011 March 15; 410(2): 272-280. Published as: Anal Biochem. 2011 March 15; 410(2): 272-280.
Abbreviations used: SDS-PAGE, sodium dodecyl sulfate-polyacrylamide gel electrophoresis; LC, liquid chromatography; MS/MS, tandem mass spectrometry; TIC, total ion current; TEV, tobacco etch virus; CV, column volume. Abbreviated detergent names are listed in Table1; genes are identified by HUGO designations.
Published as: Anal Biochem. 2011 March 15; 410(2): 272-280.Sponsored DocumentSponsored DocumentSponsored Document
Published as: Anal Biochem. 2011 March 15; 410(2): 272-280.
We thank
Sponsored Document Sponsored Document
T he nanotechnology industry is growing at a rapid rate, with the increased design and development of novel engineered nanomaterials (NM), with diverse and wide ranging applications not only in industry, but also as consumer products and in the field of medicine. The growing production and utilization of NM has inevitably resulted in increased occupational, clinical, and consumer exposure to these substances and is likely to lead to an accumulation of NM in the environment. However, as of yet the effects these engineered substances have on human health and the environment especially, in the long term, still remain largely unknown.
Over the last 5À6 years there has been a steady increase in studies focusing on the toxic effects of NM, 1À5 but this does not reflect the exponential growth in the nanotechnology industry. Thus, the first report by the Royal Society and Royal Academy of Engineering Report in 2004, has been followed with several others including the European Scientific Committee on Emerging and Newly Identified Health Risks Report in 2006 and the DEFRA report in 2007, followed by another European Scientific Committee on Emerging and Newly Identified Health Risks Reports and a European commission joint research center institute for health and consumer protection report in 2009, 6À10 all of which continue to emphasize the need for further study into the safety of NM.
Traditional assays designed to quantify and characterize cellular damage, induced following exposure to exogenous agents, have been largely optimized for chemical compounds. However, given the unique physiochemical properties associated with NM, we cannot assume that they can be tested in the same way. For example, the possibility of direct interaction between NM and experimental assay components has the potential to result in false or misleading information, which could be a complicating factor in safety assessments. Some such instances have been documented in the literature, with reports demonstrating that single walled carbon nanotubes interact with a number of fluorometric and colorimetric dyes, to give unexpected results in cell viability assays. 11À13 Furthermore, boron nitride nanotubes have been shown to interfere with the MTT cell viability test. 14 It must also be stressed that due to the unique properties of each type of NM, we are currently unable to predict ABSTRACT: Due to the unique physicochemical properties of nanomaterials (NM) and their unknown reactivity, the possibility of NM altering the optical properties of fluorometric/ colorimetric probes that are used to measure their cyto-and genotoxicity may lead to inaccurate readings. This could have potential implications given that NM, such as ultrafine superparamagnetic iron oxide nanoparticles (USPION), are increasingly finding their use in nanomedicine and the absorbance/ fluorescence based assays are used to assess their toxicity. This study looks at the potential of dextran-coated USPION (dUSPION) (maghemite and magnetite) to alter the background signal of common probes used for evaluating cytotoxicity (MTS, CyQUANT, Calcein, and EthD-1) and oxidative stress (DCFH-DA and APF). In the present study, both forms of dUSPION caused an increase in MTS signal but a decrease in background signal from calcein and 3'-(p-aminophenyl) fluorescein (APF) and no effect on CyQUANT and EthD-1 fluorescence responses. Magnetite caused a decrease in fluorescence signal of DCFH, but it did not decrease fluorescence signal in the presence of the reactive oxygen species-inducer tert-butyl hydroperoxide (TBHP). In contrast, maghemite caused an increase in fluorescence, which was substantially reduced in the presence of the antioxidant N-acetyl cysteine. This study emphasizes the importance of considering and controlling for possible interactions between NM and fluorometric/ colorimetric dyes and, most importantly, the oxidation state of dUSPION that may confound their sensitivity and specificity.
behavior, thus tests on these substances must be done on an NM by NM basis.
NM are defined as substances with at least one dimension smaller than 100 nm, with different physiochemical properties compared to their micrometer sized counterparts due to their high surface area. In some cases the small dimensions make NM more chemically reactive, with particle size inversely proportional to bioactivity and toxicity. 15À19 High surface area can also change the strength and electrical conductivity of the material, while the quantum effects associated with NM result in unique optical, electrical, and magnetic behavior. For example, when smaller than 20À30 nm, iron oxide nanoparticles (NP) become superparamagnetic; a property that makes this particular material very useful in a number of biomedical applications including magnetic drug targeting, as a contrast agent to enhance MRI imaging and in magnetic tumor ablation through hyperthermia. 20 Given the potential clinical applications of ultrafine superparamagnetic iron oxide nanoparticles (USPION), evaluation of their safety is critical. Several studies already exist in the literature suggesting that iron oxide nanoparticles are toxic to cells and induce oxidative stress. For example, in 2003 Berry et al. showed that human dermal fibroblast cells treated with either dextran coated-or uncoated-magnetite NP exhibit cell death and reduced proliferation. 21 Another study reports that exposing iron oxide NP to human microvascular endothelial cells induces reactive oxygen species (ROS) production, that leads to the remodelling of microtubules and subsequently to increased cell permeability. 22 In this study we investigated whether common assays used for the measurement of oxidative stress, cell viability, and cell growth are compatible with dextran coated ultrafine superparamagnetic iron oxide NP (dUSPION), to measure these parameters. We examined the interactions of dUSPION with cell viability assays: 3-(4,5-dimethylthiazole-2-yl)-5-(3-carboxymethoxyphenyl)-2-(4-sulfophenyl)-2H-tetrazolium (MTS), calcein, CyQUANT, and ethidium homodimer (EthD-1) in a cell free system. We also performed similar studies for oxidative stress assays: 2 0 ,7 0dichlorofluorescein-diacetate (DCFH-DA) and 3 0 -(p-aminophenyl) fluorescein (APF). Furthermore, using the antioxidant N-acetyl cysteine (NAC), we examined the potential of dUSPION to initiate ROS production in a cell-free system, and using tert-butyl hydroperoxide (TBHP) as a source of ROS, we examined the potential antioxidant properties of magnetite.
Materials. dUSPION were purchased from Liquids research, Bangor, UK. RPMI 1640, horse serum, and Hanks balanced salt solution (with NaHCO 3 , without phenol red, calcium chloride, and magnesium sulfate) were purchased from Gibco, UK. Sodium hydroxide, glucose, sodium phosphate monobasic, and sodium phosphate dibasic were purchased from Fisher Scientific, UK. N-Acetyl-L-cysteine (NAC) and dimethyl sulphoxide (DMSO) were purchased from Sigma-Aldrich, UK. Tissue culture black microplates (96 well) were purchased from Greiner Bio-one, UK, and clear 96-well tissue culture microplates were purchased from Nunc, UK. CellTiter 96 Aqueous One Solution reagent was from Promega UK, Southampton, UK. CyQUANT probe, live/Dead Viability/cytotoxicity Kit, DCFH-DA, and APF were purchased from Invitrogen molecular probes, Paisley, UK.
dUSPION were supplied in suspension in water at a concentration of 10 mg/mL. dUSPION was diluted to the appropriate concentrations in distilled water and vortexed for 10 s immediately before use.
Dynamc Light Scattering (DLS). The hydrodynamic particle size of dUSPION samples were obtained by DLS. The measurements were performed using a Malvern 4700 spectrometer (Malvern instruments Ltd., UK) either in RPMI with 1% horse serum or in Hepes-buffered (20 mM) Hanks balanced salt solution with glucose (5 mM) (pH 7.4). Data is presented as the average values of 15 readings.
The samples were dried on an Indium substrate and examined with a PHI Quantera SXM(TM) (Ulvac-phi, Inc., Japan). All data points were acquired using a beam spot size of 200 um, 40 W, and 15 kV, under a pressure of 5 Â 10 À9 Torr. The electron source was Al monochromatic with a tilt angle of 45°, operating at 26 eV. Maghemite was acquired from 700 to 720 eV, using 70 sweeps with a bandpass energy of 26 eV. Oxygen was acquired from 525 to 537 eV, using 35 sweeps with a bandpass energy of 26 eV. Carbon was acquired from 278 to 293 eV, using 25 sweeps and a bandpass energy of 26 eV. Survey scans were completed from 0 to 1100 eV using 3 sweeps with a bandpass energy of 140 eV. Each sample was acquired on its own to prevent the possible contamination from previous samples.
The z-potential values of the dUSPION were determined by Zetasizer 2000 (Malvern instruments Ltd., UK). The nanoparticles were prepared in water, and the z-potential values are presented as the average readings of 10 experiments.
dUSPION samples for TEM were prepared by dispersion in methanol, then drop-casting on holey carbon TEM support films (Cu-grids) and air-dried. TEM was performed using a Philips/FEI CM200 field emission gun TEM fitted with an Oxford Instruments ultrathin window EDX detector and ISIS software plus a Gatan Imaging Filter (GIF200) with Digitialmicrograph software. The microscope was operated at 197 keV.
The first step for all assays was loading of dUSPION concentration range onto 96-well plates. After addition of the probes specified below, the fluorescence or absorbance was measured on a POLARStar Omega plate reader (BMG Labtech, Aylesbury, UK). For all assays each dose of dUSPION was performed in triplicate within the plate, and each plate was performed in triplicate on three different days, thus accounting for both intra-and interplate variability, respectively.
The MTS assay it is based on the reduction of the tetrazolium compound MTS and an electron coupling reagent (phenazine ethosulfate; PES) into a soluble formazan product. This conversion takes place only in the presence of metabolically active cells, utilizing the mitochondrial dehydrogenase enzyme. The formazan product can be measured by absorbance at 490 nm, which is directly proportional to the number of live cells in culture and can thus be used for determining the number of viable cells in proliferation or cytoxicity assays. A 20 μl portion of CellTiter 96 Aqueous One Solution reagent was added to each well of a 96-well plate already loaded with 100 μl of dUSPION at different concentrations; then plates were incubated in a humidified incubator at 37 °C for 1 h, and the absorbance was measured at 490 nm.
A CyQUANT probe was prepared per the manufacturer's instructions for use, 200 μl of this was added to the wells of a 96-well plate already loaded with dUSPION, and fluorescence was measured after 5 min (fluorescence excitation and emission at 480 and 520 nm, respectively).
Live/Dead (calcein/EthD-1) Assay. The Live/Dead Viability/ cytotoxicity Kit was used with final concentrations of 10 μM Calcein or 20uM EthD-1 added to the appropriate volumes of dUSPION in a 96-well plate. The fluorescence was read after 1 h. For calcein fluorescence, the excitation and emission wavelengths utilized were 485 and 530 nm respectively, while for EthD-1, the wavelengths were 530 and 645 nm. The principle of using calcein is that the cell's ubiquitous esterase activity converts the virtually nonfluorescent, cell-permeant calcein AM, to the highly fluorescent calcein, which is retained within the cell. On the other hand, EthD-1 is used to detect dead cells as it cannot enter through the intact plasma membrane of live cells; it can, however, easily enter damaged cells. Upon binding to cellular nucleic acids, EthD-1 increases in fluorescence intensity 40-fold producing a bright red fluorescence detected at 635 nm.
DCFH-DA assay is based on the principle that upon internalization, the diacetate (DA) portion of the hydrophobic dye is cleaved by intracellular esterases. The resulting DCFH is nonfluorescent until it is oxidized by ROS to its highly fluorescent product DCF. Initiation of the DCFH-DA assay requires this DA portion of the molecule to be cleaved, and in acellular systems, this cleavage can be achieved via chemical means using sodium hydroxide (as in the present study) or using media. The excitation of the DCF molecule at 485 nm emits green fluorescence at levels proportional to the amount of ROS present, which can be detected at 520 nm. DCFH-DA was dissolved in DMSO to a concentration of 1 M and was further diluted to the appropriate concentration with Hepes-buffered (20 mM) Hanks balanced salt solution with glucose (5 mM) (pH 7.4). Before using DCFH-DA in a cell free system, chemical cleavage of the diacetate (DA) portion was necessary by incubation with 10 mM NaOH in the dark, at room temperature, for 30 min. The resulting DCFH was neutralized with 25 mM phosphate buffer (1:1 sodium phosphate monobasic: sodium phosphate dibasic) (pH 7.4) and the solution was kept in the dark, on ice until use. A 2 μM portion of DCFH was then added to the wells of a 96-well plate previously loaded with dUSPION, and fluorescence was measured over 1 h with fluorescence excitation and emission at 480 and 520 nm, respectively.
To determine if the increases in fluorescence signal observed with maghemite was indeed due to the generation of ROS induced by dUSPION (as opposed to dUSPION interaction with assay components exclusively), further experiments were performed in the presence of 2 mM NAC, which was applied to the plates with dUSPION, prior to DCFH. The concentration of NAC used (2 mM) was chosen, as preliminary studies showed this concentration to be sufficiently potent to reduce dUSPION induced increases in DCFH signal. Also, to determine whether the decrease in fluorescence observed with magnetite was due to interactions with DCFH or whether magnetite actually has antioxidant properties, experiments were done to see whether the presence of magnetite can prevent or reverse TBHP-induced oxidative stress. For this, TBHP (25 mM) was applied to the plates with dUSPION and compared to wells loaded with TBHP alone, to look for a reduction in fluorescence.
For all DCFH experiments, because of the dynamics of the DCFH fluorescence over time, time zero readings were subtracted from time 60 min readings.
APF was added to a dUSPION preloaded 96-well plate giving a final concentration of 8 μM, and fluorescence was measured over 1 h (fluorescence excitation and emission at 480 and 520 nm, respectively).
A one-way ANOVA with a two-sided Dunnett's post hoc test was performed for each data point (n = 3) comparing each one to its relevant untreated control. Fisher's exact test was used to compare each dose of dUSPION treatment with NAC to its relevant zero NAC control. For all graphs, data is presented as percentage of control without dUSPION inclusion.
Several recent studies have investigated the potential of engineered NM to induce cytotoxicity, oxidative stress, and genotoxicity. However, results from these studies are not always consistent, and consequently, a great degree of uncertainty regarding the true toxicity of NM still exists. 1,2,4,23À28 One potential explanation for this uncertainty is the lack of standardized protocols and a deficiency in appropriate controls when using certain assays for studying the toxic effects of NM. 23,28À30 We have previously shown that standard DNA damage assays also need modification when dealing with NM as opposed to chemicals for which they were originally optimized. 28 When considering the best approach for characterization of NM, it must be recognized that due to their unique physicochemical properties it cannot be assumed that NM can be tested in the same way as chemicals and that there may be some confounding factors skewing the results, which may result in misinterpretation of data sets. In fact, previous studies have shown that carbaceous NM and nanotubes interact with a range of colorimetric and fluorometric probes used for testing cytotoxicity and oxidative stress, including MTT, neutral red, IL-8 cytoset ELISA, almar blue, WST-1, and Coomasie blue assays, and is thought to be due to the adsorbing properties of NM resulting in false readings. 11À13,31À34 Interactions between NM and assay components are particularly problematic when these test systems are central to assessing NM safety. Thus, where colorimetric and fluorometric dyes are to be relied on for experimental test systems, potential alteration of background signal due to interference imparted by the NM must be considered, the importance of which is demonstrated in the present study using dUSPION.
The physicochemical features of dUSPION were assessed under experimental conditions (Table 1), and as shown in Figure 1, both maghemite and magnetite dUSPION were spherical with a core diameter of ∼10 nm. However, the latter dUSPION exhibited a slightly more pronounced degree of agglomeration.
When using the MTS assay in a cell-free system, both dextrancoated maghemite and magnetite showed no significant change in absorbance levels, between concentrations of 1 Â 10 À3 and 10 μg/mL as illustrated in Figure 2a. However, 100 μg/mL of both maghemite and magnetite dUSPION samples caused a significantly dramatic (9.5-and 6.5-fold, respectively) increase in background absorbance in an acellular system (p < 0.05) (Figure 2a). For investigating whether the optical properties of dUSPION have a direct effect on absorbance readings, the absorbance of dUSPION alone at the wavelength required for the MTS assay was investigated. Interestingly, the results showed that 100 μg/mL of dUSPION are capable of significantly increasing the absorbance readings compared to the control level in the absence of the MTS reagent (p < 0.05) (Figure 2b).
The MTS assay is a simple and sensitive colorimetric method that has been used in the past to quantitate NM induced cytotoxicity, including zinc oxide, titanium dioxide, and silica-and alkoxy silane coated iron oxide NP. 35,36 However, in the current study it is clearly demonstrated that dUSPION induce a substantial increase in absorbance at the wavelengths required for the MTS assay, thereby severely confounding the sensitivity and specificity of the assay for quantifying cell viability in response to dUSPION exposure. This could potentially lead to misinterpretations of biological response. This suggests that MTS can be used to evaluate viability of cells treated with dUSPION only if it is taken into consideration that higher concentrations of dUSPION might affect background signal. The present study is not alone in demonstrating NM-induced tetrazolium-based assay interference. Studies have also shown that carbon nanotubes can adsorb another common tetrazolium compound, MTT, used for cytotoxicity studies, onto their surface leading to a quenching and an alteration in absorbance. 11,12 However, the effect in the case of dUSPION depends on the type of probe and the oxidation state of the dUSPION used.
Incubation of dUSPION with the fluorescent probe calcein (also frequently used to quantify cell viability), resulted in a dosedependent decrease in the intensity of the resultant fluorescent signal at 520 nm with 1 Â 10 À2 μg/mL maghemite, reaching significance at 10 μg/mL (p < 0.05; Figure 3). A similar profile was also observed with magnetite, but only the highest concentration (100 μg/mL) significantly reduced the calcein fluorescent signal in a cell free system (p < 0.05). The reduction in fluorescence intensity observed at the higher dUSPION concentrations suggests quenching is induced by the NP that is independent of their oxidative status.
The precise mechanism involved in the assay interferences observed is not well understood. It is evident that dUSPION alters the optical properties of the assay probes. It is known that carbon NM such as single walled carbon nanotubes (SWCNT) can adsorb dyes onto their surface, likely through van der Waals forces which subsequently quench or alter their absorbance or fluorescent properties. 11,34 It is not known if dUSPION adsorb colorimetric and fluorometric dyes in the same manner as SWNCT. According to Worle-Knirsch et al. in 2006 and later verified by Casey et al. in 2007, SWNCT interact with insoluble MTT-formazan crystals that are formed after MTT reduction by cellular enzymes. 11,34 However the present study was in a cell free system, thus dUSPION are unlikely to interact with MTS in exactly the same way as SWNCT interact with MTT. However, it is possible that dUSPION somehow adsorb calcein and MTS onto their surface leading to quenching of fluorescence in the case of calcein and enhancement of absorbance readings in the case of MTS. Interestingly, a dUSPION solution alone examined without the MTS assay components significantly increased absorbance readings at 490 nm to levels similar to that seen with the MTS assay components. This suggests that the increased MTS response is mostly due to the contribution of dUSPION's own optical properties (Figure 2b). Potentially, factors such as surface chemistry, fabrication process, or types of surfactants used to disperse the NM (in this study, dextran) may also play a role in governing the interactions and degree of interference with colorimetric and fluorometric dyes. However, in the present study, dextran alone without the MTS assay components did not increase absorbance levels at 490 nm (data not shown).
An alternative cell viability test system, the CyQUANT assay is a sensitive technique used for the determination of cell numbers in culture, and unlike MTT or MTS, this assay does not depend on cellular metabolic activity. This assay is based on the CyQUANT GR dye fluorescing only when bound to cellular nucleic acids (DNA and/or RNA) in lysed cells. The current study shows that in a cellfree system increasing concentrations of maghemite or magnetite did not interfere with the assay (up to 100 μg/mL; Figure 4). Thus, this assay could be used as an alternative to the MTT or MTS assays. Similarly, neither form of dUSPION used in this study interfered with EthD-1 fluorescence in a cell free system (Figure 5).
Interestingly, as opposed to the other probes tested in this study, the function of which is reliant on chemical reactions, both CyQUANT and EthD-1 work through enhancing fluorescence intensity upon binding to nucleic acids. While dUSPION appear to present limited interference when the assay is dependent on a physical change, such as binding to nucleic acids, it appears that if the assay relies on a chemical reaction, the dUSPION may be interacting with the assay components directly or may be interfering at some step in the chemical reaction.
Fluorescence-based dyes are also key reporters for oxidative stress, for example the fluorometric probe DCFH-DA is widely used for the detection of intracellular ROS. 37,38 As illustrated in Figure 6a, there was a dose-dependent decrease in fluorescence intensity of DCFH with increasing concentrations of magnetite, which was significant at concentrations of 1, 10, and 100 μg/mL (p < 0.05; Figure 6a). TBHP and increasing concentrations of magnetite demonstrated a synergistic effect with a dose dependent increase in fluorescence, reaching significance at 100 μg/mL of magnetite (p < 0.05; Figure 6a). Our results suggest that this decrease in DCF signal is not due to any antioxidant effect that magnetite may have, since magnetite did not reduce TBHPinduced increases in DCF fluorescence. Thus, it is likely that in a cell free system magnetite quenches the fluorescence response at higher doses, possibly through adsorption of the probe onto their surface. Somehow, in the presence of a ROS-inducer such as TBHP, this effect is reversed. The present results again demonstrate the unpredictability of responses when using such assays to measure NM safety and the importance of performing preliminary tests to check for assayÀNM interactions.
Interestingly, maghemite presented the opposite response, causing an increase in fluorescence intensity from 1 μg/mL. At subsequent concentrations, the increase in fluorescence signal was dramatic, equating to a 60-fold elevation at 10 μg/mL and a 40fold increase in background fluorescence when 100 μg/mL maghemite was used (Figure 6b). To determine whether this increase in signal was indeed caused by oxidative stress, the experiment was repeated in the presence of the antioxidant NAC. The maghemite-induced increase in fluorescence signal at 1, 10, and 100 μg/mL was indeed found to be significantly reduced (p < 0.05). Thus, demonstrating that maghemite induces oxidative stress in the acellular system (as opposed to interacting with the dye itself) which can be substantially reduced using NAC (Figure 6b). This suggests that maghemite has a much higher oxidative potential than magnetite. It is possible that in the same manner as magnetite, maghemite is able to quench fluorescence response at low concentrations. However, due to its higher oxidative potential, at higher concentrations the massive increase in oxidative species production masks any fluorescence quenching effect that the maghemite may have, resulting in an overall increase in fluorescence response There are several examples in the literature of the fluorescent probe DCFH-DA being used for the quantification of oxidative stress induced by NM, as this is one of the primary mechanisms associated with adverse cellular responses to NM. Some examples include iron oxide NP exposed to mesenchymal stem cells and HeLa (human cervival carcinoma) cells, SWCNT-induced ROS in HaCaT (human keratinocyte) cells, ambient ultrafine particles, cationic polystyrene nanospheres, TiO 2 , fullerol NP and carbon black in RAW 264.7 phagocytic cells, and silver nanoparticles in human hepatoma and skin keratinocytes. 4,39À42 It is not clear from these studies whether the possibility of confounding factors, which we have found to be associated with NM, have been taken into consideration, and whether the appropriate controls have been included.
The distinct difference in the oxidative potential of maghemite and magnetite may be related to the oxidative state of the iron ions in the complex. In maghemite (Fe 2 O 3 ) iron ions are mostly Fe 3þ , while in magnetite (Fe 3 O 4 ) they are a mixture of Fe 3þ and Fe 2þ with a Fe 2þ /Fe 3þ ratio of 0.435 43 (Table 1). It is possible that Fe 2þ surface ions undergo Fenton reaction by reacting with any H 2 O 2 that may be available within the aqueous environment of the assay, to produce a hydroxyl radical. H 2 O 2 can also react in a Fenton-like reaction with Fe 3þ to generate [Fe III OOH] 2þ which can go on to generate OOH 3 , OH 3 , or OH À , it has also been suggested that reaction of Fe 3þ with H 2 O 2 generates superoxide. It is known that Fe 2þ is more reactive than Fe 3þ ; however, in the present study it seems that the Fenton-like reaction involving Fe 3þ is more potent. One possible explanation for this is the size of dUSPION agglomerates. Particle sizing using DLS suggests that in the buffer used for the DCFH-DA assay (Hepesbuffered (20 mM) Hanks balanced salt solution with glucose (5 mM)), magnetite forms bigger agglomerates than maghemite, thus magnetite would have less exposed surface area and potentially less ions available to react (Table 1). It is also possible that maghemite is a more stable molecule than magnetite and consequently does not release as many iron ions to react with the assay components. Fenton and Fenton-like reactions are known to generate different ROS and intermediate species (see below). 44 Thus, another possible explanation for the differences observed when using the DCFH-DA assay is that DCFH-DA may not be as sensitive in detecting ROS produced by magnetite as compared to ROS produced by maghemite.
Fe
Similar to DCFH-DA, APF is a fluorometric probe used for the detection of oxidative species. APF is oxidized by free radicals to yield a highly fluorescent product which can be detected at 520 nm. Concentrations of maghemite and magnetite between 1 Â 10 À3 and 1 μg/mL did not have a significant effect on the resultant intensity of the APF fluorescence signal. However, 10 and 100 μg/mL maghemite and 100 μg/mL magnetite caused a significant decrease in background APF fluorescence signal (Figure 7), which could be due to adsorption onto the surface of the NPs, thus quenching the fluorescence response. This is in contrast to the results seen when using DCFH with maghemite, which reported increased oxidative stress at 1À100 μg/mL (Figure 6). This may be because APF is oxidized by fewer free radical species than DCFH as it is much more selective, only detecting the hydroxyl radical and peroxynitrite anion. Thus, APF may not be detecting the specific ROS produced by maghemite, which are possibly ROS products of Fenton-like reactions. The contrasting results obtained from using these two different oxidative stress assays again highlights the difficulties and complexity of testing the safety of NM.
Although this study sheds light on NP-USPION interactions in an acellular environment, one important point is that the undesirable interactions demonstrated in the present study may or may not be mimicked exactly in a cellular milieu. It is quite likely that these interactions will still occur in the culture media when the cells are exposed to NP and the test reagent as demonstrated by Zhang and colleagues, who have shown that the brown color of the USPION led to higher cell viability readings. 45 Additionally factors in a cellular system, such as media components or cell debris, may modulate the interactions observed in the acellular system. Alternatively, intracellular masking of NP with proteins and other metabolites following cellular uptake could result in a dampened interaction.
In conclusion, not all standard biological assays are compatible with dUSPION and this holds great significance because these NM are routinely used in various biomedical applications subsequent to toxicity testing that utilizes colorimetric/fluorometric probes such as those used in the current study. The present study shows that colorimetric assays such as MTS, and fluorometric probes including calcein, DCFH-DA, and APF can interact with dUSPION especially at higher doses. Therefore, control experiments are essential to establish this threshold for interaction prior to the use of such probes for assessment of cell viability and oxidative stress responses following exposure. Where such interference is detected, alternative test systems should be used and may even require those that do not rely on quantitating colorimetric or fluorometric changes. A number of studies can be found in the literature which have used colorimetric and fluorometric assays, but do not indicate whether the possibility of test system/NM interactions have been controlled for. Thus, in some of these cases, it is possible that results may be misleading due to interactions between the NM and the selected test system, thereby confounding interpretation. Test system/NM interactions may represent a source for some of the conflicting observations in the current literature, in reports assessing apparently the same material but with different experimental systems. Additionally, it is important to note that even subtle differences in oxidation state of the metal oxide NP is sufficient to result in major differences in their ability to interfere with fluorometric dyes. Thus, until we develop a more comprehensive understanding of the parameters that influence such interactions, test system validation assessments are necessary on a NM-by-NM basis.
dx.doi.org/10.1021/ac200103x |Anal. Chem. 2011, 83, 3778-3785
S.M.G. and N.S. are joint first authors for this work. This work is supported by funds from the
The effect of varying short-chain alkyl substitution of the indole nitrogens on the spectroscopic properties of cyanine dyes was examined. Molar absorptivities and fluorescence quantum yields were determined for a set of pentamethine dyes and a set of heptamethine dyes for which the substitution of the indole nitrogen was varied. For both sets of dyes, increasing alkyl chain length resulted in no significant change in quantum yield or molar absorptivity. These results may be useful in designing new cyanine dyes for analytical applications and predicting their spectroscopic properties.
Cyanine dyes are a class of conjugated, fluorescent molecules with polymethine chromophores composed of an odd number of carbon atoms. These dyes exhibit unusually long-wavelength absorbance and fluorescence relative to the size of their chromophores, typically absorbing light in the visible to near infrared (NIR) region. 1 These compounds were originally utilized as sensitizing additives to photographic emulsions, but their unique structural and photophysical characteristics have since proven useful for a wide variety of other applications requiring photosensitive materials, such as optical recording media and solar cells. [2][3][4] Additionally, cyanine dyes can be used as fluorescent labels of both proteins and DNA, thereby greatly enhancing the sensitivity of fluorescence detection for these types of biomolecules. [5][6][7] NIR-absorbing cyanine dyes are particularly wellsuited for use as fluorescent labels of proteins and nucleic acids, as there is no interfering autofluorescence from biomolecules at these long wavelengths, and have been applied to both in vitro analytical studies and in vivo biomedical imaging. [5][6][7][8][9] Due to the tremendous utility and versatility of this class of compounds, significant research efforts are being directed at developing new cyanine dyes functionalized for specific applications and optimizing the properties of these dyes.
In the process of developing new dyes, it is important to determine how varying the heteroaromatic ring nitrogen substituents influences spectroscopic behavior. Cyanine dye structures are commonly modified at these positions to enhance binding interactions, make the dyes pH sensitive, or improve their solubility in various solvents. Understanding how these modifications may influence the absorption characteristics and quantum efficiency is important for designing new compounds with specific functional and spectroscopic characteristics. Cyanine dyes in the excited singlet state can decay back to the ground state through four major pathways: fluorescence, intersystem crossing, internal conversion, and photoisomerization. Of the radiationless decay processes, it has been suggested that photoisomerization is the most significant, followed by internal conversion. 10,11 The extent of photoisomerization has been shown to be dependent on dye rigidity, 12 which may be influenced by both backbone structure and side chain substitution. By introducing rigidifying structures or rotationhindering bulky substituents, photoisomerization would be expected to decrease with a corresponding increase in quantum yield. 13,14 The purpose of this study was to investigate the effect of varying alkyl group length substitution of the indole ring nitrogen on the molar absorptivities and quantum yields of cyanine dyes. Molar absorptivities (ε) of each dye were calculated as per the Beer-Lambert law, and fluorescence quantum yields (φ) were determined by a relative method. Two classes of cyanine dyes were studied: pentamethine cyanine dyes and ring-stabilized heptamethine cyanine dyes. These two different dye "backbones", which differ in terms of the substitution and length of the polymethine chain as well as the heterocyclic moieties at either end of the polymethine chain, were chosen as models to ensure that the observed results are widely applicable to a range of cyanine dyes, rather than peculiar to one subgroup of dyes. The pentamethine cyanine dyes studied all had a 4,5:4′,5′-Dibenzo-3,3,3′,3′tetramethylindadicarbocyanine backbone, the structure of which is shown in Figure 1.
The heptamethine cyanine dyes studied all had a 2-[2-[2-Chloro-3-[(1,3-dihydro-3,3-dimethyl-1-propyl-2H-indol-2-ylidene)ethylidene]-1-cyclohexen-1-yl]ethenyl]-3,3-dimethyllindolium backbone, the structure of which is shown in Figure 2.
Absorbance spectra were measured using a Perkin-Elmer Lambda 20 UV-Visible Spectrophotometer (Perkin-Elmer Incorporated, Waltham, MA) interfaced to a PC, with a spectral bandwidth of 2 nm. Fluorescence spectra for the pentamethine cyanine dyes were obtained using a Shimadzu RF-1501 Spectrofluorophotometer (Shimadzu Scientific Instruments, Columbia, MD) interfaced to a PC, with the spectral bandwidths for both excitation and emission set to 10 nm and the sensitivity set to "high". Fluorescence spectra for the heptamethine cyanine dyes were obtained using a ISS K2 Multifrequency Phase Fluorometer (ISS Inc., Champaign, IL) interfaced to a PC, with the spectral bandwidths set to 10 nm. The excitation source used for the ISS K2 Fluorometer was an external 690 nm class IIIB laser (100 mW, S/N 901290, Lasermax Inc., Rochester, NY). Disposable absorbance cuvettes and quartz fluorescence cuvettes with pathlengths of 1.00 cm were used for absorbance and fluorescence measurements, respectively. All calculations were carried out using Microsoft Excel (Microsoft Corporation, Redmond, WA).
Pentamethine dyes 666 ($98.0%) and 829 ($99.5%) were obtained from Organica Feinchemie GmbH (Wolfen, Germany). IR-676 iodide (97%) was obtained from Spectrum Info Limited (Kiev, Ukraine). Rhodamine 800 chloride (R800) (Fluorescence Reference Standard, Sigma-Aldrich, St. Louis, MO) was also obtained for use as a reference standard in the determination of the quantum yield of the pentamethine dyes. Heptamethine dyes IR-780 iodide (99%) and IR-786 perchlorate (98%) were obtained from Aldrich Chemical Co. (Milwaukee, WI) and Sigma-Aldrich, respectively. Indocyanine green (ICG) (lot GG01, 82.0% purity, TCI America, Portland, OR) was obtained for use as a standard in the determination of the quantum yield of the heptamethine dyes. The purchased dyes were used without further purification.
Pentamethine dye MHI-85 and heptamethine dye MHI-71 were synthesized in our lab following near infrared dye syntheses described in the literature. 15
MHI-85 was synthesized as illustrated in Equation 1. The pentacarbocyanine dye MHI-85 was obtained by the condensation reaction between benz[e]indolium salt and malonaldehyde bis(phenylimine) monohydrochloride under basic conditions.
The purified product consisted of dark purple-blue crystals, mp 244-246 °C, yield 85%; 1 HNMR (300 MHz, DMSO-d6): δ = 8.46 (t, J = 12.0 Hz, 2H), N MHI-85 N N O 3 S CH 2 (CH=NPh)2 • HCl Ac 2 O -∆ O 3 S SO 3 Na -(Eq. 1)
NaOAc 8.25 (d, J = 8.4 Hz, 2H), 8.08 (d, J = 3.6 Hz, 2H), 8.05 (d, J = 3.6 Hz, 2H), 7.78 (d, J = 8.4 Hz, 2H), 7.68 (t, J = 7.5 Hz, 2H), 7.51 (t, J = 7.5 Hz, 2H), 6.67 (t, J = 12.0 Hz, 1H), 6.43 (d, J = 13.8 Hz, 2H), 4.25 (s, 4H), 2.60-2.50 (m, 4H), 1.97 (s, 12H), 1.85-1.75 (m, 8H). MS (ESI+): calcd. for C 41 H 45 N 2 S 2 O 6 + [M-Na] + 725.2719; found 725.2688. MHI-71 was synthesized as illustrated in Equation 2. The heptacarbocyanine dye MHI-71 containing cyclohexene in the middle was synthesized by condensing the salt of Fischer base with Vilsmeier-Haack reagent under basic conditions to produce the dye.
The purified product consisted of iridescent golden-green crystals, mp 219-221 °C, yield 80%; 1 HNMR (300 MHz, CDCl 3 ): δ = 8.37 (d, J = 13.0 Hz, 2H), 7.45-7.37 (m, 4H), 7.30-7.20 (m, 4H), 6.24 (d, J = 13.0, 2H), 4.22 (t, J = 6.0 Hz, 4H), 2.75 (t, J = 6.0 Hz, 4H), 2.08-1.98 (m, 2H), 1.90-1.80 (m, 4H), 1.74 (s, 12H), 1.56-1.46 (m, 4H), 1.04 (t, J = 6 Hz, 6H). MS (ESI+): calcd. for C 38 H 48 N 2 Cl + [M-I] + 567.3521; found 567.3506.
Stock solutions of the dyes were prepared by weighing the solid on a 5-digit analytical balance directly into a brown glass vial and adding methanol (MeOH) (LC-MS Chromasolv Grade, Sigma-Aldrich, St. Louis, MO) via a class A volumetric pipette (Kimble/Kontes, Vineland, NJ.). The contents of the vial were vortexed for 20 seconds, then sonicated for 5 minutes to ensure complete dissolution. The stock solutions were protected from light and stored in the freezer when not in use.
Stock solutions were used to prepare five to six samples in methanol with concentrations ranging from 0.25-10 µM. Samples were prepared in 5.00 (±0.02) and 10.00 (±0.02) mL volumetric flasks using a 5-50 µL Micropipette 821 and a 200-1000 µL Pipetman (P1000) micropipette (Gilson, Inc., Middleton, WI.). The absorbance spectrum of each sample was measured using the Perkin Elmer Lambda 20 Spectrophotometer, and the absorbance at the wavelength of maximum absorbance (λ MAX AB ) was determined. The absorbance values (A) of each sample at λ MAX AB were plotted as a function of dye concentration (C), and the linear regression equation was computed.
Standards were chosen with wavelengths of maximum emission within 10 nm of those of the unknowns to prevent errors resulting from wavelength dependent variation in fluorimeter response. Samples of the dyes and their respective standards were prepared from stock solutions such that their absorbance at λ MAX AB was less than 0.1 (to prevent the inner filter effect in fluorescence measurements). The absorbance and fluorescence spectra of each sample were obtained concurrently to minimize experimental error from photobleaching and potential solubility issues, and for all scans, the standard was run both prior to and following the unknowns (to ensure no change in instrumental response over the course of the runs). For both the pentamethine and heptamethine dyes, duplicate absorbance scans were obtained and the absorbance values at both the λ MAX AB and λ EXC were averaged. The emission spectra of the pentamethine dyes were measured in triplicate using the RF-1501 fluorimeter with the excitation wavelength set to 620 nm. The emission spectra of the heptamethine dyes were measured in triplicate using the ISS-K2 fluorimeter with a 690 nm excitation wavelength. For both sets of dyes, the area under each fluorescence curve was calculated and corrected for the Rayleigh peak area (if necessary). The average fluorescence peak areas were then calculated for each sample.
N I Cl NHPh PhHN Cl N Cl N I MHI-71 NaOAc EtOH ∆ (Eq. 2)
Molar absorptivities of each dye (in methanol) were computed from the slope of the linear regression plots of absorbance versus concentration. Absorbance values greater than or equal to 2.0 were excluded from these data sets. The molar absorptivities (ε) were then calculated at λ MAX AB from the least squares slopes of the respective data sets, as per Beer's law.
Provided in Table 1 is a summary of the found λ MAX AB values and average calculated molar absorptivities (ε) for the pentamethine cyanine dyes and R800 as a reference sample. Also included in Table 1 are the standard deviations and percent relative standard deviations of the calculated molar absorptivities. The similarities in the λ MAX AB values of the cyanine dyes indicate the similarity in substitution. For dyes 666, 829, and IR-676, the substituents are all electrondonating alkyl groups. MHI-85 exhibits virtually no shift in absorbance maximum relative to those found for the alkyl-substituted dyes 666 and 829, indicating that the sulfonate moiety is far enough removed from the chromophore that its electron-withdrawing effects have insignificant influence. The molar absorptivities of the dyes at λ MAX AB do not vary greatly amongst themselves and do not follow any apparent trend based on substitution. Butyl-substituted dye 666 exhibited the greatest molar absorptivity, followed by methyl-substituted IR-676, followed by isopentylsubstituted dye 829. Butylsulfonato-substituted MHI-85 had the lowest observed molar absorptivity. The differences between dyes 666, 829, and IR-676 are not statistically significant, as the molar absorptivities of these dyes fall within the outer limits of each others ranges of standard deviation. However, the dif- ferences in molar absorptivity between MHI-85 and both dyes 666 and IR-676 are statistically significant, if relatively small. The lower molar absorptivity of butylsulfonato-substituted MHI-85 relative to the other dyes may be due to substituent chain length, effects of the sulfonate moiety, or the presence of minor impurities. The molar absorptivity values showed good precision, with reasonably low percent relative standard deviations (1.7%-7.4%) for all of the dyes. Provided in Table 2 is a summary of the found λ MAX AB values and average calculated ε values at the λ MAX AB for the heptamethine cyanine dyes and ICG as a reference sample. Also included in Table 2 are the standard deviations and percent relative standard deviations of the calculated molar absorptivities. The similarities in the λ MAX AB values amongst the cyanine dyes indicate the similarity in substitution; all are substituted with electron-donating alkyl groups. As with the pentamethine dyes, the molar absorptivities of the heptamethine dyes at λ MAX AB do not vary greatly amongst themselves and do not follow any apparent trend based on substitution. Propyl-substituted IR-780 exhibited the greatest molar absorptivity, followed by methyl-substituted IR-786, followed by butyl-substituted MHI-71. Only the differences in molar absorptivity between MHI-71 and IR-786 and between MHI-71 and IR-780 are statistically significant (the ranges of molar absorptivity specified by the standard deviations are mutually exclusive). The molar absorptivity values showed good precision, with reasonably low percent relative standard deviations (0.88%-8.45%) for all of the dyes.
Provided in Figures 3 and 4 are representative comparisons of the absorbance and emission spectra of the pentamethine and heptamethine dyes, published quantum yield of the standard (φ S ), as per the following equation. In this equation, the indices S and U refer to the standards and the unknowns, respectively.
Methanol was used as the solvent for both the standards and the unknowns; therefore no correction for solvent refractive index was necessary in this equation. This calculation assumes negligible variation in instrument response within the range of emission wavelengths exhibited by the unknowns with respect to the standard emission wavelengths.
Provided in Table 5 are the average quantum yields of the pentamethine dyes calculated relative to R800 as determined from multiple studies, along with their standard deviations and percent relative standard deviations. Reproducibility of results following duplicate determinations of quantum yield was good for dyes MHI-85 and IR-676, with percent relative standard deviations of 5.6 and 5.1 percent, respectively; accordingly no further determinations were made. However, reproducibility was poor for dyes 666 and 829, accordingly the quantum yield of dye 666 was determined two additional times, and the quantum yield of dye 829 respectively. For both sets of dyes, the fluorescence and absorbance spectra are reasonably good mirror images of one another (with the exception of the Soret peak visible at lower wavelengths in the absorbance spectrum), and the Stokes' shifts provided in Tables 3 and 4 are relatively small, ranging from 20-23 nm (321-348 cm -1 ). This indicates minor structural changes between the ground and excited singlet states of these dyes.
The fluorescence quantum yields of each of the cyanine dyes were calculated relative to the standard from their respective average fluorescence peak areas (F), average absorbances at λ EXC (A), and the was determined one additional time in an attempt to improve reproducibility. Following these further studies, a significant improvement in reproducibility was observed for dye 666 but not dye 829. Nonetheless, even taking into account the high percent relative standard deviations, the average quantum yields of the pentamethine dyes did not vary significantly with increasing alkyl N-substitution. Additionally, it appears that the addition of solubility-enhancing anionic sulfonate groups to these alkyl N-substituents has no significant effect on the quantum yield.
Provided in Table 6 are the average quantum yields of the heptamethine dyes calculated relative to ICG as determined from duplicate studies, along with their standard deviations and percent relative standard deviations. Reproducibility of results following duplicate determinations of quantum yield was good for all of the dyes, with IR-786, IR-780, and MHI-71 having percent relative standard deviations of 2.8, 4.5, and 6.8 percent, respectively; accordingly no further determinations of quantum yield were made. Overall, the quantum yields of the heptamethine dyes were much lower than those of the pentamethine dyes, as expected, 16 and as with the pentamethine dyes, the average quantum yields of the heptamethine dyes did not vary significantly with increasing alkyl N-substitution.
Based on these results, it can be generalized that increasing the chain length of short-chain alkyl substituents on the heterocyclic indole nitrogens has little to no effect on the quantum yields and molar absorptivities of cyanine dyes, and the addition of sulfonate groups to the ends of these alkyl N-substituents also does not change the quantum yield. The lack of effect on quantum yield may be attributable to the fact that the N-substituents are not directly conjugated to the chromophore and therefore have little to no effect on internal conversion-type energy loss, and that additionally, these short chain substituents do not provide adequate steric hindrance to interfere sufficiently with photoisomerization to cis-cyanine. The implications of these results are significant to the design and synthesis of new cyanine dyes. As previously mentioned, cyanine dyes are frequently modified at these heterocyclic nitrogen positions to enhance binding interactions, make the dyes pH sensitive, or improve solubility. Accordingly, provided the functional group is not directly attached to the chromophore, but instead "bridged" by a short alkyl substituent, introducing this group should have little effect on the quantum yield of the dye relative to a structurally similar dye lacking this functionality. For example, if one wished to modify a dye with desirable spectroscopic properties to be more water soluble or to bind more strongly with DNA or protein by adding a charged or polar functionality, this modification could be carried out without concern for loss or change of the desired spectral characteristics. Additionally, these results suggest that quantum yields of new cyanine dyes can be roughly predicted prior to their synthesis and characterization based on quantum yields of structurally similar preexisting dyes, provided that such data exists. publish with Libertas Academica and every scientist working in your field can read your article "I would like to say that this is the most author-friendly editing process I have experienced in over 150 publications. Thank you most sincerely." "The communication between your staff and me has been terrific. Whenever progress is made with the manuscript, I receive notice. Quite honestly, I've never had such complete communication with a journal." "LA is different, and hopefully represents a kind of scientific publication machinery that removes the hurdles from free flow of scientific thought." Your paper will be: • Available to your entire community free of charge • Fairly and quickly peer reviewed • Yours! You retain copyright http://www.la-press.com
Analytical Chemistry Insights 2011:6
We would like to thank
This manuscript has been read and approved by all authors. This paper is unique and is not under consideration by any other publication and has not been pub-
lished elsewhere. The authors and peer reviewers of this paper report no conflicts of interest. The authors confirm that they have permission to reproduce any copyrighted material.
A novel approach to prepare homogeneous PbS nanoparticles by phase-transfer method was developed. The preparatory conditions were studied in detail, and the nanoparticles were characterized by transmission electron microscopy (TEM) and UV-vis spectroscopy. Then a novel lead ion-selective electrode of polyvinyl chloride (PVC) membrane based on these lead sulfide nanoparticles was prepared, and the optimum ratio of components in the membrane was determined. The results indicated that the sensor exhibited a wide concentration range of 1.0 Â 10 À5 to 1.0 Â 10 À2 molÁL À1 . The response time of the electrode was about 10 s, and the optimal pH in which the electrode could be used was from 3.0 to 7.0. Selectivity coefficients indicated that the electrode was selective to the primary ion over the interfering ion. The electrode can be used for at least 3 months without any divergence in potential. It was successfully applied to directly determine lead ions in solution and used as an indicator electrode in potentiometric titration of lead ions with EDTA.
Lead is ubiquitous in the environment. In recent years, because of the increasing use of lead and its serious hazardous effect to human health, much effort has been placed on the development of ion-selective electrodes (ISEs) for detecting of lead ions. Most of ISEs are polyvinyl chloride (PVC) membrane electrodes. Their dynamic response is generated by dispersing different active material as an ion carrier in a PVC matrix.
A number of diverse ligands, viz., N 0 -dibenzyl-1,4,10,13-tetraoxa-7,16-diazacyclooctadecane (Gupta, Jain, and Kumar 2006), N,N 0 -bis(salicylidene)-2,6-pyridinediamine (Jeong et al. 2005), meso-tetrakis-(2hydroxy-1-naphthyl) porphyrin (Ardakany 2004), capric acid (Mousavi, Barzegar and Sahari 2004), benzyl disulphide (Abbaspour and Tavakol 1999), 5,5 0 -dithiobis-(2-nitrobenzoic acid) (Rouhollahi, Ganjali, and Shamsipur 1998), tetrabenzyl pyrophosphate (Xu and Katsu 2000), 9,10-anthraquinone derivative (Tavakkoli 1998), and 3,4,4a,5-tetrahydro-3-methylpyrimido-[1,6-a] benzimidazole-1(2H) thione (Jain, Sondhi, and Rajvanshi 2003) were used to prepare Pb 2þ sensors. Besides these, crown ethers and calixarenes were also widely investigated as sensing materials, including benzo-15-crown-5 (Sheen and Shih 1992), diazacrown ether (Yang et al. 1997), 18-crown-6 (Zareh, Ghoneom, and El-Aziz 2001), N,N 0 -dimethylcyanodiaza-18-crown-6 (Ganjali 2002), sym-dibenzo-16-crown-5 ethers (Su, Chang, and Liu 2001), 4 0 -vinylbenzo-15-crown-5 homopolymer (Ganjali et al. 1998), lariat crown ethers (Hasse 2001), 1,10-dibenzyl-1,10-diaza-18-crown-6 (Mousavi 2000), thiacrown ether (Shamsipur, Ganjali, and Rouhollahi 2001), calix[n] arene phosphine oxide derivatives (Cadogan 1999), di-and tetrathioamide calyx[4]arene (Malinowska 1994), thiophosphorylated calyx[6]arene (Wroblewski 1996), 4-tert-butylcalix[6]arene (Bhat, Ijeri, and Srivastava 2004), calixarene carboxyphenyl azo derivative (Lu, Chen, and He 2002), 2,12-dimethyl-7,17-diphenyltetrapyrazole and 5,11-dibromo-25,27-dipropoxy-calix[4]arene (Jain et al. 2006), and 4-tert-butylcalix[4]arene (Gupta, Mangla, and Agarwal 2002). In addition, nanoparticles used as ionophores have ever been reported, such as nanosized PbO powders (Li et al. 2005). However, a lead ion-selective electrode based on homogeneously nanosized PbS particles has never been reported.
In this work, a novel approach to prepare homogeneous PbS nanoparticles by a phase-transfer method is first described, followed by characterization of the nanoparticles TEM and UV-vis spectra. Performances of lead ion-selective PVC electrode based on PbS nanoparticles prepared were also studied. The results, reported in the present communication, showed that this sensor based on PbS nanoparticles exhibited good sensitivity and selectivity toward lead ions and could therefore be used as a selective sensor for its quantification.
High-molecular-weight polyvinylchloride (PVC) and dioctylphthalate (DOP) were obtained from Aldrich. Analytical reagent-grade tetrahydrofuran (THF), nitric acid, and sodium hydroxide were obtained from Shanghai Chemical Reagent Corporation. Solutions of metal (nitrates) were prepared in doubly distilled water and standardized by the reported methods whereever necessary. Working solutions of different concentrations were prepared by diluting 0.1 molÁL À1 stock solutions.
The microstructure and morphology of the nanoparticles were characterized by TEM (model 800, Hitachi) and UV-vis spectrophotometry (model UV-1601PC, Shimadzu), respectively. The potential measurements were carried out with a digital pH meter (model pHS-3D).
A 60-mL portion of 1 Â 10 À4 molÁL À1 lead acetate was placed in a beaker and adjusted to pH 6.0 with 1 molÁL À1 sodium hydroxide under continuous stirring. The solution was then transferred to a 150-mL separatory funnel. Then 60 mL of 1.5 Â 10 -4 molÁL À1 H 2 Dz=CCl 4 solution was added, and the mixed solution was oscillated thoroughly for 20 min. The color of the organic phase changed from deep green to pink, proving that lead ions were transferred from water phase to the organic phase by extraction. The two phases were stratified thoroughly after 2 h, and the transparent organic phase was transferred to an Erlenmeyer flask. Then 60 mL of 1 Â 10 -4 molÁL À1 CH 3 CSNH 2 =CH 3 CH 2 OH solution was added slowly in a dropwise fashion into the flask under vigorous stirring, while the dropping speed was controlled to about 10 drops per minute. The mixed solution was stirring for 12 h at room temperature, and then the PbS nanoparticles formed in the solution.
An amount of 0.1000 g of PVC powder was put into a 100-mL beaker and dissolved in 10 mL of THF. Then 0.1000 g of DOP was added and mixed homogeneously. Then the solution was put into a Petri dish (90 mm in diameter and 18 mm in height). Fifty mL of solution containing PbS nanoparticles was also added. The resulting solution was left overnight at room temperature for evaporation of the organic solvent. A transparent membrane of about 0.2 mm in thickness was obtained.
A cell assembly of the following type was used: Ag=AgCl j 0:01mol Á L À1 Pb 2þ ; pH 6 PVC membrane k test solution; pH 6 j Ag=AgCl reference electrode
The electrodes were immersed directly in the test solution at 25 AE 0.2 C. The electric potential of lead acetate standard solutions (with concentrations from 1 Â 10 -7 to 1 Â 10 À1 molÁL À1 ) were measured by a pHS-3D potentiometer.
All pH adjustments were made with diluted solutions of HNO 3 or NaOH.
TEM has been used rather extensively to measure the nanoparticle size in a direct and visual manner. Figure 2 shows the TEM micrographs of PbS nanoparticles. The majority of the particles' sizes fall into the range of 40 to 60 nm in diameter. It is noted that most of the particles have compact spherical structures with proportional size distribution. Although some portion aggregates during the sample preparation, the overall distribution is still very well. The nanoparticles can be stored from 3 to 5 days at room temperature without any obvious aggregation. Nanosized particles generally exhibit threshold energy in the optical absorption measurements because of the size-specific band gap structures (Wroblewski et al. 1996;Bhat, Ijeri, and Srivastava 2004;Lu, Chen, and He 2002;Jain et al. 2006), which is reflected by the blue shifting of the absorption edge (from near-infrared to visible) with decreasing particle size (Chen, Templeton, and Murray 2000;Steigerwald and Brus 1990;Weller 1993;Zhang 1997;Borrelli and Smith 1994;Fendler 1987. From the optical spectra of PbS nanoparticles (Fig. 3), it is seen that the H 2 D z =CCl 4 solution and Pb 2þ ÀH 2 D z =CCl 4 solution showed rather large absorption band at 629 nm and 517 nm, respectively. Otherwise, the PbS nanoparticles showed a broad but not strong peak at 437 nm and obvious blue shifting, which confirmed the production and remarkable quantum effect of PbS nanoparticles.
The concentration of lead ions was a key factor to the prepared PbSnanoparticles. If the solution contained a high concentration of Pb 2þ
2848 W. Song et al.
ions, the reaction speed with sulfide ions was too fast to distribute in time, resulting in that PbS nanoparticles that easily aggregated. In contrast, if the concentration of Pb 2þ ions solution was low, the resulting PbS particles would be small. The response of nano-PbS-PVC membrane electrodes had poor performances. Therefore, 1 Â 10 À4 molÁL À1 lead ion was adopted.
The influence of the acidity on concentration of Pb 2þ after 20 min of extraction was investigated. The UV-vis spectrum of Pb 2þ ÀH 2 D z =CCl 4 solution was shown in Fig. 4a. The results showed that the H 2 D z could extract Pb 2þ from the water phase to the CCl 4 organic phase in pH 4 to 10. The maximum intensity of absorption at 517 nm is shown in Fig. 4b. The optimal pH value was 6. The solution of PbS nanoparticles prepared at this pH had good stability and could be stored for a long time.
The influence of the ratio between H 2 D z and CCl 4 was also investigated, as illustrated in Fig. 5. The optimal ratio was 1:1.5. Considering the solubility of H 2 D z in CCl 4 and the concentration of Pb 2þ ions, 1.5 Â 10 -4 molÁL À1 H 2 D z =CCl 4 solution was used.
Three methods were used to obtain sulfide ions: Na 2 S=H 2 O solution, CH 3 CSNH 2 =H 2 O solution, and CH 3 CSNH 2 =CH 3 CH 2 OH solution. When the Pb 2þ -H 2 D z =CCl 4 solution was mixed with Na 2 S=H 2 O or CH 3 CSNH 2 =H 2 O solution, two phases were formed. The reaction only took place at the interface. The sulfide ions in water were so bulky that plenty of large PbS particles were produced. Otherwise, the CH 3 CSNH 2 =CH 3 CH 2 OH solution could be mixed thoroughly with Pb(&)ÀH 2 Dz=CCl 4 solution.The reaction, therefore, occurred in the homogeneous phase. Furthermore, the speed of sulfide ions released from thioacetamide in anhydrous ethanol was not fast, and thus homogeneous PbS nanoparticles could be obtained. The TEM images showed that these nanoparticles had small diameters and well-proportioned size distribution. The optimal concentration of thioacetamide in anhydrous ethanol was 1.0 Â 10 À4 molÁL À1 from the experimental results.
Membrane composition was investigated to evaluate the performance of the lead ion-selective electrode based on nano-PbS-PVC membrane. The DOP was a plasticizer that could increase the membrane flexibility. 2850 W. Song et al.
Without DOP, the membrane would turn hard and fragile. As a result, the electrode had no response signal and lead ion. With too much DOP, the membrane would be viscid and its mechanical intensity would decrease. When the mass of PVC powder to DOP was 0.1000 g respectively, a 0.2-mm-thick membrane was formed after THF evaporated. The membrane had appropriate thickness, good flexibility, and high mechanical intensity. It was found that the thickness of 0.2 mm was optimal through the experiments. The performance of the ion-selective electrode developed was excellent.
In accord with the responses of the generally adopted ion sensor, the internal solution could affect the sensor response when the membrane internal diffusion potential was appreciable. Thus, the effect of activity of the internal solution on the functioning of the membrane sensors was studied by measuring the potentials at varying activity of internal solution, viz. 1.0 Â 10 À2 , 5.0 Â 10 À2 , and 1.0 Â 10 À1 molÁL À1 Pb 2þ (Fig. 6). Best results in terms of slope and working concentration range were obtained with internal solution of activity 1.0 Â 10 À2 molÁL À1 . Thus, the activity of the internal solution was kept at 1.0 Â 10 À2 molÁL À1 in all studies.
As an active component in the membrane, nano-PbS directly affected the electrode performance in a dose-dependent manner. Various volumes of 2852 W. Song et al. PbS nanoparticles, viz. 10, 20, 30, 40, 50, and 60 mL of solution containing PbS nanoparticles, were added to prepare six membranes. The performances of those electrodes were tested in a series of lead ion standard solutions from 1.0 Â 10 À7 to 1.0 Â 10 À1 molÁL À1 . It was seen from Fig. 7 that the electrode made of membrane containing 30 or 40 mL of PbS nanoparticle solution had better performance.
With 40 mL of PbS nanoparticles as ionophores and DOP as a plasticizer, the response of Pb 2þ selective electrode affected by pH was studied (Fig. 8). The pH value of 10 À3 molÁL À1 Pb(NO 3 ) 2 test solutions was adjusted with HNO 3 and NaOH. As illustrated in Fig. 8, the potentials remained constant when the pH value was kept from 3.0 to 7.0. Outside this range, the electrode response changed drastically. This was probably due to lead hydroxide formation and response of the electrode to hydrogen ions at low pH values.
The response time of the electrodes was found to be approximately 10 s when the electrodes were directly immersed in 10 À3 or 10 À4 molÁL À1 solution; the response times of the membrane electrodes were 6 and 8 s, respectively. Even if the concentration of Pb 2þ ions was 10 À5 molÁL À1 , the response time would also be stable within 15 s. The lifetime of the electrodes was determined by recording its potential at an optimum pH value and plotting its calibration curve each day. The parameters, such as the slope, working range, and response time of the electrode, were found to be reproducible. Also, the membrane could be used over a period of 3 months without observing any significant drift in the parameters. After this period, a slight deviation was observed in response time and slope, which could be corrected by re-equilibrating the membrane with 1.0 Â 10 À7 molÁL À1 lead ion solution for 2-3 days. With this treatment, the assembly could be used over a period of about 1 more month and then it was replaced by a fresh membrane.
The modified fixed interference method as suggested by Viteri and Diamond (1994) was used with 1.0 Â 10 À2 molÁL À1 interfering ions to determine selectivity coefficients of the proposed sensor. Selectivity parameter data for various ions are presented in Fig. 9 and Table 1. A value of selectivity coefficient less than 1 indicates that the electrode was selective to the primary ion over the interfering ion. However, it was important to mention that the smaller value of selectivity coefficient was, the higher selectivity of the electrode. In this respect, the selectivity coefficients for K þ and Na þ are not very small. Although the sensor is selective even over these two ions, the order of selectivity is not very high. It suggested that lower concentrations of K þ and Na þ would not cause interference, but higher levels would cause interference.
The analytical application of the electrode was investigated as an indicator electrode in the potentiometric estimation of Pb 2þ solution by titrating 25 mL of 1.0 Â 10 À4 molÁL À1 Pb(NO 3 ) 2 against 1.0 Â 10 À3 molÁL À1 EDTA solution. The pH of the solution was maintained at 6.0 throughout the titration with HNO 3 and NaOH. The titration plot does not have a standard sigmoid shape (Fig. 10), which may be due to some interference caused by Na þ ions of the disodium EDTA salt (Gupta and Kumar 1999;Gupta et al. 1999). However, the sharp breakpoint corresponded to the stoichiometry of Pb 2þ -EDTA complex show the efficacy of the proposed electrode in the potentiometric estimation of Pb(II).
In this work, homogeneous lead sulfide nanoparticles were successfully prepared by the phase-transfer method. TEM and UV-vis spectroscopy showed that PbS nanoparticles had good spherical structure with proportional size distribution, which could be used as ion carriers in the development of lead ion-selective electrode. The sensor prepared by PbS nanoparticles exhibited good reproducibility and fast response time and could be employed for more than 3 months. It could also be used as an indicator electrode in the potentiometric titration of lead ions with EDTA.
Otherwise, the results of potentiometric selectivity indicated that most of metal ions would not affect the selectivity of the lead electrode seriously. Therefore, the proposed sensor is a good addition to the existing list of the lead ion-selective sensors reported and can be used for real sample analysis.
The infrared (IR) receptors in the pit organ of crotaline snakes are very sensitive to temperature. The sensitivity to IR radiation is much greater in crotaline snakes than in boid snakes because they have a thermosensitive membrane suspended in a pair of pits that comprise the pit organ. The vasculature of the pit membrane, which is located near IR-sensitive terminal nerve masses, the IR receptors, supplies the blood necessary to provide cooling and the energy and oxygen that the IR receptors require. The ophthalmic and maxillary branches of the trigeminal nerve innervate the pit membrane. In crotaline snakes, the trigeminal ganglion (TG) is divided into the ophthalmic and maxillomandibular ganglia; a prominent septum further separates the two divisions of the maxillomandibular ganglion. The TG neurons in the ophthalmic ganglion and the maxillary division of the maxillomandibular ganglion relay IR sensation to the brain. This article reviews the IR-sensitive pit organ and trigeminal sensory system structures in crotaline snakes.
Animals can detect temperature changes in their environment with special thermoreceptors, which aid in hunting, feeding, and survival in many organisms, including crotaline and boid snakes [1], vampire bats [2,3], fire-seeking beetles [4,5], certain butterflies [6,7], and blood-sucking bugs [5,8]. Crotaline and boid snakes have specialized infrared (IR) receptors in pit organs that enable them to detect, locate, and apprehend prey when combined with other sensory systems [1]. IR sensitivity in crotaline and boid snakes has likely evolved from the somatic sensory system, which evolved to sense IR radiation, similar to the vision sensation mechanism [9]. IR radiation sensitivity is much greater in crotaline and chamber opens widely to the exterior, whereas the inner cavity interacts with the external air via a sphincter-controlled pore rostral to the eye [11]. Fig. 1B shows a diagram of a crotaline pit organ in cross-section. In comparison, boid snakes have receptors in their labial scales, either without specialized structures or in the fundus of specialized labial pits (Fig. 1C) [1,10,12].
IR radiation sensitivity is much greater in crotaline snakes than in boid snakes [10,12,13], and IR sensation in crotaline snakes requires lower temperature changes and energy thresholds of 0.003°C and 10.75 μW/cm 2 , respectively, and has a greater detection distance (66.3 cm) than that in boid snakes [10,13], due to receptor depth and pit organ anatomy [9]. In the crotaline snake Trimeresurus flavoviridis, the pit organ membrane is approximately 15 μm thick [11] and is surrounded by air on both sides, preventing heat loss by conduction to surrounding tissues and allowing most of the IR radiation incident on the pit to heat the membrane [11]. Approximately 7,000 IR-sensitive trigeminal sensory axon endings are distributed throughout the membrane, which excite their nerve fibers when warmed [9]. Similar heatsensitive nerve endings cover the bottom of each pit in boid snakes [9], because their receptors are not suspended, as in crotaline snakes. Unlike the receptors in the floor of the labial pits or scales, the pit membrane of crotaline pit organs is highly innervated and has a dense capillary bed that projects throughout the network of mitochondria-rich terminal nerve masses (TNMs), supplying blood for cooling and providing energy and oxygen to the TNMs (Fig. 2A) [9,14]. The pit membrane vasculature is finer, flatter, and more convoluted than that of other sensory organs, including the skin and retina [14]. Scanning electron microscopy has revealed the three-dimensional morphology of the fundamental structures that comprise the IR receptor system in crotaline snake pit membranes (Fig. 2). The TNMs are arranged in one layer directly beneath the outer epithelium of the pit membrane (Fig. 2A, B). A group of myelinated nerve fibers branch to innervate the TNMs at a point farthest from the outer epithelium, below the TNMs and capillary network (Fig. 2A, C). The intimate apposition of the blood capillaries with the TNMs and their feeder nerve fibers is shown in Fig. 2D. Viewed from the inner chamber, the myelinated fibers are beneath the TNMs and become unmyelinated as they bend upward to form the array of sensors (Fig. 2E).
In crotaline snakes, the pit organ has unusual vasculature and numerous IR-sensitive TNMs [11,14,15]. Due to the extensive nature of the vasculature and its proximity to the TNMs [14,15], chemicals administered via the blood reach the TNMs. This morphology has been examined by studying the IR receptor response to vasoactive chemicals, which occurs either by a vasoactive effect on the pit membrane vasculature or by a chemical effect on the IR receptor thermoreceptor channels, such as the transient receptor potential (TRP) vanilloid family [16][17][18]. Gracheva et al. [19] and Panzano et al. [20] found that pit organs respond to temperature using the heat-activated cation channel TRP ankyrin 1 (TRPA1). They suggested using TRPA1 and other TRP channels as new genetic and physiological markers to delineate evolutionary relationships between vertebrates and invertebrates. Further studies are required to confirm if IR receptors have other TRP channels.
In mammals, the trigeminal nerve connects sensations including touch, pressure, pain, and temperature from the facial, maxillary, and mandibular regions to the brain. Crotaline snake pit organs are also innervated by the trigeminal nerve, which contains a very pure population of unimodal warm nerve fibers [1]. Three branches of the trigeminal nerve innervate the pit membrane: one ramus ophthalmicus branch and two ramus maxillaries branches. These trigeminal ganglion (TG) neuron fibers connect ipsilaterally in the lateral descending nucleus of the medulla oblongata [21,22]. This system is independent of the common trigeminal sensory system found in these snakes and other vertebrates. Other reports have examined the IR afferent system, including the nucleus reticularis caloris [23] and optic tectum [21,24].
In amphibians and mammals, the ophthalmic and maxillomandibular ganglia are fused to form one complex trigeminal or semilunar ganglion. Conversely, the two ganglia remain separate in most reptiles [25]. A prominent septum further separates the two divisions of the maxillomandibular ganglion in crotaline snakes [26]. In a horseradish peroxidase (HRP) tracing study, HRP cell labeling from the pit organ to the TG revealed HRP-positive IR-sensory TG neurons in the ophthalmic ganglion and the maxillary division of the maxillomandibular ganglion but not in the mandibular division [27]. Fig. 3 shows a diagram of the TG structure in crotaline snakes. Based on its specialized structure, several immunohistochemical studies have described the distribution of various neuronal signals in the TG, including calcitonin gene-related peptide [28], neuronal and endothelial nitric oxide synthases [29,30], and protein kinase C delta [31].
Some animals and insects have thermoreceptors, which aid hunting, feeding, and survival. IR radiation sensitivity is much greater in crotaline and boid snakes than in other thermosensitive animals. Furthermore, crotaline snakes have a more sensitive IR receptor than boid snakes because they have a thin thermosensitive membrane suspended between a pair of pits, located in the loreal region, which comprises the pit organ. The pit membrane vasculature, located near IRsensitive terminal nerve masses, the IR receptors, supplies blood for cooling and provides the energy and oxygen required by the IR receptors. The ophthalmic and maxillary branches of the TG nerves innervate the pit membrane. In crotaline snakes, the TG is separated into the ophthalmic and maxillomandibular ganglia and a prominent septum further separates the two divisions of the maxillomandibular ganglion. TG ganglion neurons in the ophthalmic ganglion and the maxillary division of the maxillomandibular ganglion, but not in the mandibular division, connect IR sensation to the brain. Further studies are required to expand what is known about the thermosensitive TRP channels in the crotaline pit organ IR receptor.
www.acbjournal.com www.acbjournal.org
www.acbjournal.com www.acbjournal.org doi: 10.5115/acb.2011.44.1.8
The author thanks
Sickle cell disease (SCD) is one of the commonest severe inherited disorders, but specific treatments are lacking and the pathophysiology remains unclear. Affected individuals account for well over 250,000 births yearly, mostly in the Tropics, the USA, and the Caribbean, also in Northern Europe as well. Incidence in the UK amounts to around 12-15,000 individuals and is increasing, with approximately 300 SCD babies born each year as well as with arrival of new immigrants. About two thirds of SCD patients are homozygous HbSS individuals. Patients heterozygous for HbS and HbC (HbSC) constitute about a third of SCD cases, making this the second most common form of SCD, with approximately 80,000 births per year worldwide. Disease in these patients shows differences from that in homozygous HbSS individuals. Their red blood cells (RBCs), containing approximately equal amounts of HbS and HbC, are also likely to show differences in properties which may contribute to disease outcome. Nevertheless, little is known about the behaviour of RBCs from HbSC heterozygotes. This paper reviews what is known about SCD in HbSC individuals and will compare the properties of their RBCs with those from homozygous HbSS patients. Important areas of similarity and potential differences will be emphasised.
Like homozygous HbSS individuals, individuals heterozygous for HbS and HbC (HbSCs) suffer from sickle cell disease (SCD) [1][2][3][4][5][6]. The condition in HbSC patients (here called HbSC disease cf. HbSS disease in homozygotes) not only has some overlap with that seen in HbSS patients, but also has distinctive laboratory and clinical features identifying it as a separate entity [6][7][8]. Although HbSC disease is one of the commonest significant genetic diseases worldwide, it is comparatively neglected with very few laboratory or clinical studies addressing the condition directly. Thus, whilst extensive research has been carried out on understanding SCD in HbSS patients, little relates specifically to the pathogenesis in HbSC patients. In clinical trials of potential novel therapies for SCD, HbSC patients are often specifically excluded. Furthermore, most clinical and laboratory features of HbSC disease have been inferred from studies of HbSS, which may not be appropriate. This paper addresses the pathophysiological differences shown by SCD in HbSS and HbSC patients and the diversity in their clinical complications. Particular reference is paid to the transport abnormalities of the RBC membrane.
All SCD patients have the abnormal haemoglobin HbS in their red cells instead of the normal adult HbA [9][10][11][12]. HbS results from a single base mutation in codon 6 of the βglobin gene which causes a single amino acid substitution in position β6 (glutamic acid → valine, with net loss of one negative charge). Homozygous HbSS patients have two copies of the altered gene. The mutation arose in West Africa, where the high prevalence of HbS appears to be due to selection pressure conferred by a relative resistance to malaria. Malaria resistance has also increased the prevalence Anemia of a second abnormal Hb, HbC, which like HbS represents one of the most prevalent forms of abnormal human Hb. HbC also has a single mutation/amino acid change at the same position in β-Hb, but with lysine replacing glutamic acid (hence net loss of two negative charges). These changes in protein charge may alter how the different Hbs interact and modulate transporter function at the RBC membrane [13]. The charge differences are also used for electrophoretic tests for abnormal Hb, although care must be exercised to exclude certain non-SCD haemoglobinopathies which may mimic HbSC. Homozygous HbCC individuals show few disease symptoms apart from a mild haemolytic anaemia [6]. Heterozygotes of HbA with either HbS or HbC are also largely asymptomatic. Coinheritance of HbS and HbC to produce HbSC heterozygotes, however, results in a clinically significant disease similar, but not identical, to that in HbSS individuals [6][7][8]. Although globally HbSC heterozygotes represent about a third of SCD cases, their distribution is by no means uniform. HbC appears to have originated in Burkina Faso [6] where HbSC cases may outnumber those of HbSS. In other areas, such as the Middle East and India, HbSC cases are rare. In this context, it is worth pointing out that estimates of the frequency of different haemoglobinopathies are likely to be inexact, relying on outdated or incorrect information [14].
All cases of SCD, including those of HbSC disease, are characterised by shortened red cell life span and chronic anaemia, together with recurrent episodes of more acute vaso-occlusion, tissue ischaemia, and increased mortality [12]. Affected individuals have a poor quality of life with numerous complications, for example, pain, cerebrovascular disease (strokes), renal and pulmonary damage, leg ulcers, and priapism [2]. An important feature of SCD is that the clinical scenario is notably heterogeneous-patients may present with mild forms of the disease which rarely require medical intervention or alternatively with more severe complications warranting frequent hospitalisation and aggressive management. Presumably modifier genes and/or environmental factors are significant, but although this area is now receiving considerable attention, it remains poorly understood at present [15][16][17].
In most cases, HbSC disease is clinically milder than HbSS disease and the various complications of SCD usually occur less often or later in life [8]. For example, leg ulcers and other chronic vascular manifestations occur infrequently. Loss of splenic function is relatively delayed, preserving red cell scavenging and thereby possibly affecting disease complications. Nevertheless, HbSC disease still has a significant impact on patients who show haemolytic anaemia, organ failures (stroke, renal failure, chronic lung disease.), and increased mortality (with a median survival of 60 years for males in USA) [15,18]. Pregnant women sometimes develop complications having been hitherto asymptomatic [19]. Complications also occur in children with the risk of stroke in childhood being about 100 times greater than that in the general population [18]. Furthermore, in HbSC heterozygotes, some of the serious complications of SCD (such as osteonecrosis) are as common as for HbSS patients and some (e.g., proliferative sickle retinopathy and possibly acute chest syndrome) occur more frequently [8]. This is also apparent for some central auditory and vestibular problems [19].
Additionally, HbSC is haematologically distinct from HbSS, with higher Hb levels (but lower levels of HbF), lower rates of haemolysis and lower white cell counts [8]. Some of these features are well illustrated in clinical and haematological observations on patients from our clinics (see [20,Table 1]). These distinctive features imply that individuals with HbSC disease should be treated as a discrete subset of SCD patients.
Currently there is very little specific information on the pathophysiology and management of HbSC disease, with much being inferred from studies of HbSS patients. Differences in pathogenesis between HbSC and HbSS disease are expected, however. Understanding them will be important in the management of HbSC patients and may also contribute to a better appreciation of the condition in homozygous individuals.
Although the underlying molecular defect of SCD is long established, how HbS results in the clinical complications remains poorly understood. The chronic anaemia and acute ischaemic episodes both are associated with altered rheology and increased adhesiveness of both RBCs and vascular endothelium [21]. RBCs are more fragile and more readily scavenged from the circulation, contributing to the chronic anaemia, whilst microvascular occlusion is also encouraged causing the acute ischaemic events characteristic of SCD. Intravascular haemolysis is observed and the consequent release of Hb to circulate freely in plasma contributes to the vasculopathy, probably by scavenging nitric oxide (NO) and causing a functional deficiency of that molecule [22,23]. Some authors divide the disease complications of SCD into two broad categories, with sequelae caused either predominantly by altered RBC rheology and elevated blood viscosity (e.g., pain, osteonecrosis, and acute chest syndrome) or by intravascular haemolysis and NO scavenging (e.g., pulmonary hypertension, stroke, priapism, and leg ulceration) [24][25][26]. In any event, polymerisation of HbS on deoxygenation is central to anaemia, vaso-occlusion, and haemolysis-although complete deoxygenation may not be needed, especially in the case of hyperdense RBCs with high cell [Hb], such as some of those found in HbSC individuals. Formation of long rods of HbS distorts RBC shape, reduces deformability, and increases viscosity, thus compromising vascular red cell rheology [27]. Other key events in the pathogenesis have been identified. First, red cell volume is critically important [28]. The increased cation permeability of HbS-containing red cells results in solute loss with water following osmotically. Consequently, [HbS] increases. As the rate of HbS polymerisation upon In RBCs from homozygous HbSS individuals, high cation permeability is accounted for by three main pathways [28,37]. Under oxygenated conditions, the KCl cotransporter (KCC) is highly active. It is overexpressed in HbSS cells compared to HbAA ones and does not become quiescent as RBCs mature. It is stimulated further by low pH (reduction in extracelluar pH from 7.4 to 7). Under deoxygenated conditions, KCC remains active-again unlike the situation in HbAA RBCs [40]. In addition, two other pathways are observed. The deoxygenation-induced cation conductance (or P sickle ) is activated as HbS polymerises. It mediates entry of Ca 2+ . Elevation in intracellular Ca 2+ then leads to activation of the third pathway, the Ca 2+ -activated K + channel, or Gardos channel. These three pathways result in solute loss, cell shrinkage and dehydration, and consequent increase in [HbS]. They thereby contribute to pathogenesis of sickle cell disease. They are also likely to be involved in solute loss from RBCs of patients heterozygous for HbS and HbC (HbSC genotype), though details are lacking and differences in their behaviour compared to that in HbSS cells are expected.
deoxygenation is proportional to a very high power of [HbS], a small reduction in cell volume and hence increase in [HbS] markedly encourages HbS polymerisation [27]. Second, red cell stickiness is also increased [21,[29][30][31]. This results, at least in part, from exposure of phosphatidylserine (PS) on the outer bilayer of the membrane [32]. Exposed PS is prothrombotic and increases adherence of red cells to macrophages and endothelium, contributing to chronic anaemia, haemolysis, and vaso-occlusion [33]. Again, the cause of PS exposure is not clear, but sickling-induced Ca 2+ entry may play an important role [34]. Third, SCD represents an inflammatory state with raised levels of cytokines, chronic elevation of leukocyte counts, shortened leukocyte half-life, and abnormal activation of granulocytes, monocytes, and endothelium [35,36]. The resulting cytokine stimulation of endothelial cells increases their adhesiveness to sickle RBCs [35,36]. How these various changes interact to produce the symptoms of SCD represents a major research challenge. In addition, the extent to which these various mechanisms are involved in disease pathogenesis would be expected to differ between HbSS homozygotes and HbSC heterozygotes. For example, reduced intravascular haemolysis in the latter may ameliorate NO scavenging.
Increased membrane permeability of HbS-containing red cells contributes to SCD pathogenesis by promoting Ca 2+ entry, KCl loss with water following osmotically, and hence RBC dehydration [28,34,37]. In HbSS cells, the involvement of three pathways has been proposed: the KCl cotransporter (KCC), the deoxygenation-induced cation conductance (or P sickle ), and the Ca 2+ -activated K + channel Anemia (or Gardos channel, KCNN4) [28]. These three systems are illustrated schematically in Figure 1. The first of these, KCC (likely KCC1 and KCC3 isoforms), is more active and abnormally regulated in HbSS cells [38][39][40]. Mean activity is enhanced >10-fold in unstimulated cells with several stimuli increasing activity further. In normal RBCs, cell swelling is an important trigger of KCC activity [41]. For HbSS cells, however, intracellular pH is probably the most important stimulus in vivo, with KCC activity reaching a peak at about pH 7 [38,42]. The transporter also responds to O 2 tension [43]. In normal red cells, high levels of O 2 are required for KCC activity, with the transporter becoming inactivated at low O 2 . By contrast, in HbSS cells, the transporter remains active during full deoxygenation, thereby allowing it to respond to low pH in hypoxic areas (like active muscle beds) [40] (Figures 1 and 2). KCC is regulated by phosphorylation, through cascades of conjugate protein kinases and phosphatases [44], with differences apparent in HbSS cells compared with HbAA ones, but at present these are poorly defined. The relative deficiency of intracellular Mg 2+ in HbSS cells [45,46] probably acts to increase KCC activity by altering the activity of these regulatory enzymes.
The second pathway, P sickle , is apparently unique to HbS-containing red cells [28,34]. It is activated to a variable extent by deoxygenation, HbS polymerisation, and shape change [47,48] (Figures 1 and 2). P sickle has the characteristics of a nonspecific cation channel [34]. An anion permeability is controversial, whilst, more recently, it has been proposed as permeable under certain conditions to nonelectrolytes [49]. The main effect of P sickle is probably the increased Ca 2+ entry [49,50] and possibly the Mg 2+ loss [45]. Raised intracellular Ca 2+ has several roles which include phospholipid scrambling [51]. It will also activate the third pathway responsible for HbS cell dehydration, the Gardos channel [52] (Figures 1 and 3). The Gardos channel is then capable of mediating very rapid efflux of K + with Cl -following for electroneutrality and water osmotically.
These mechanisms cause solute loss and HbSS cell shrinkage. Episodes may be short lived and produce only modest degrees of solute loss. But they may occur repeatedly during the lifetime of the RBCs, often during deoxygenationinduced sickling events. Accordingly, HbSS cells show an increase in MCHC of a few percent compared to normal red cells (c.34 g•dL -1 cf. 33 in HbAA cells, density approx 1.085 g•mL -1 ), but importantly there is a large range about this mean with many dense cells (>1.095 g•ml -1 , MCHC c.38 g•dl -1 ), some of which are exceedingly dense (1.125 g•ml -1 , c.50 g•dl -1 ) [53]. A significant feature of HbSS RBCs is their marked heterogeneity, with certain subpopulations possibly more important in pathogenesis [28]. The densest HbSS cells are mainly older ones, presumably following repeated episodes of solute loss [54]. Reticulocytes are mostly low density (c.26 g•dl -1 ), as they are in normal individuals [55]. However, there is a small fraction of young, dense HbSS cells, the so-called fast-track reticulocytes, which become rapidly dehydrated on deoxygenation while still young [28,56]. Although this technique measures a K + influx, because of the high K + content of RBCs, net solute movement through the transport systems will be outwards. KCl cotransport activity was calculated as the Cl -dependent K + influx, Gardos channel activity as the clotrimazole (5 μM)-sensitive K + influx, and P sickle as the Cl --independent K + influx (Cl -substituted with NO 3 -). Sickling, P sickle , and Gardos channel activation occurs in deoxygenated conditions-as for HbSS RBCs-but KCC activity is low when O 2 is removed (as in RBCs from HbAA cells). Data taken from [64].
RBCs from HbSC patients also show K + loss, raised MCHC and haemolytic anaemia with reticulocytosis [5,[57][58][59]. The properties of HbSC RBCs, however, differ in important respects from those of HbSS cells. In HbSC cells, K + loss and dehydration are markedly more pronounced [58,59]. MCHC is particularly high, at about 37 g•dl -1 (cf. 33 g•dl -1 in HbAA individuals; 33-34 g•dl -1 for the reversibly sickled fraction of HbSS patients) [57,58]. Whilst most reticulocytes from normal HbAA and HbSS are characteristically low density (26 g•dl -1 ), HbSC reticulocytes are mainly high density (MCHC c.34 g•dl -1 ) [5,58]. Usually older RBCs are denser; however the monotonic decrease in reticulocyte count with increasing cell density observed for red cells from HbSS patients (as well as HbAA and HbAS individuals) does not occur in HbSC patients [5,60]. Instead, HbSC reticulocytes are fairly evenly distributed across the different RBC densities [5], or perhaps even more concentrated in the denser fractions [60]. This has been taken as evidence that a significant proportion of young HbSC cells begin their lives with a high density [5], rather than undergoing a more gradual dehydration observed in HbSS cells upon repeat Test potentials from -80 to +80 mV were applied for 300 ms in 10 mV increments from a holding potential of -10 mV. Measurements were made using Na + -containing bath and pipette solutions. Data taken from [20]. See [66] for experimental details. The conductance of RBCs from HbSC patients is high and increases further on deoxygenation.
episodes of sickling. In this respect, perhaps the majority of HbSC reticulocytes behave like the "fast-track" reticulocytes of HbSS patients [56]-cells which dehydrate rapidly on leaving the bone marrow-but this remains to be established. It also raises the question as to what constitutes RBCs in the less dense HbSC fractions. Can shrunken HbSC cells regain lost solute and increase their volume? If so, what is the mechanism and what are transport systems involved?
As heterozygotes, HbSC cells contain both HbS and HbC, in approximately equal amounts (i.e., 50%). This contrasts with the lower HbS content (c.40%) found in sickle trait HbAS cells [5]. Crystals of HbC are sometimes present in oxygenated RBCs. In contrast to HbS polymers, these deposits are lost on deoxygenation [60]. Because of the high HbS content and polymerisation, HbSC cells also show a deoxygenation-induced sickling shape change. In this case, however, rather than the HbSS sickles and holly leaf forms, deoxygenated HbSC cells show multifolded shapes such as "pita breads" and "tricorns" [60], perhaps because of the high surface area to volume ratio subsequent to their more marked dehydration. How HbS and HbC interact has also received some attention. Using different Hb mixtures, a direct interaction between the two Hbs appears to only slightly enhance HbS polymerisation. Much more important in HbSC disease is RBC dehydration and consequently the high MCHC [5,61,62]. High MCHC and lower levels of HbF may have an effect on the extent and kinetics of HbS polymerisation whilst concurrent Hb mutations (such as βthalassaemia) may also play a significant role.
It is therefore critical to understand fully the mechanisms by which these RBCs shrink, but our understanding of the mechanisms involved remains uncertain. Oxygenated HbSC cells have elevated KCC activity that is stimulated by low pH and swelling [13,60]. Cytoplasmic protein concentration has been suggested as the "volume" sensor of RBCs [63]. It is therefore intriguing to speculate that high KCC activity in oxygenated HbSC cells may result from the presence of HbC crystals which would lower the total concentration of soluble Hb, as occurs in swollen RBCs. In effect, the cells "think" that they are swollen and so activate mechanisms to lose solutes and water, namely, KCC. On the other hand, deoxygenated HbSC cells also show increased K + efflux, to an extent Anemia apparently greater than that observed in deoxygenated HbSS cells [58]. Which pathway mediates the flux in deoxygenated conditions, however, has not been established. If P sickle is involved, given the lower [HbS] of HbSC cells, it is not clear why it should be activated to a greater extent than in HbSS cells. KCC and the Gardos channel represent obvious alternative pathways.
A number of manoeuvres which reduce reticulocyte density may provide evidence for the transport pathways involved in their dehydration. Both Cl -removal and deoxygenation shift HbSC reticulocytes to lower densities, consistent with solute retention following inhibition of KCC [60]. Hypotonic swelling of HbSC cells also reduces the deoxygenation-induced K + loss [58], perhaps through reduction in [HbS] removing a P sickle -like element of K + flux.
In preliminary studies, we have observed high KCC activity in oxygenated unfractionated HbSC cells, which was almost completely inhibited on deoxygenation [64]. Thus, KCC in HbSC cells behaved like that in HbSS cells at high O 2 tension and like that in HbAA cells when tension was reduced [64] (Figure 3). We also found activation of a deoxygenation-induced Cl --independent K + flux [64], a deoxygenation-induced nonelectrolyte permeability [65] and a deoxygenation-induced rise in K + conductance in patch-clamp experiments [66] (Figure 4), namely, a P sicklelike permeability, together with activation of the Gardos channel. In this context, it is interesting that HbC has a higher affinity for the RBC membrane than either HbA or HbS [59] leading to an early suggestion that HbSC interaction is involved in modulating RBC permeability.
It is apparent, however, that our understanding of the permeability of HbSC cells requires further investigation.
Understanding dehydration is particularly relevant for HbSC cells. The solubility of deoxygenated HbS is about 17 g•dl -1 compared to 70 g•dl -1 for HbA. As HbS represents only about half the total Hbs in HbSC cells, a relatively small decrease in MCHC (from an RBC total of 37 to 33 g•dl -1 ) will prevent HbS polymerisation [61] while retaining the functionally important discocyte morphology. As HbS constitutes about half the Hbs in these RBCs, this would mean a fall in [HbS] from 18.5 to 16.5. In comparison, in HbSS cells, a reduction of MCHC to <25 g•dl -1 is required, by which time RBCs will be spherocytic, and cell swelling per se will adversely affect rheology [58]. Notwithstanding their relevance to dehydration and sickling, the permeability of HbSC cells has not been well studied nor compared in detail with that of HbSS cells. Several areas require more careful investigation. The interaction between different Hbs, membrane target sites regulating permeability, the transport pathways involved, the role of cell density, oxygenation, volume, and pH presents a complex pattern of modalities controlling solute content and hence cell density and MCHC. Control by phosphorylation remains mainly unexplored. The challenge ahead lies to define the most important stimuli and how they interact to determine cell volume. A major therapeutic goal is the ability to prevent HbSC cell dehydration or to promote rehydration.
The authors thank the
Here we term red cells from HbSS patients as HbSS cells, those from HbSC individuals as HbSC cells.
Anesthesia options for upper extremity surgery include general and regional anesthesia. Brachial plexus blockade has several advantages including decreased hemodynamic instability, avoidance of airway instrumentation, and intra-, as well as postoperative analgesia. Prior to the availability of ultrasound the risks of complications and failure of regional anesthesia made general anesthesia a more desirable option for anesthesiologists inexperienced in the practice of regional anesthesia. Ultrasonography has revolutionized the practice of regional anesthesia. By visualizing needle entry throughout the procedure, the relationship between the anatomical structures and the needle can reduce the incidence of complications. In addition, direct visualization of the spread of local anesthesia around the nerves provides instant feedback regarding the likely success of the block. This review article outlines how ultrasound has improved the safety and success of brachial plexus blocks. The advantages that ultrasound guidance provides are only as good as the experience of the anesthesiologist performing the block. For example, in experienced hands, with real time needle visualization, a supraclavicular brachial plexus block has changed from an approach with the highest risk of pneumothorax to a block with minimal risks making it the ideal choice for most upper extremity surgeries.
Anesthesia options for upper extremity surgery include general anesthesia, regional anesthesia, or a combination of the two. In the past general anesthesia was frequently the method of choice for upper extremity surgery due to lack of training and experience with regional anesthesia as well as fear of complications including vascular puncture, local anesthetic toxicity, pneumothorax, and patient discomfort [1]. Needle placement utilizing the paresthesia technique or peripheral nerve stimulator could be a time-consuming process leading to operating room delays and patient discomfort. However, the advantages of general anesthesia, including control of the airway and familiarity of the technique by the majority of anesthesiologists are overshadowed by the clear benefits of regional anesthesia. These include intraoperative, as well as postoperative analgesia [2,3]. In addition, regional anesthesia results in excellent muscle relaxation during surgery, decreased opioid requirements and their potential side effects, greater hemodynamic stability, increased efficiency in the operating room by avoiding the time required to awaken and extubate the patient, reduced PACU stay, a decrease in unplanned hospital admission for pain control, as well as greater patient satisfaction [2,3]. The most significant advantage of regional anesthesia for surgery of the upper extremity is the prolonged postoperative analgesia that a nerve block can provide. The pain relief following brachial plexus blockade with long-acting local anesthetics such as bupivacaine, ropivacaine and levobupivacaine has resulted in patients being discharged home on the day of surgery as opposed to a planned or unplanned overnight admission [3,4].
Thompson and Rorie performed cadaveric studies to map out the brachial plexus anatomy [5]. The fifth through eighth cervical and first thoracic nerve exit through the intervertebral foramina and travel along the groove formed by the transverse processes of their corresponding vertebrae. After exiting the transverse processes the roots of the brachial plexus travel between the anterior and the middle scalene muscles, identified as the interscalene groove [1]. The authors reported that the brachial plexus is confined by a continuous fascial sheath formed by the deep cervical fascia and that the fascia is continuous from emergence of the nerve roots to the axilla. Distally the fascia folds inwards to form separate compartments for each nerve. For example, at the cord level, local anesthetic injected around the posterior cord may not spread to include the lateral or medial cords [6][7][8] whereas an axillary brachial plexus block is performed by identifying the individual nerves and blocking each one separately [9][10][11]. While an interscalene block is performed by means of a single injection technique, the more distal approach to the brachial plexus, the less likely that a single injection technique will result in complete blockade of the upper extremity [12][13][14].
Ultrasound probes (transducers) act as both a transmitter and receiver of sound waves. The probes are classified as either high (10-15 MHz), midrange (5-10 MHz), or low (<5 MHz) frequency. High-frequency probes provide highresolution images but lack depth of penetration compared to low-frequency probes [15]. Both frequency types are available with a wide or a narrow footprint. High-resolution linear transducers are most suitable for imaging superficial structures such as the brachial plexus in the interscalene, supraclavicular, and axillary regions. The lower frequency curved transducer is preferable when the anatomical structures are deeper than 4 cm, for example, when performing an infraclavicular block [16]. Prior to the use of ultrasound, block needle placement was achieved using a blind approach with the nerve stimulation or paresthesia technique. Ultrasound imaging has revolutionized the practice of regional anesthesia in that the operator can visualize and identify nerves and blood vessels as well as the needle during its passage through the tissues. Abnormal anatomy can also be recognized [17]. In addition, direct visualization of the spread of local anesthetic decreases the risk of intravascular injection, local anesthetic toxicity, pneumothorax, and a failed block [18]. It is important to remember, however, that the success of an ultrasound-guided block is dependent upon the skill and experience of the anesthesiologist. Anesthesiologists performing brachial plexus anesthesia under ultrasound guidance must first become comfortable with identifying anatomical structures as well as visualizing the needle during the block performance. In experienced hands, the benefits of performing a peripheral nerve block with real-time ultrasound imaging of needle placement and local anesthesia spread include decreased performance as well as onset time, a decreased dose of local anesthetic required to achieve a successful block, and an increase in block success rate [19][20][21][22][23][24]. In a systematic review and meta-analysis of randomized controlled trials comparing ultrasound guidance with electrical neurostimulation for peripheral nerve blocks, Abrahams et al. confirmed these aforementioned benefits of ultrasound. However, the authors concluded that larger studies are needed to determine whether ultrasound can decrease the number of complications [25].
The success of a regional anesthesia program is dependent on patient education and the support of the surgical team.
To put a patient at ease it is desirable for the surgeon to inform the patient about the possibility of receiving a brachial plexus block for his or her surgery prior to the day of surgery. A patient that has been informed beforehand is often more amenable to accepting regional anesthesia. In addition, the training, education, and skill of the individual performing the block are of paramount importance. To this end, both the American as well as the European Society of Regional Anesthesia (ASRA and ESRA) hold annual meetings as well as numerous workshops throughout the year to educate and train individuals in the art of regional anesthesia [26]. There are few absolute contraindications to a brachial plexus block. These include patient refusal, local anesthetic allergy, infection at the site of needle entry, and the presence of infected lymph nodes in the axilla or supraclavicular region [7,27]. Deep blocks, for example, an infraclavicular approach, as well as blocks in the vicinity of a noncompressible artery (e.g., supraclavicular) should not be performed in coagulopathic patients [6,7]. A patient that is unable to cooperate secondary to decreased mental status is also an absolute contraindication. Regional anesthesia is not contraindicated in patients that have a pre-existing stable neurological deficit or chronic neurological disease provided that the condition is well documented [28,29].
It is up to the anesthesiologist to decide, based on each individual patient's risk benefit ratio whether performance of a peripheral nerve block in the presence of pre-existing nerve damage is indicated [30]. An informed cooperative patient is an essential factor in ensuring safe and effective regional anesthesia. Following a brachial plexus block, it is essential that the affected extremity be immobilized and protected until loss of sensation and proprioception have resolved. Patients and family members should receive clear instructions regarding the anticipated duration of the block and how to transition to oral analgesia at home to avoid the sudden onset of pain.
De Andres and Sala-Blanch state that it is essential to understand both the topographic anatomy and cross-sectional anatomy of each anatomic zone of the brachial plexus [15]. They describe the brachial plexus as being divided into three zones: the supraclavicular region in the posterior triangle of the neck, the infraclavicular region deep to the pectoralis muscles in the anterior chest, and the axillary region. The level of needle entry in one of these zones will determine the extent, limitations, and potential complications of a brachial plexus block. Prior to the use of ultrasound, the likelihood of encountering the spinal cord, the lung, and major vessels such as the subclavian and vertebral arteries with the more proximal approaches (interscalene and supraclavicular) was a concern [31]. Ultrasound has minimized these risks provided that the needle tip as well as the spread of local anesthetic is constantly visualized throughout performance of the block [20,32]. The choice of which technique to use is dependent on the surgical procedure, the comfort and expertise of the anesthesiologist, and patient-associated factors such as sepsis in the axilla. In the latter case a more proximal approach is desirable [30].
The interscalene approach to the brachial plexus is the technique of choice for surgical procedures of the shoulder. It is inappropriate for surgeries involving the medial aspect of the upper extremity due to inconsistent blockade of the lower trunk (C8 and T1) [33,34]. At the level of the cricoid cartilage the brachial plexus trunks appear as three distinct hypoechoic areas between the anterior and middle scalene muscles [1,35]. It should be emphasized that the large doses of local anesthetic traditionally used for an interscalene block with the neurostimulation technique result in a 100% incidence of ipsilateral phrenic nerve paralysis due to blockade of the 3rd, 4th, and 5th cervical nerve roots. This may decrease the patient's FRC by 25% [36] and may therefore not be suitable for patients with emphysema and other chronic lung diseases with decreased pulmonary reserve. Ultrasound imaging improves the interscalene approach primarily by being able to visualize the spread of local anesthetic within the fascia surrounding the trunks. This direct visualization decreases the amount of local anesthetic needed to provide surgical anesthesia [37,38]. Decreasing the volume of local anesthetic to 10 mL or 5 mL resulted in a significant decrease in the incidence of hemidiaphragmatic paresis [37,39,40]. Kapral et al. report that ultrasound guidance improves both the quality of the nerve block and shortens the time of onset of sensory blockade [22].
The supraclavicular block was traditionally performed for surgeries of the upper extremity below the shoulder. Liu et al., however, recently reported that ultrasound-guided supraclavicular blocks are effective and safe for shoulder arthroscopy [41]. The supraclavicular approach has several advantages over the more distal approaches including rapid onset of the block, more complete blockade of the nerves supplying the upper extremity (with the exception of the intercostobrachial nerve) due to the compact arrangement of the trunks of the brachial plexus at this level [32,42]. Prior to the use of ultrasound the supraclavicular approach was frequently avoided, particularly in ambulatory surgeries, due to the increased risk of pneumothorax and, to a lesser extent, direct vascular puncture of the subclavian, superficial (transverse) cervical, suprascapular, or dorsal scapular arteries with subsequent local anesthetic toxicity and cardiovascular collapse [43,44]. Ultrasound has improved the safety of a supraclavicular block as the anesthesiologist can now visualize the subclavian artery, the first rib, as well as the dome of the lung. Placement of the needle and spread of the local anesthetic can now be seen in real-time resulting in resurgence in the use of this block [17,45]. Chan et al. examined the supraclavicular region in 40 patients and reported that in all cases the nerves of the brachial plexus appeared as hypoechoic nodules in clusters lateral, posterior, and cephalad to the subclavian artery [46]. The authors also concluded that if the needle is seen at all times and not inserted beyond the first rib, then the risk of a pneumothorax in a supraclavicular block is essentially eliminated. Williams found that supraclavicular nerve blocks were performed faster with ultrasound guidance when compared with nerve stimulation (5 versus 10 min) [20]. Ultrasound guidance has increased the safety profile of the supraclavicular approach so that in experienced hands this may be the block of choice for most upper extremity surgeries below the shoulder.
The infraclavicular approach is indicated for surgeries of the arm and hand. Compared to the supraclavicular approach, the risk of pneumothorax is significantly reduced and is virtually eliminated with the use of ultrasound. In addition, the phrenic nerve is not blocked with this approach [47]. Compared to the axillary approach, the infraclavicular block targets the brachial plexus at the level of the cords which surround the second part of the axillary artery and are proximal to the takeoff of the musculocutaneous, axillary, and medial brachial cutaneous nerves. This may result in a higher success rate of complete blockade with a single injection technique [48]. The infraclavicular anatomy may, however, be more difficult to visualize under ultrasound guidance particularly in obese patients. Perlas et al. found that compared to the interscalene, supraclavicular, axillary, and midhumeral approaches in which the brachial plexus was visualized 100% of the time, in only 27% of patients were they able to visualize the infraclavicular brachial plexus [1,49]. This difficulty is due to the relative depth of the brachial plexus in the infraclavicular approach compared to all other approaches to the brachial plexus. A low-frequency probe with its greater tissue penetration may facilitate performance of this block [47]. In an ultrasound-guided infraclavicular block the lateral, posterior, and medial cords are seen in close proximity to the axillary artery and vein. The posterior cord is usually blocked first. If the spread of local anesthetic does not surround the lateral and medial cords, then all three cords must be blocked individually to obtain complete blockade of the upper extremity. As with the supraclavicular and axillary approaches, the intercostobrachial nerve will have to be blocked separately in the axilla to achieve anesthesia of the inner aspect of the upper arm.
Axillary blocks are performed for procedures of the elbow, distal arm, and hand. Prior to the use of ultrasound, the axillary approach was the most common approach to the brachial plexus due to the safety of this technique. As with the infraclavicular block, the risk of phrenic nerve paresis is avoided and the risk of pneumothorax is eliminated. The high success rate without respiratory compromise makes this block desirable in patients with reduced lung capacities and chronic pulmonary diseases [50]. Contraindications to the axillary approach include inability to abduct the arm to the position necessary to perform the block, localized infection in the axilla, or enlarged axillary lymph nodes [51]. A major advantage of using ultrasound in an axillary approach is the ability to confirm blockade of the musculocutaneous nerve [52]. Because of the anatomical variance of the musculocutaneous nerve in relation to the axillary artery, failure to block this nerve with a perivascular approach in not uncommon [53]. At the axillary level, the terminal branches of the brachial plexus (median, ulnar, and radial nerves) are situated close to the axillary artery and veins with the two axillary veins situated medial to the artery. There is, however, a great deal of variation in the distribution of these three nerves in relation to the artery [54]. The four nerves are easily visualized utilizing ultrasonography. The ultrasound guided axillary approach has been shown to both decrease block failure rate and time of onset of sensory blockade compared to the transarterial technique [14]. The success of US guided axillary blocks depends on the multiple needle approach in which each nerve is identified individually and spread of local anesthetic is observed around the median, ulnar, radial, and musculocutaneous nerves.
Individual terminal nerve blocks can be performed at the midhumeral, elbow, forearm, or wrist either by design or as a rescue block [7]. These more peripheral nerve blocks may be performed to achieve postoperative analgesia while at the same time maintaining more proximal control of the upper extremity.
Postoperative pain and nausea are the leading causes of unplanned hospital admission after ambulatory surgery [55]. Orthopedic upper extremity surgery is reported as having a high incidence of severe pain [56]. One of the clear benefits of regional anesthesia over general anesthesia for upper extremity surgery is the postoperative pain relief a long-acting local anesthetic can provide. The choice of local anesthetic is determined by the duration of surgery, necessity of motor blockade, urgency of neurological assessment after surgery, and the anticipated requirement for postoperative analgesia.
In brachial plexus nerve blocks short-, intermediate-, and long-acting local anesthetics can be chosen. Bupivacaine, ropivacaine, and levobupivacaine are equally effective in surgeries in which extended postoperative analgesia would be beneficial. In comparison to bupivacaine, however, ropivacaine and levobupivacaine are the long-acting local anesthetics of choice due to their decreased cardiotoxicity [57][58][59]. It is important for the patient to be informed of the anticipated duration of the local anesthetic so that he or she will not be concerned about the length of time it takes for the block to wear off. It is also essential that patients be instructed regarding protection of the extremity until sensation has completely returned. Finally, patients should be instructed to take their prescribed oral analgesics at the earliest sign of pain to mitigate against the analgesic gap that may otherwise develop. Additional methods to improve and or prolong postoperative analgesia include insertion of a brachial plexus catheter to provide continuous regional analgesia [51,60] as well as the use of multimodal analgesia.
In the multimodal approach use of a long-acting peripheral nerve block in combination with acetaminophen, NSAIDs (when not contraindicated), and oral opioid analgesics will minimize the total opioid requirements and their resulting side effects [61,62].
The various approaches to the brachial plexus afford the anesthesiologist the ability to provide both excellent intraoperative anesthesia as well as postoperative analgesia with minimal complications and increased patient satisfaction following upper extremity surgery. The advantages over general anesthesia are numerous when performed in skilled hands. Ultrasound guidance with real-time needle visualization in relation to anatomic structures and target nerves makes regional anesthesia safer and more successful. With ultrasound guidance in experienced hands, brachial plexus blockade can lead to decreased block performance and onset time, increased success rate and decreased rate of complications. These advantages result in increased operating room efficiency, as well as increased patient and surgeon satisfaction.
Recently, a general picture has been proposed of how long, and to what extent, native protein structure can be retained in the gas phase. [1a] In particular, molecular dynamics simulations suggest that salt bridges and ionic hydrogen bonds on the protein surface can transiently stabilize the global fold shortly after desolvation. [1b] However, the use of native mass spectrometry [2] for studying protein solution structure is still controversial, mostly because site-specific experimental gasphase data [3] is scarce. Here we report electron capture dissociation (ECD) [4] data on the gas-phase structures of the three-helix bundle protein KIX [5] (Figure 1) that indicate
substantial preservation of the native solution structure on a timescale of at least 4 s. We demonstrate that in the gas phase, the most stable regions are those stabilized by salt bridges and ionic hydrogen bonds.
Figure 2 shows site-specific yields of c and zC fragment ions [6] from ECD of (M + n H) n+ ions of KIX (see Figure S1 in the Supporting Information) formed by electrospray ionization (ESI). [7] For the 7 + ions, separated c and zC products were observed only from backbone cleavage near the termini (residues 1-13 and 89-91), but not from the threehelix bundle region, which forms a globular fold around a hydrophobic core (residues 16-88). [5] This observation is consistent with intramolecular interactions in the three-helix bundle region preventing separation of c and zC backbonecleavage products [3a-c] in the gaseous 7 + ions. Collisional activation of the 7 + ions (laboratory-frame energy: 28 eV) prior to ECD effected only marginal unfolding near the N terminus (see Figure S2 in the Supporting Information), revealing a notable stability of the three-helix bundle in the absence of solvent.
For the 8 + ions (Figure 2), the appearance of cleavage products from the N-terminal ends of helices a1 (residues 16-30) and a2 (residues 42-61) indicates partial unfolding, with helix a1 separating from the bundle, and helices a1 and a2 starting to unravel from their N-terminal ends. Unraveling of a1 and a2 continues in the 9 + ions, while helices a2 and a3 appear to largely retain their native antiparallel bundle structure. Separation of a2 and unraveling of a3 (residues 65-88), also from its N-terminal end, is evident from the fragmentation pattern observed for the 10 + ions. However, c-and zC-ion yields in the 65-88 region remained relatively small for the 10 + and 11 + ions, suggesting that partially intact a3 helix structure limits fragment ion separation. Further increasing the precursor ion charge gave increased c-and zC-ion yields and unfolding, similar to ECD data for Ubiquitin [3c] (see Figure S3 in the Supporting Information), with the fragmentation pattern of the 16 + KIX ions being largely unselective with respect to backbone cleavage site.
The data in Figure 2 provide substantial evidence for a correlation between the solution-and gas-phase structures of KIX. This supposition is corroborated by ECD of 12 + ions generated by nano-ESI from a solution (in H 2 O at pH 4.5) that better resembles the native protein environment, [8] which gave decreased c-and zC-ion yields in the a2 and a3 regions (see Figure S4 in the Supporting Information), along with a smaller total fragment ion yield (37 %) relative to that resulting from ECD of 12 + ions from ESI of solutions in H 2 O/CH 3 OH (80:20) at pH 4 (total fragment-ion yield: 49 %; see Figure S3 in the Supporting Information).
The temporal stability of nativelike KIX 7 + ions was studied by introducing a delay between ion trapping and structural probing by ECD. However, the ECD fragmentation patterns showed no significant differences for delay times of 1 ms and 2 s (see Figure S5 in the Supporting Information). To [5] expedite possible structural transitions, we next activated the gaseous 7 + ions by 28 eV collisions (see Figure S2 in the Supporting Information) prior to ion trapping. Despite the increase in ion internal energy, the fragmentation patterns from ECD with delays of 1 ms, 2 s, and 4 s (Figure 3) are strikingly similar. [9] Apparently, the three-helix bundle structure of KIX is sufficiently stabilized by specific noncovalent interactions that outweigh the loss of hydrophobic bonding in the gas phase.
Figure 4 a shows integrated c-and zC-ion yields for helix regions a1, a2, and a3 versus precursor ion charge. The data exhibit sigmoidal behavior, with transition charge values (at 50 % of the plateau value) of 9.2, 10.7, and 12.4 for a1, a2, and a3, respectively. This order of helix stability (a3 > a2 > a1) in the gas phase agrees with that in solution as determined by NMR spectroscopic experiments. [10] However, in solution, each helix unfolds cooperatively, [10] whereas the gasphase data (Figure 1) show incremental unraveling from their N-terminal ends. This behavior is also reflected in the site-specific transition charge values from analysis of site-specific c-and zC-ion yields (see Figure S6 in the Supporting Information), which generally increase from the N to the C terminus (Figure 4 b). Transition charge values for cleavage sites between helix regions (31-41, 62-64) are similar to values for adjacent helix ends, indicating that helix separation does not precede helix unraveling.
Although the ECD data in Figures 2 and 3 demonstrate extensive preservation of the native solution structure in the 7 + ions, its stabilization in the gas phase must be based on interactions other than hydrophobic bonding. [3d,e] These include neutral [11] and ionic [1b, 12] hydrogen bonds, charge-dipole interactions, [13] and salt bridges. [1b, 14] Figure 5 shows helices a1, a2, and a3 with all basic (H, K, R) and acidic (D, E) residues highlighted in color. The density of charged residues is smallest for a1 (5 out of 15 residues, 0.33) and largest for a3 (14 out of 24 residues, 0.58); a2 exhibits an intermediate density of 0.4 (8 out of 20 residues). Importantly, the charge density values correlate (r = 0.9775) with transition charge values (as a measure of helix stability in the gas phase) for a1, a2, and a3 (Figure 6 a). This observation strongly suggests that interactions involving charged residues, that is, ionic hydrogen bonds and salt bridges, largely determine helix stability in the gas phase.
Close inspection of the native KIX structure revealed that one (D17/H21), three (R42/E45, K52/E55, K53/ D57), and six (R65/D66, E67/ H70, E74/K75, K78/E82, K81/ E84, E85/R88) intrahelix salt bridges can stabilize helices a1, a2, and a3, respectively (Figure 5). The density of salt bridges correlates (r = 0.9999) with transition charge values (Figure 6 b) even better than the density of charged residues, suggesting that salt bridges are major determinants for protein structural stabilization in the gas phase. However, this conclusion does not exclude additional stabilization by ionic hydrogen bonds as well as charge-dipole interactions. In particular, interaction of the positive net charge at the C-terminal end of helix a3 (Figure 5) with its electric dipole moment can further stabilize the a3 helix structure, [13] and is consistent with helix unraveling from the N-terminal end.
Stabilization of the global fold by interactions between the three helices probably involves helix dipole/dipole interactions; [15] the antiparallel helices a2 and a3 with larger dipole moments than that of the shorter helix a1 separate and unfold last. Additional stabilization of tertiary structure by ionic hydrogen bonding between charged residues and backbone amides [1b] is indicated by the scatter of site-specific transition charge values (Figure 4 b).
We show here that electrostatic interactions can compensate for the loss of hydrophobic bonding and stabilize the native three-helix bundle structure of KIX in the gas phase on a timescale of at least 4 s. Among these interactions, salt bridges were found to play a dominant role. However, a high number of surface-exposed charged residues alone does not guarantee protein stability in the gas phase: equine Cytochrome c has 24 basic and 12 acidic residues, [3a] with the number of salt bridges on the protein surface increasing from 6 in solution to an average value of 17.3 in the gas phase within 10 ps after desolvation, [1b] yet its native fold disintegrates on a timescale of milliseconds. [3e, 16] The outstanding stability of gaseous KIX ions observed in this study must be attributed to the combination of favorable electrostatic interactions, including salt bridges, neutral and ionic hydrogen bonds, as well as charge-dipole interactions. Whether or not native mass spectrometry can reveal information about the solution structure of a protein critically depends on the timescale of the experiment [1a] and the extent of intramolecular stabilization by electrostatic interactions. KIX is the first protein for which site-specific ECD data indicate preservation of the solution structure in the gas phase. We propose KIX as a model protein for the evaluation of new and emerging methodology for the structural probing of gaseous proteins.
KIX protein (91 residues, GSHMGVRKGW HEHVTQDLRS HLVHKLVQAI FPTPDPAALK DRRMENLVAY AKKVEGD-MYE SANSRDEYYH LLAEKIYKIQ KELEEKRRSR L) was expressed in Escherichia coli cells by using a plasmid that included the CBP KIX coding region [5] (residues 586-672; residue 586 corresponds to residue 5 in this study) and purified by Ni-affinity and size-exclusion chromatography. [10] The purified protein was desalted as described previously. [17] Solution pH was adjusted by addition of acetic acid. Experiments were performed on a 7 T Fourier transform ion cyclotron resonance (FT-ICR) mass spectrometer (Bruker) equipped with an ESI source (flow rate: 1.5 mL min À1 ) and a hollow dispenser cathode operated at 1.6 A for ECD. The desolvation gas temperature was 200 and 150 8C for 80:20 and 50:50 H 2 O/CH 3 OH solutions, respectively. Before ion trapping, precursor isolation (using radiofrequency waveforms), and irradiation with low-energy (< 1 eV) electrons for 17-50 ms in the FT-ICR cell, ions were accumulated in the hexapole ion cells for 0.3-2.0 s. Ion activation prior to ECD was realized in the second hexapole by energetic collisions with Ar gas. Between 250 and 500 scans were added for each ECD spectrum. ECD fragment ion yields were calculated as percentage values relative to all ECD products excluding aC/y ions, [6] considering that backbone dissociation of a parent ion gives a pair of complementary c and zC ions (100 % = 0.5 [c] + 0.5 [zC] + [other products], in which other products are reduced molecular ions and products from loss of small neutral species from the latter).
Angew. Chem. Int. Ed. 2011, 50, 873 -877 2011 Wiley-VCH Verlag GmbH & Co. KGaA, Weinheim
www.angewandte.org 2011 Wiley-VCH Verlag GmbH & Co. KGaA, Weinheim Angew. Chem. Int. Ed. 2011, 50, 873 -877
Angew. Chem. Int. Ed. 2011, 50, 873 -877 2011 Wiley-VCH Verlag GmbH & Co. KGaA, Weinheim www.angewandte.org
Anti-VEGF (vascular endothelial growth factor) therapy with the monoclonal antibody bevacizumab can cause gastrointestinal (GI) perforations. In recent years it became apparent that GI perforations also occur during treatment with antiangiogenic tyrosine kinase inhibitors (TKIs). It is of clinical importance to consider (vague) abdominal complaints during antiangiogenic treatment as a sign of a GI perforation. To illustrate this serious complication, we report four cases of antiangiogenic treatment related GI perforations. In three cases this was due to antiangiogenic TKI treatment. Reported risk factors of GI perforations due to bevacizumab include the presence of a primary tumor in situ and recent history of endoscopy or abdominal radiotherapy. Pathology assessments of surgical removal of the perforated intestinal part reveal that perforations are predominantly seen at the tumor or anastomotic site, in case of carcinomatosis or diverticulitis or when GI obstruction or an intra-abdominal abcess is present. Whether the same risk factors may be involved in antiangiogenic TKI related GI perforations is unknown. The underlying mechanisms responsible for GI perforation during antiangiogenic treatment is unknown, but disturbance of host cell homeostasis of immune cells as well as platelet-endothelial cell interactions may play an important role. In conclusion, while clinical awareness that antiangiogenic treatment can cause GI perforations is critical for current medical practice, it is also very important to get more insight in its underlying mechanisms so that this life-threatening complication may be prevented in the near future.
Malignant tumors depend on the formation of new blood vessels from the pre-existing vasculature for their growth and dissemination [1]. This process, called angiogenesis, is regulated by pro-and antiangiogenic factors. One of the main angiogenic factors is vascular endothelial growth factor (VEGF), which exerts its function by activation of VEGF tyrosine kinase receptors [2]. Multiple agents that target these angiogenic growth factor signaling pathways have been developed. Since these agents only interfere with growth factor signaling pathways in proliferating endothelial cells, serious toxicities from these agents were not expected. Normally, more than 99% of the endothelial cells are quiescent in the absence of malignancy and angiogenesis only occurs during wound healing or in the menstrual cycle [3]. However, in contrast to preclinical tumor models, incidental severe toxicities were observed during clinical development of these agents. For example, incidences of 1.5-5.4% were reported on GI perforations induced by treatment with the humanized monoclonal VEGF-antibody bevacizumab [4,5]. Only a few cases of GI perforations have been reported for antiangiogenic tyrosine kinase inhibitors (TKIs) such as sunitinib or sorafenib.
In this report we present four cases of antiangiogenic treatment related GI perforations, of which in three cases an antiangiogenic TKI was responsible for this complication. In addition, we discuss current views on the potential risk factors and mechanisms of antiangiogenic treatment related GI perforations.
A 74-year-old man with a medical history of right hemicolectomy and hepatectomy for metastasized colon carcinoma was treated in adjuvant setting with oxaliplatin, capecitabine and bevacizumab. Because of rectal blood loss during the second chemotherapy cycle, a colonoscopy with subsequent band ligation of observed hemorrhoids was performed. Three weeks later, the patient was admitted to the hospital with persistent diarrhea, severe anal pain and malaise. Body temperature and blood pressure were normal, but pulse frequency was increased (105 bpm). Anal examination was very painful, but no abnormalities were palpable. Laboratory and faeces examination as well as abdominal and chest X-rays revealed no abnormalities. At colonoscopy multiple deep colonic and perianal ulcers were found and considered as drug induced enterocolitis (Fig. 1). Therefore capecitabine treatment was immediately terminated. Despite this treatment interruption, the patient got worse and subsequently a laparotomy was performed. At the site of the previously placed band ligations (3 weeks before), peri-anal and -rectal necrotic cavities connected to the anal canal were found. The patient recovered within a few weeks after extended necrotectomies, a Hartmann-procedure and antibiotic treatment. No further adjuvant chemotherapy was administered.
A 68-year-old man with a medical history of metastasized renal cell carcinoma (RCC) started treatment with sorafenib, after previous nephrectomy and immunotherapy (interferon-alpha), upon disease progression. Sorafenib treatment resulted in a rapid partial response. However, the patient developed fever and abdominal pain 5 months after start of sorafenib treatment and his condition deteriorated within hours. The patient suffered from diarrhea and substantial rectal bleeding. Physical examination revealed fever, abdominal pain and hepatomegaly. Laboratory results showed anemia and signs of inflammation (hemoglobin 11.8 g/dl (normal value between 13.5 and16.5 g/dl), C-reactive protein (CRP) 57 mg/l (normal value between 0 and 10 mg/l) and Leukocyte counts 9.5 9 10 9 /l (normal value between 4.0 and 10 9 10 9 /l). Computed tomography (CT), performed because of progressive diarrhea together with substantial rectal bleeding with a decrease in hemoglobin to 8.5 g/dl, revealed colonic perforation into the necrotic liver metastasis (Fig. 2). Based on these findings, sorafenib treatment was terminated and antibiotics were prescribed. Because a surgical resection of these necrotic liver metastases was impossible, a terminal ileostomy with a slime fistula was constructed. Within a few days the patient recovered rapidly and could be discharged from the hospital. Two months later, when the patient was fully recovered from this episode, an mTOR inhibitor was prescribed because of disease progression.
Bevacizumab plus an antiangiogenic TKI After optimal interval debulking and extensive treatment with standard chemotherapy, a 67-year-old woman with advanced ovarian cancer and extended peritonitis carcinomatosa participated upon progression in a phase I trial with bevacizumab combined with an experimental antiangiogenic TKI. Targets of this TKI include VEGFR-2 and PDGFR. One month after start of treatment she was admitted to the hospital because of progressive pain in groin and lower left abdominal part accompanied by fever and elevated inflammation parameters (CRP 290 mg/l, Leukocyte counts 13.8 9 10 9 /l). A CT-scan revealed a necrotic tumor mass in the pelvis with, secondary to the tumor response, retroperitoneal perforation. Surgical resection of this retroperitoneal complication was not feasible and optimal palliative care was initiated.
Sunitinib A 62-year-old woman with a medical history of metastasized RCC, for which she underwent surgery and radiotherapy, started sunitinib upon disease progression. During the second 6 week treatment cycle, she developed pain at lower back and bottom accompanied by fever. At physical examination, a peri-anal fistula was found and laboratory results supported systemic inflammation (CRP [ 500 mg/l). Magnetic Resonance Imaging revealed widespread perianal abscesses and fistulas. As extensive surgical resection was no reasonable option, optimal palliative care was provided.
The incidence of sunitinib or sorafenib related GI perforation is unknown, since only few cases were reported in trials [6][7][8][9][10][11][12][13] and case reports [14][15][16][17][18][19][20]. Because of the potential serious outcome, it would be extremely helpful if we could predict patients at-risk on basis of risk factors and underlying biological mechanisms. In addition, more insight in these underlying mechanisms is important to develop potential novel agents with an improved toxicity profile.
Since the first observations of GI perforation during bevacizumab treatment, the risk factors of primary tumor in situ and recent history of endoscopy or abdominal radiotherapy [5,[21][22][23][24][25][26] were described. Pathological findings, frequently associated with observed perforations, include perforation at the tumor or anastomotic site, abdominal carcinomatosis, diverticulitis, GI obstruction and intra-abdominal abcess [5,[22][23][24]. Different biological mechanisms of bevacizumab related perforations have been theorized, which we have outlined in the next part. Whether the same risk factors and mechanisms may be involved in TKI related perforation is unknown, but seems very likely, because both type of agents inhibit VEGF signaling. We have summarized in Table 1 that for both type of agents, gastrointestinal perforations were reported in diverse parts of the gastrointestinal tract. In the reported cases for sunitinib and sorafenib, tumor cells at the site of perforation and previous radiation treatment were frequently mentioned similar to bevacizumab reports [6,[13][14][15][16][17].
Possible mechanisms of GI perforations due to angiogenesis inhibition Tol et al. [27] suggested a relationship between bevacizumab treatment and ulcer development, which may eventually cause a GI perforation. In a phase III study with 755 patients receiving chemotherapy with bevacizumab plus or minus cetuximab, twelve GI perforations were observed of which four were located in an ulcer. The high incidence of ulcers in this study (1.3 vs. 0.1% in the general population), the occurrence of perforations early in treatment, the established role of VEGF in ulcer healing [28][29][30] and the inhibitory effect of bevacizumab on wound healing support their hypothesis. Since the majority of perforations were located at the primary tumor site, preexistent mucosal lesions were expected as preferential localizations. In another report it was speculated that bevacizumab induced VEGF inhibition might result in the cholesterol emboli syndrome (CES), which may consequently give rise to GI perforations due to mesenteric ischemia [31]. Hypertension in combination with eosinophilia is a feature of CES. All three out of twenty-two prospectively observed patients who developed hypertension during bevacizumab treatment had atherosclerotic risk factors, an increased heart rate and eosinophilia at onset of hypertension. In this report it was hypothesized that CES might cause all acute bevacizumab related complications in atherosclerotic patients, including GI perforations as a consequence of mesenteric ischemia.
Alternatively, Saif et al. [22] postulated that GI perforation is caused directly by regression of normal blood vessels in the GI tract, induced by excessive VEGF inhibition. The authors extrapolated data from animal models in which VEGF inhibition has shown to reduce vascular density in the small intestinal villi as well as in other organs [32]. In a recent editorial on the risk of bevacizumab associated GI perforation in ovarian cancer it was speculated that bevacizumab induces necrosis of malignant ovarian cells that invade the bowel serosa resulting in GI perforation [4]. In addition, in this editorial it was suggested that increased pressure due to abdominal carcinomatosis or adhesions from prior surgeries might lead to micro-perforations in vulnerable areas of the bowel, with subsequent delayed healing due to bevacizumab. Finally, loss of nitric oxide (NO) release due to VEGF inhibition, leading to decreased blood flow to the splanchnic vasculature, was proposed to result in bowel infarction and perforation at areas with marginal blood supply.
On account of early closure of the ORBIT trial, evaluating bevacizumab treatment in platinum resistant ovarian cancer, tumor involvement of the bowel was suggested [33]. Five out of 44 patients developed GI perforation and showed radiographic evidence of bowel involvement at study entry. A significant association of GI perforations with increased number of prior chemotherapy regimens (respectively three) and a non-significant relation with bowel wall thickening/obstruction were found. In contrast, in another study with twenty-five heavily pretreated (median of five prior chemotherapy regimens) patients with advanced ovarian cancer, treatment with bevacizumab did not cause any GI perforations [34].
We recently discussed the role of platelets in antiangiogenic treatment related toxicity [35,36]. Platelets contain VEGF in their a-granules which they secrete upon activation and on the other hand VEGF activation of the endothelium results in platelet binding and subsequent activation [37][38][39]. In addition, we found that bevacizumab is taken up by platelets, leading to VEGF neutralization [35]. Since VEGF is an endothelial cell survival factor [2,40,41], we postulated that the subsequent disturbed platelet-endothelial cell interaction might be involved in GI perforation, disturbed wound healing and bleeding complications [36]. The platelet-endothelial cell homeostasis may be disturbed by antiangiogenic treatment. Therefore an increased leakiness and extravasation of inflammatory cells may cause submucosal inflammation and subsequent ulcer formation.
It is of clinical importance to study underlying biological mechanisms of bevacizumab related GI perforation. In addition, it is expected that these underlying mechanisms and risk factors might account for antiangiogenic TKI treatment as well. Risk factors of tumors at the primary site and recent history of endoscopy or abdominal radiotherapy should be taken into account before treatment initiation with angiogenesis inhibitors. In ovarian cancer patients it is recommended to consider the number of prior chemotherapy regimens and abdominal surgeries and to exclude tumor involvement of the bowel by physical examination and CT-scan upon start of treatment with angiogenesis inhibitors. Endoscopic evaluation is advised in patients with symptoms possibly related to GI ulcer during treatment [27]. In addition, based on this report, rubber band ligation should be prevented until bevacizumab or TKI treatment is interrupted or terminated. The third case of GI perforation during combined bevacizumab and TKI treatment emphasizes a possible increased perforation risk related to combination treatment with antiangiogenic agents with different biological mechanisms. Although most of the current preclinical and clinical knowledge on potential underlying mechanisms of angiogenesis inhibitor induced gastrointestinal perforations is on bevacizumab, based on preclinical and clinical studies potential underlying mechanisms as described may hold true for TKIinduced perforations as well.
In conclusion, we would like to advocate to include GI perforation in the differential diagnoses, when patients complain of (vague) abdominal pain during treatment with TKIs as well as with bevacizumab.
Acknowledgments
The authors declare that they have no conflict of interest.
Predation is a major selective force for the evolution of behavioural characteristics of prey. Predation among consumers competing for food is termed intraguild predation (IGP). From the perspective of individual prey, IGP differs from classical predation in the likelihood of occurrence because IG prey is usually more rarely encountered and less profitable because it is more difficult to handle than classical prey. It is not known whether IGP is a sufficiently strong force to evolve interspecific threat sensitivity in antipredation behaviours, as is known from classical predation, and if so whether such behaviours are innate or learned. We examined interspecific threat sensitivity in antipredation in a guild of predatory mite species differing in adaptation to the shared spider mite prey (i.e. Phytoseiulus persimilis, Neoseiulus californicus and Amblyseius andersoni). We first ranked the players in this guild according to the IGP risk posed to each other: A. andersoni was the strongest IG predator; P. persimilis was the weakest. Then, we assessed the influence of relative IGP risk and experience on maternal strategies to reduce offspring IGP risk: A. andersoni was insensitive to IGP risk. Threat sensitivity in oviposition site selection was induced by experience in P. persimilis but occurred independently of experience in N. californicus. Irrespective of experience, P. persimilis laid fewer eggs in choice situations with the high-rather than low-risk IG predator. Our study suggests that, similar to classical predation, IGP may select for sophisticated innate and learned interspecific threat-sensitive antipredation responses. We argue that such responses may promote the coexistence of IG predators and prey.
Predation risk is a major selective force for the evolution of behavioural characteristics of prey (Lima & Dill 1990;Kats & Dill 1998). During their life, most prey species are faced with multiple predator species posing different levels of predation risk (Sih et al. 1998). Additionally, there can be large temporal and spatial variation in predator species composition in a given predator-prey community (Lima & Dill 1990;Lima & Bednekoff © 2011Elsevier Ltd. 1999). Consequently, predation risk is a highly variable component in the life of prey. Prey that overreacts by responding to each predator encounter irrespective of the risk posed by the predator would incur a fitness decrease, because antipredation behaviour is commonly traded off against foraging and/or mating and/or reproduction. Conversely, underestimation of predation risk may have dramatic consequences for prey (its death) or lead to a fitness decrease of prey by lowering its reproductive success. Hence, prey should be able to assess the magnitude of predator threat and adjust its behaviour accordingly (Sih 1982(Sih , 1986;;Helfman 1989). The degree of predation risk may be determined by numerous interrelated but hierarchically structured factors such as predator species, sex or life stage. On top of this hierarchy is predator species recognition, because it allows discrimination of predatory from nonpredatory species and high-risk from low-risk species. Interspecific threat-sensitive prey responses are well documented in both aquatic (Kiesecker et al. 1996;Botham et al. 2008) and terrestrial (Stapley 2003;Edelaar & Wright 2006;Blumstein et al. 2008;Monclus et al. 2009) communities in classical predator-prey interactions.
Predation among species competing for shared resources is called intraguild predation (IGP;Polis et al. 1989). IGP differs from classical predation in various aspects. Like classical predation, intraguild (IG) predators may gain energy by food intake, but unlike classical predation they also gain from eliminating a competitor and potential predator of themselves and/or their offspring. For many predators IG prey has a lower profitability, defined as energy content divided by handling time (e.g. Charnov 1976), than extraguild prey, as measured in predator survival, growth, development or oviposition (ladybirds: Kagata & Katayama 2006;mirid bugs: Provost et al. 2006;predatory mites: Schausberger & Croft 2000a;spiders: Matsumura et al. 2004). Although the nutrient composition of IG prey per se seems favourable for IG predators (Denno & Fagan 2003;Matsumura et al. 2004), subduing, capturing and killing IG prey is energetically costly because IG prey are predators themselves and in mutual IGP may counterattack their predators. In extreme cases, the energetic costs of IGP may even exceed the benefits obtained (predatory mites: Lawson-Balagbo et al. 2008). In addition, true predators (carnivores and omnivores), and consequently potential IG prey, are usually less abundant than classical prey such as herbivores (Begon et al. 1996), reducing the predator-prey encounter frequency and making predator adaptations that help exploit IG prey more efficiently less likely than adaptations that improve exploitation of classical prey. Hence, although omnipresent in natural and artificial food webs (Polis et al. 1989;Arim & Marquet 2004) IGP events should be rarer than classical predation events. It is not known whether IGP is enough of a selective force to evolve interspecific threat-sensitive antipredation behaviours in IG prey, as is common in classical predation.
A precondition for displaying effective threat-sensitive antipredation responses is the ability to recognize a given predator species by direct and/or indirect cues. In general, recognition of predator species can be innate or learned, or a combination of both, ranging from strictly innate recognition unaffected by experience (e.g. tadpoles: Gallie et al. 2001) to innate recognition modifiable by experience (e.g. fish: Hawkins et al. 2008) to recognition only after experience (e.g. rodents: Kindermann et al. 2009). However, most studies on interspecific threat-sensitive antipredation behaviour do not allow discrimination between innate and learned predator recognition because of studying wild animals in the field (e.g. Edelaar & Wright 2006;Blumstein et al. 2008) or using either wild-caught or experienced experimental animals (Kiesecker et al. 1996;Stapley 2003;Botham et al. 2008;Monclus et al. 2009). Unambiguous evidence for learned threat sensitivity is rare and mostly relates to intraspecific threat sensitivity in aquatic or semiaquatic animals such as larval mosquitoes, tadpoles, fishes and water striders in classical predator-prey interactions (Ferrari et al. 2005(Ferrari et al. , 2008;;Ferrari & Chivers 2009;Hirayama & Kasuya 2009). Evidence for learned interspecific threat sensitivity in IGP and terrestrial animals is lacking.
We studied the influence of predation risk and experience on antipredation behaviour within a guild of three predatory mite species: Phytoseiulus persimilis, Neoseiulus californicus and Amblyseius andersoni (Acari: Phytoseiidae). All three species are plant-inhabiting predators of phytophagous mites (e.g. McMurtry & Croft 1997;De Moraes et al. 2004). They naturally co-occur in the Mediterranean basin (De Moraes et al. 2004;A. Walzer, personal observation), presumably share a long coevolutionary history and interact with each other via competition for shared prey, such as the two-spotted spider mite, Tetranychus urticae (Acari: Tetranychidae), and mutual IGP. They differ in diet breadth and adaptation to, and strength in competition for, spider mites. Phytoseiulus persimilis is a highly specialized predator of spider mites and the strongest competitor; N. californicus is a generalist predator with a ranked preference for spider mites and an intermediate competitor while A. andersoni is a generalist predator poorly adapted to utilize spider mites as prey and therefore the weakest competitor for spider mites (McMurtry & Croft 1997). Regarding the propensity to engage in IGP, the ranking is reversed, with A. andersoni being a highly aggressive IG predator, followed by the intermediate N. californicus. Phytoseiulus persimilis is a comparably weak IG predator, only occasionally preying on other predatory mites (Walzer & Schausberger 1999;Schausberger & Croft 2000b). Thus, owing to a joint natural history and presumable differences in the IGP risk posed to each other, the three predatory mites are perfectly suitable animals to test for threat-sensitive anti-IGP behaviours.
As with many other animals (Polis et al. 1989), IGP among phytoseiid mites is mutual but asymmetric with respect to size. Small/younger juveniles are usually preyed upon by larger/ older juveniles and/or adult females, whereas adult females and eggs are relatively invulnerable to IGP (Croft et al. 1996;Schausberger 1997;Walzer & Schausberger 1999;Schausberger & Croft 2000b). The larva is the smallest and least mobile juvenile stage and most in danger of falling victim to larger IG predators (Walzer & Schausberger 1999;Schausberger & Croft 2000b). The IGP risk for larvae may be reduced by antipredation behaviours of the larvae themselves (e.g. Schausberger 2003) or by maternal investment in reducing larval predation risk (Faraji et al. 2001;Walzer et al. 2006). Here, we focused on the latter. Foraging gravid phytoseiid females facing IG predators have several possibilities for decreasing the predation risk of their offspring. They may avoid prey patches with IG predators and choose predator-free patches for oviposition (e.g. Walzer et al. 2006); they may kill potential IG predators of their offspring (e.g. Schausberger & Croft 2000b); or they may reduce/postpone oviposition in the presence of IG predators (e.g. Montserrat et al. 2007;Abad-Moyano et al. 2010a) and resume oviposition when conditions improve.
We investigated the influence of IGP risk and IG predator experience of IG prey on the above-mentioned three maternal strategies to reduce offspring IGP risk. In the first experiment, we assessed the relative risk that P. persimilis, N. californicus and A. andersoni posed to each other in IGP, allowing categorization of each species as a low-or high-risk IG predator of another species. In the second experiment, we scrutinized prey patch choice and oviposition behaviour of IG predator-naïve and predator-experienced females of P. persimilis, N. californicus and A. andersoni when confronted with low-and high-risk IG predators.
Phytoseiulus persimilis, N. californicus and A. andersoni are indigenous, co-occurring species in Sicily (De Moraes et al. 2004). All three species were sampled from herbs and trees in the state of Trapani in 2007. About 20-30 specimens of each species were used to initiate populations reared in the laboratory. Rearing arenas consisted of plastic tiles resting on water-saturated foam cubes in plastic boxes half-filled with water. The edges of the tiles were covered with moist tissue paper to confine the predators to the rearing arenas. Cotton wool fibres under coverslips served as shelter and oviposition sites for A. andersoni. To prevent contamination of the predator populations, an adhesive (Raupenleim, Avenarius Agro, Wels, Austria) was applied to the rim of the plastic boxes and the boxes were placed in a tray containing water with dishwashing detergent. The predators were fed in 2-3-day intervals with T. urticae, reared on whole bean plants (Phaseolus vulgaris), by adding bean leaves infested with spider mites (for P. persimilis, N. californicus) or by brushing spider mites from infested leaves (for A. andersoni) onto arenas.
In the first experiment, we assessed the IGP risk posed by adult females to larvae to categorize each species as a low-or high-risk predator of another species (Schausberger & Croft 2000b). We measured the attack probability and attack latency of IG predator females on IG prey larvae confined in closed acrylic cages. Each cage consisted of a cylindrical cell (15 mm in diameter and 3 mm in height) with a fine-mesh screen at the bottom and closed on the upper side with a microscope slide (Schausberger 1997). Treatments were all possible combinations of P. persimilis, A. andersoni and N. californicus as IG predators and prey. Gravid females randomly taken from the rearing units were singly placed into closed cages and starved for 12 h. Subsequently, a single heterospecific larva was added to each cage. Each female and each larva was used only once. Each cage was checked for the first successful attack of the predator female on larval prey (killed and sucked out) every 10-15 min for 6 h at the longest. Each treatment (predator-prey combination) was replicated 19-20 times.
In the second experiment, we assessed prey patch choice and oviposition behaviour of IG predator-naïve and predator-experienced females of P. persimilis, N. californicus and A. andersoni given a choice between a spider mite patch with and without cues (traces of foraging females and their eggs) from a low-or high-risk IG predator. We used a full 3 (IG prey female species) × 2 (naïve and experienced) × 2 (cues of high-and low-risk IG predator) factorial design. Each choice situation was replicated 20-31 times.
To obtain naïve and experienced prey females for experiments, three groups of 13-15 eggs each of A. andersoni, N. californicus and P. persimilis were randomly taken from the rearing units and placed on separate leaf arenas for development. Each leaf arena consisted of a detached bean leaf (5 × 5 cm) placed upside down on a water-saturated foam cube in a plastic box half-filled with water. Water-saturated cellulose strips (1 cm height) at the edge of the leaf confined the arena and prevented the mites from escaping. For each species, one group of eggs was placed on a leaf with only spider mites (to be used as naïve IG prey females in experiments), and the second and third groups of eggs were placed on separate leaves with spider mites and either five low-or high-risk IG predator females, respectively (to be used as experienced IG prey females in experiments). The developmental progress of IG prey was observed and spider mite prey replenished daily to exclude food competition. The IG predator females mainly preyed on spider mites but also killed a few IG prey individuals and produced eggs. Mortality of IG prey was about 10-15% higher in rearing units with IG predators than without, but the set-up provided for random IGP (A. Walzer & P. Schausberger, unpublished data). Therefore, during development IG prey to be used as experienced females in experiments were exposed to all direct (predators, their eggs, chemical footprints and/or metabolic waste products) and indirect (killed conspecifics and spider mites) volatile and/or tactile chemosensory cues possibly indicating IG predator
Walzer and Schausberger Page 4 Published as: Anim Behav. 2011 January ; 81(1): 177-184.
Sponsored Document Sponsored Document presence (Walzer et al. 2006;Montserrat et al. 2007). Gravid IG prey females appeared after 10 days and were then subjected to choice experiments.
Each experimental choice unit consisted of two similarly sized leaflets (5-7 cm 2 ) taken from trifoliate bean leaves placed upside down on a foam cube in a plastic box (15 × 10 cm and 4 cm high) half-filled with water. The leaflets were connected by a wax bridge (1 × 0.5 cm; Walzer et al. 2006). One leaflet only harboured spider mites and the other harboured spider mites plus cues (traces of foraging females, such as metabolic waste products and/or chemical footprints, and their eggs) of either the low-or the high-risk IG predator. The setup was designed to simulate a natural scenario in which a gravid female searches for a prey patch and oviposition site among leaves on a branch within a plant. During preparation of the prey patches before the choice experiment took place, we blocked the bridge with a strip of moist tissue paper. Prey patches were created by placing 30 juvenile and four to seven adult T. urticae females on each leaflet. After 24 h either no or five low-risk or five high-risk IG predator females were added and allowed to produce eggs. After a further 24 h the T. urticae females and the IG predator females were removed and spider mite densities (juveniles and eggs) adjusted, by adding or removing individuals using a brush, to identical predetermined levels on leaflets with and without IG predator cues. IG predator eggs were reduced to five per leaflet and their position was marked by a tiny watercolour dot on the leaf surface to ease identification of eggs produced by the experimental IG prey females. To account for species-specific prey stage preferences (eggs versus mobile prey; Blackwood et al. 2001) and prey needs (Vanas et al. 2006) we adjusted the density and composition of each spider mite patch to 60 eggs and 20 juveniles for P. persimilis, 30 eggs and 10 juveniles for N. californicus and 30 eggs and 30 juveniles for A. andersoni. Each patch allowed the IG prey female to reach the maximum oviposition rate during the 24 h experimental period and leave enough prey for her offspring to reach adulthood (Vanas et al. 2006). Therefore, differences in prey patch choice and/or oviposition behaviour of IG prey females can be delimited to the presence or absence of IG predator cues. Through the abovementioned pre-experimental procedure, the presence of an IG predator in a given spider mite patch was indicated by traces left by the foraging and ovipositing IG predator (e.g. metabolic waste products and/or chemical footprints), eggs laid and killed spider mites (Grostal & Dicke 1999).
Before we ran the choice experiments, each experimental IG prey female was singly placed into a closed acrylic cage (described above) and starved for 12 h. Only females producing at least one egg during the starvation period were used for experiments. Subsequently, each female was singly released in the middle of the wax bridge and given a choice between a prey patch with only spider mites and a prey patch with spider mites and cues of either a low-risk or a high-risk IG predator (determined for each species in experiment 1). Each choice unit and each IG prey female was used only once. The position of the IG prey female was checked eight times during the experiment (immediately after release (first choice), and then after 1, 2, 3, 4, 5, 6 and 24 h). After 24 h we recorded eggs deposited in each prey patch and predation by the IG prey female on IG predator eggs.
Statistical analyses were carried out for each species separately using SPSS 15.0.1 (SPSS, Chicago, IL, U.S.A.). In experiment 1, larval survival functions (combination of cumulative survival and survival time), used as indicators of the relative IGP risk posed by the other two species, were analysed by the Kaplan-Meier procedure and pairwise Breslow tests (Bühl 2008). In experiment 2, the influence of experience (naïve versus experienced) and predation risk (low versus high risk) on the residence frequency of IG prey females in prey patches with IG predator cues during the experiment (presence in the prey patch with IG predator cues out of eight observation points) was analysed using generalized linear models (GLM; counts of events; binomial distribution with logit link function; Bühl 2008). Within each choice situation, the numbers of eggs laid in the prey patch with and without predator cues were compared by Wilcoxon signed-ranks tests owing to non-normality of data. The influence of experience and predation risk on total egg production (eggs laid in both prey patches combined; log-transformed before analysis to meet variance homogeneity and improve normality) and the number of IG eggs preyed upon by the IG prey females were analysed using two-factorial (experience and predation risk) ANOVAs. The proportion of eggs laid by the IG prey females in prey patches with IG predator cues (eggs in patch with IG predator cues out of total egg production; females producing no eggs were removed from the analysis) were compared among treatments using GLMs (counts of events; binomial distribution with logit link function; Bühl 2008).
For each species, the survival functions of larvae differed significantly between the two IG predator species (Fig. 1; Breslow tests within Kaplan-Meier analyses). Based on these differences we categorized the two IG predators of a given species as low-and high-risk predators: for A. andersoni, low-risk P. persimilis and high-risk N. californicus (chi-square test: χ 1 2 = 4.507, P = 0.034); for N. californicus, low-risk P. persimilis and high-risk A. andersoni (χ 1 2 = 9.743, P = 0.002); for P. persimilis, low-risk N. californicus and high-risk A. andersoni (χ 1 2 = 13.875, P < 0.001).
Amblyseius andersoni-Prey patch choice by A. andersoni females was influenced by experience but not by predation risk (Table 1). Experienced females resided in prey patches with IG predator cues less often than naïve females (Fig. 2).
Within choice situations, experienced females confronted with cues of the high-risk predator deposited more eggs in the patch with predator cues (Wilcoxon signed-ranks exact test: Z = -2.822, N = 28, P = 0.005). In the other choice situations the numbers of deposited eggs did not differ between patches with and without predator cues (P > 0.05; Fig. 3). Total egg production was not affected by experience, predation risk or the interaction of the two sources of variation (ANOVA experience: F 1,96 = 0.159, P = 0.691; predation risk: F 1,96 = 0.353, P = 0.554; experience*predation risk: F 1,96 = 0.889, P = 0.348). The proportion of eggs laid in the prey patch with IG predator cues was marginally influenced by experience (GLM: Wald χ 1 2 = 3.446, P = 0.063) but not by predation risk (Wald χ 1 2 = 0.001, P = 0.995) or the interaction between experience and predation risk (Wald χ 1 2 = 1.057, P = 0.304). Experienced females laid a slightly higher proportion of eggs in prey patches with IG predator cues than naïve females (Fig. 3).
The IGP rates of A. andersoni were not influenced by experience (ANOVA: F 1,96 = 1.999, P = 0.161) and predation risk (F 1,96 = 0.317, P = 0.575) as main effects but were influenced by the interaction of the two sources of variation (F 1,96 = 3.955, P = 0.050). Naïve females killed similar numbers of eggs of the low-and high-risk IG predator (mean ± SE = 0.9 ± 0.3, N = 20 versus 1.2 ± 0.3, N = 25), whereas experienced females killed more eggs of the lowthan of the high-risk IG predator (1.0 ± 0.2, N = 28 versus 0.4 ± 0.1, N = 28).
Neoseiulus californicus-The influence of experience on prey patch choice by N. californicus females was dependent on predation risk (Table 1, Fig. 2). The females were similarly distributed between prey patches with or without IG predator cues, except experienced females subjected to the choice situation with the high-risk IG predator A. andersoni, which were more often found in the prey patch with only spider mites (Table 1, Fig. 2).
Within choice situations, experienced females confronted with cues of the high-risk predator deposited more eggs in the prey patch with only spider mites (Wilcoxon signed-ranks exact test; Z = -2.658, N = 31, P = 0.008). In the other choice situations, the numbers of deposited eggs did not differ between patches with and without predator cues (P > 0.05; Fig. 3). Total egg production was unaffected by experience and predation risk (ANOVA: experience: F 1,110 = 0.112, P = 0.738; predation risk: F 1,110 = 0.349, P = 0.556; experience*predation risk: F 1,110 = 0.253, P = 0.616; Fig. 3). Irrespective of experience, the proportion of eggs laid in prey patches with predator cues was higher in choice situations with the low-risk IG predator than in those with the high-risk IG predator (GLMs: experience: Wald χ 1 2 = 0.121, P = 0.728; predation risk: Wald χ 1 2 = 8.928, P = 0.003; experience*predation risk: Wald χ 1 2 = 2.710, P = 0.100; Fig. 3).
The IGP rates of N. californicus were not influenced by experience (ANOVA: F 1,111 = 0.002, P = 0.951) and predation risk (F 1,111 = 0.612, P = 0.330) as main factors but were influenced by the interaction of these two sources of variation (F 1,111 = 3.010, P = 0.032). Irrespective of predation risk, experienced females killed similar numbers of eggs of the low-and high-risk IG predator (mean ± SE = 0.6 ± 0.2, N = 27 versus 0.5 ± 0.1, N = 30), whereas naïve females killed more eggs of the high-than low-risk IG predator (0.8 ± 0.2, N = 27 versus 0.3 ± 0.1, N = 31).
Phytoseiulus persimilis-The influence of predation risk on prey patch choice by P. persimilis females was dependent on experience (Table 1, Fig. 2). Experienced females resided more often in the predator-free prey patch in choice situations with the high-risk IG predator compared with the other choice situations (Fig. 2).
Within each choice situation P. persimilis deposited more eggs in the prey patch without predator cues than in the prey patch with predator cues (Wilcoxon signed-ranks exact test: naïve female/harmless predator: Z = -3.106, N = 26, P = 0.002; naïve female/harmful predator: Z = -3.382, N = 22, P = 0.001; experienced female/harmless predator: Z = -2.519, N = 27, P = 0.012; experienced female/harmful predator: Z = -4.193, N = 23, P < 0.001; Fig. 3). Total egg production was affected by predation risk but not by experience (ANOVA: experience: F 1,94 = 0.354, P = 0.533; predation risk: F 1,94 = 7.533, P = 0.007; experience*predation risk: F 1,94 = 0.045, P = 0.833). Irrespective of experience, P. persimilis females produced fewer eggs in choice situations with the high-risk IG predator than in choice situations with the low-risk IG predator (Fig. 3). Predation risk (GLM: Wald χ 1 2 = 7.377, P = 0.007) but not experience (GLM: Wald χ 1 2 = 2.647, P = 0.104) influenced the proportion of eggs laid by P. persimilis in prey patches with IG predator cues. However, the significant interaction between predation risk and experience (GLM: Wald χ 1 2 = 6.554, P = 0.010) indicates that experience decreased the proportion of eggs laid in the patch with predator cues in choice situations with the high-risk IG predator but not in choice situations with the low-risk IG predator (Fig. 3).
The IGP rates of P. persimilis (mean ± SE = 0.11 ± 0.06, N = 26 versus 0.12 ± 0.06, N = 29 and 0.04 ± 0.04, N = 22 versus 0.09 ± 0.06, N = 23 eggs of the high-and low-risk IG predator for naïve and experienced females, respectively) were unaffected by experience and predation risk (ANOVA: experience: F 1,94 = 0.599, P = 0.441; predation risk: F 1,94 = 0.189, P = 0.665; experience*predation risk: F 1,94 = 0.132, P = 0.718).
Published as: Anim Behav. 2011 January ; 81(1): 177-184.
Sponsored Document Sponsored Document
Our study shows that IGP is a sufficiently strong force to select for interspecific threat sensitivity in antipredation behaviours. It documents innate and learned interspecific threatsensitive anti-IGP responses and provides experimental evidence for learned threat-sensitive antipredation in a strictly terrestrial animal (e.g. Kats & Dill 1998). We observed three maternal strategies to reduce predation risk of offspring. These strategies are common in classical predator-prey interactions but are less well documented for IGP, that is, killing predators (for classical predation
: Saito 1986; for IGP: Walzer & Schausberger 2009), decreasing and postponing oviposition (for classical predation: Skaloudova et al. 2007; for IGP: Montserrat et al. 2007) and oviposition site avoidance (for classical predation: Grostal & Dicke 1999; Murphy 2003; for IGP: Walzer et al. 2006). However, the strategies adopted, threat sensitivity and the influence of predator experience differed between species. First, IG prey females of all three species killed IG predator eggs. IG egg predation was most pronounced in A. andersoni and negligible in P. persimilis. Only naïve N. californicus females behaved in a threat-sensitive manner in IG egg predation. Second, both naïve and experienced P. persimilis females laid fewer eggs in the presence of the high-risk IG predator than in the presence of the low-risk IG predator, indicating innate threat sensitivity. Egg production by A. andersoni and N. californicus was unaffected by predation risk and experience. Third, both N. californicus and P. persimilis were threat sensitive in oviposition site selection. Both laid a lower proportion of eggs in prey patches with the high-risk predator than in those with the low-risk IG predator. However, threat sensitivity in oviposition site selection was induced by experience in P. persimilis but not in N. californicus.
Amblyseius andersoni females were threat insensitive in prey patch selection and oviposition behaviour. A possible interpretation is that A. andersoni is not able to discriminate between low-and high-risk IG predators. More likely, the lack of threat sensitivity in A. andersoni was specific to the IGP risks posed by P. persimilis and N. californicus, respectively. The risks posed by P. persimilis and N. californicus to A. andersoni were lower and their difference smaller than those in the other IG predator-prey combinations. Therefore, it could be that the overall predation risk was too low and the difference between risks too small to trigger a threat-sensitive response in A. andersoni. Alternatively or additionally, experiment 1 indicates that larvae of A. andersoni are better able to escape from or defend themselves against IGP than are the larvae of N. californicus and P. persimilis (see also Zhang & Croft 1995). Experienced A. andersoni females preferred to deposit their eggs in prey patches with cues of the high-risk IG predator. At first glance such behaviour seems maladaptive. However, IG predator eggs are not only potential future predators and competitors of offspring but also an alternative prey for A. andersoni. Phytoseiid mites are a higher quality prey for A. andersoni than are T. urticae, allowing rapid juvenile development and sustained oviposition (Schausberger & Croft 2000a). Therefore, experience may have enhanced acceptance and utilization of IG predator eggs as food by A. andersoni females. Higher predation by experienced females on eggs of P. persimilis than N. californicus may be explained by differing nutritional quantity of single eggs. Eggs of P. persimilis are about one-third larger than those of N. californicus (Croft et al. 1999). Whether the eggs of N. californicus and P. persimilis also differ in nutritional quality is unknown.
Both N. californicus and P. persimilis females responded in a threat-sensitive manner in prey patch selection and oviposition behaviour. However, the relative contribution of innate and learned components to threat sensitivity differed between the two species. Only N. californicus experienced with the high-risk IG predator avoided residence in the patch with IG predator cues. By contrast, both naïve and experienced P. persimilis females avoided patches with IG predator cues, but only experienced females were threat sensitive in patch selection. In P. persimilis, oviposition site selection matched prey patch selection. Irrespective of predation risk and experience, P. persimilis laid more eggs in predator-free patches, but experience induced threat sensitivity and fine-tuned oviposition site selection.
To our knowledge, only one previous study dealt with threat-sensitive oviposition site selection influenced by experience in a classical predator-prey interaction. Water striders, Aquarius paludum insularis learned to adjust the oviposition depth to the risk of egg parasitism (Hirayama & Kasuya 2009). The total number of eggs laid by P. persimilis was lower in choice situations with the high-risk IG predator than in those with the low-risk IG predator. Proximate explanations for decreased egg production in the presence of a high-risk IG predator include reduced feeding and/or egg retention (Montserrat et al. 2007;Abad-Moyano et al. 2010a). In N. californicus, oviposition site selection and prey patch selection did not exactly match. Experience induced threat sensitivity in prey patch selection, whereas threat sensitivity in oviposition site selection was innate and unmodified by experience. A likely explanation is that offspring are much more vulnerable to IGP than are the adult females themselves (Croft et al. 1996), rendering oviposition site selection a much stronger selective force than prey patch selection.
The species-specific strategies to reduce IGP on offspring and their complexities and magnitudes reflect the species ranking in mutual predation (experiment 1; Schausberger & Croft 2000b). They further indicate that IGP, which contains elements of both predation and competition, and not merely competition for spider mites, has been the selective force driving the evolution of these behaviours. If competition for spider mites was the driving force, the weakest competitor A. andersoni would have had the strongest response, and the strongest competitor P. persimilis the least, respectively, to the presence of the other species. However, the opposite was true. The lack of threat-sensitive behaviours in A. andersoni may indicate that their juveniles are the least endangered IG prey within the guild studied. Moreover, killing potential IG predator eggs by A. andersoni IG prey females both yields nutritional benefits and relaxes offspring predation and competition risks. Cues from killed conspecifics may prevent further IG predators from entering prey patches, as for example shown for the predatory mite Iphiseius degenerans (Faraji et al. 2001). Neoseiulus californicus females were threat sensitive in IGP and oviposition site selection. However, the latter was only evident in prey patches with the high-risk IG predator A. andersoni, indicating that P. persimilis represents a negligible risk for N. californicus offspring.
Phytoseiulus persimilis is the most vulnerable to IGP within the guild studied. The females responded to both IG predators under all circumstances, indicating that both predators constitute a considerable risk for their offspring. Experience drastically increased threat sensitivity of P. persimilis, which was reflected in almost complete oviposition avoidance in prey patches with cues of the high-risk IG predator (only two of 51 eggs were placed in the patch with IG predator cues).
Threat sensitivity in IGP may have important implications for population and community dynamics. For example, graded oviposition avoidance may result in graded direct traitmediated IG interactions (Luttbeg & Kerby 2005;Preisser et al. 2005;Abad-Moyano et al. 2010b). Oviposition site selection affects spatial distribution of IG prey independent of the level of IGP risk, which in turn may trigger indirect trait-mediated effects on the shared prey (e.g. Werner & Peacor 2003;Walzer et al. 2009). The ability to perform threat-sensitive anti-IGP behaviours should have stabilizing effects on the coexistence of IG predators and prey. Currently, the general theoretical prediction deduced from IGP models is that coexistence of IG predators and prey within communities at ecological timescales is unlikely at high productivity levels (Janssen et al. 2006), but there is sparse empirical support for this assumption (e.g. Amarasekare 2008). Threat sensitivity in prey patch and oviposition site selection could be a mechanism promoting the coexistence of IG predators and prey (Heithaus 2001;Amarasekare 2008). Walzer and Schausberger Page 13 Published as: Anim Behav. 2011 January ; 81(1): 177-184. Sponsored Document Sponsored Document Sponsored Document Figure 2. The influence of experience and IGP risk on residence of (a) A. andersoni, (b) N. californicus and (c) P. persimilis females in the prey patches with spider mites and IG predator cues (eggs and traces such as metabolic waste products and/or chemical footprints of an IG predator female; mean proportion ± SE; calculated from eight observations during 24 h). Each female was given a choice between a prey patch with only spider mites and a prey patch with cues of a low-risk (white bars) or high-risk (black bars) IG predator. Walzer and Schausberger Page 14 Published as: Anim Behav. 2011 January ; 81(1): 177-184. Sponsored Document Sponsored Document Sponsored Document Figure 3. Total egg production (both patches combined) and proportion of eggs (mean ± SE) laid by (a, b) A. andersoni, (c, d) N. californicus and (e, f) P. persimilis in the prey patches with spider mites and IG predator cues (eggs and traces such as metabolic waste products and/or chemical footprints of an IG predator female). Each female was given a choice between a prey patch with only spider mites and a prey patch with spider mites and cues of a low-risk (white bars) or high-risk (black bars) IG predator. Horizontal lines indicate random choice for proportion of eggs laid (b, d, f).
Walzer and Schausberger Page 15
Published as: Anim Behav. 2011 January ; 81(1): 177-184.
Published as: Anim Behav. 2011 January ; 81(1): 177-184.
Published as: Anim Behav. 2011 January ; 81(1): 177-184.Sponsored DocumentSponsored Document Sponsored Document
Published as: Anim Behav. 2011 January ; 81(1): 177-184.
This work was funded by the
Published as:
Sponsored Document Sponsored Document
Several recent studies have documented that non-human primates can individuate objects according to property and/or kind information in much the same way as human infants do from around one year of age when they begin to acquire language. Some studies suggest, however, that only some properties are used for the individuation of food items: color, but not shape. The present study investigated whether these findings reveal a true competence problem with shape properties in the food domain or whether they merely reveal a performance problem (e.g., lack of attention to shapes). We tested 25 great apes (chimpanzees, bonobos and gorillas) in two food individuation tasks. We manipulated subjects' experience with differences in color and shape properties of food items. Results indicated (i) that all subjects, regardless of their prior experience, solved the color-based object individuation task and (ii) that only the group with previous experience with different shape properties succeeded in the shape-based individuation task. Great apes can thus be primed to take shape into account for individuating food objects, and this results clearly speaks in favor of a performance (rather than a competence) problem in using shape for object individuation of food items.
Human infants' object cognition has been shown to undergo a developmental shift around the first birthday: while from very early on, infants are capable of tracking objects according to spatiotemporal criteria, only from around 10 to 12 months do they become able to track and individuate objects according to property and/or kind information (e.g., Xu and Carey 1996; see Xu 2007 for a review). As this ability has been found to correlate with natural language comprehension (Xu and Carey 1996) and to reveal itself in linguistically supported contexts specifically (Xu 2002;Xu et al. 2005), one hypothesis is that language is necessary for the development of this very ability (Xu 2002).
Work with non-human primates (hereafter primates), however, puts that bold hypothesis into question. Rhesus monkeys and great apes have been found to individuate objects according to their property/kind much in the same way as human infants from around 1 year old do (Uller et al. 1997;Santos et al. 2002;Phillips and Santos 2007;Mendes et al. 2008): When they see an object with property X or of kind A go into an empty box and then find an object with property Y or of kind B (unexpected), they search longer than when they find the original object (expected).
However, mixed results have been found so far concerning the question of which types of properties primates use for individuating objects in different domains. In the domain of food items, the only one which has been used in object individuation studies so far, primates seem to spontaneously use color differences to individuate objects of the same kind, but not shape differences (Santos et al. 2002). That is, when they see, for example, a white food item disappear in an empty box and then find a blue one instead, they continue searching. However, if they see a round food item, they do not respond differently upon finding a triangle one than upon finding the original round one (Santos et al. 2002).
Similar behavior patterns are found in induction tasks in which primate subjects have to decide on the edibility of novel food items (for an overview, see Hauser and Spelke 2004). In one study, for example, an experimenter (E) ate, in full view of rhesus monkeys (Macaca mulatta), a food item which was new to the subjects. E then put down two food items, one identical in color but not in shape to the one E ate and another one different in color but identical in shape. Monkeys approached significantly more often the food item identical in color but different in shape compared to the item differing in color but identical in shape to the one E ate (Santos et al. 2001; see also Shutts et al. 2009).
This prevalence in the use of color properties over shape properties in object individuation and induction in the food domain stands in contrast to findings from the domain of tool use. Here, primates have been found to rely on shape properties (functionally relevant) and neglect color (functionally irrelevant) when having to choose among tools which are differentially appropriate for a given problem (e.g. Hauser et al. 2002;Santos et al. 2003;Santos et al. 2006; see also Furlong et al. 2008 andBania et al. 2009 regarding the choice of tools with functional relevant shape properties, over non-functional ones, by chimpanzees).
This raises the question of how the restriction to certain properties, namely color, in primates' individuation of food objects is to be explained, in particular in light of reverse patterns in tool induction tasks. Does this restriction reflect a true competence problem? One way such a competence problem might arise would be that primates' object individuation operates domain specifically and the domainspecific ability is wired such that shape is not a property that enters the picture in the food domain (whereas it does in the domain of tools). Alternatively, the restriction to certain properties found so far might reflect merely some kind of performance problem such that primates can use shape properties to individuate objects but do not spontaneously do so. One possibility along the latter lines would be that shape is not salient enough for primates in the food domain, but can be used for object individuation if made more salient.
To test between such competence and performance accounts, two groups of great apes were studied, with the amount of experience with shape and color properties and thus their salience being experimentally manipulated between groups. One group (the ''priming group'') was primed to attend to three different shapes and colors of food items belonging to the same kind (food pellets). The other group (the ''naı¨ve group'') remained naı ¨ve with regard to shapes and colors other than the ''regular'' ones (i.e., the regular color (brown) and shape (cylinder shaped) of pellets). Both groups then performed object individuation tasks similar to the ones previously used (Mendes et al. 2008;Santos et al. 2002), and their performance in colorbased and shape-based individuation of food items was compared.
Sixteen chimpanzees (Pan troglodytes), five bonobos (Pan paniscus) and four gorillas (Gorilla gorilla) participated in the present study (see Table 1). There were nine males and 16 females. The average age of all the males was 12 years and 4 months, and the average age of all the females was 17 years and 5 months. All subjects were socially housed at the Wolfgang Koehler Primate Research Center, Leipzig Zoo, Germany. Subjects had access to both indoor and outdoor enclosures. All enclosures were furnished with vegetation, climbing structures and visual barriers. Subjects were neither food-nor water-deprived, and they could stop participating in the task at any given moment. All subjects had previously participated in a first study investigating the use of spatiotemporal and property/kind information during an object individuation task (Mendes et al. 2008).
An opaque plastic box (40 9 40 9 34.5 cm) was used during the experiment (see Fig. 1). The box had a circular opening (approx. 8.5 cm in diameter) on its top middle part which was used by the experimenter (E) to introduce a food item (pellet). The frontal part of the box (facing the subjects) had an opening (13 cm wide 9 6 cm high) which the subjects could use to reach for the pellet. Such opening was covered with a curtain (to avoid subjects looking inside it)
and could be closed using a sliding door manipulated by E.
In order to facilitate the surreptitious introduction of the ''missing'' pellet inside the box (during unexpected trials), a horizontal sliding door was constructed on a false roof of the box (9 cm high from its top part). Subjects were not aware of that. A soft carpet was added to the floor of the box. The carpet prevented the subjects to use possible auditory cues that may have emanated from the fall of the pellets on the floor of the box. The pellets used in the current study were much harder than the food items used in Mendes et al. (2008) and therefore produced a louder sound that this way could be minimized.
In addition to the box, a plastic table (78 9 35 cm) was also used but only during the priming phase. The table was attached to a mesh window in the subjects testing room and was used to place different colored or shaped pellets on its top part. Procedure Testing was done by the same E as in our previous study (Mendes et al. 2008), and the procedure was also very similar to the one used in that study. A cameraperson helped E by timing the trials and informed E when the subject had retrieved the pellet from inside the box. The experiment comprised three phases which were always administered in the following order: Priming phase Subjects were divided into two groups. The priming group (N = 13) was exposed to three different colors and shapes of pellets. Previous to this study, subjects had never experienced pellets different from the ''regular'' ones. In contrast, the naı¨ve group (N = 12) remained naı ¨ve toward colors and shapes other than the regular ones.
During the priming phase, E sat behind a table which was attached to a mesh window in the subjects' room. Subjects received two blocks of trials; one block per each condition (color and shape). The order of presentation of the blocks was counterbalanced across subjects. Each block contained three trials in which the pellets differed within the same property (color or shape). In each trial, the three pellets were placed on the table, aligned in a row and handed one by one to the subject by E. The order of presentation of the pellets on the table was counterbalanced across the three trials of the same block. Each block was presented 24 h previous to the corresponding testing phase. That is, subjects were exposed to a color priming 24 h before the testing phase with different colored pellets and analogously for shape priming and the corresponding testing phase. On the day of the testing phase, but previous to its start, the priming group received one more priming trial with the same property as the one of the test condition to be next presented.
The colored pellets were red, blue and brownish (i.e., ''regular'') ones. The red and blue colors were obtained by adding an edible non-flavored food coloring to the ''regular'' pellets. The shaped pellets had a form of a star, moon and cylinder (i.e. ''regular''), all with equal volume. The star and moon-shaped pellets were made from crushed pieces of ''regular'' pellets, misted with water, molded into the cookie cutters and finally dried.
On the same day, the testing phase began, but immediately before its start, subjects could explore the new the box over a 40-s period. In the case of the priming group, subjects received the familiarization phase immediately after the last priming trial.
Both priming and naı¨ve groups received two conditions, a color and a shape conditions. Each condition contained four trials, two expected and two unexpected trials. The procedure was identical to the one described previously (Mendes et al. 2008-Experiment 2). In the expected trials, subjects saw E introducing pellet A inside the box, and when allowed to reach, they found pellet A. In the unexpected trials, the box was initially baited with pellet B (subjects were unaware of this manipulation) and they saw E introducing pellet A. However, pellet A was surreptitiously stored on the horizontal sliding door of the false roof of the box. Once the sliding door was opened, subjects found pellet B, different in properties but not in kind from the one they saw being hidden.
The temporal structure in both expected and unexpected trials was as follows (see Fig. 2 for a schematic illustration of the applied procedure):
First reach period. Once the subjects had found the pellet, the sliding door of the box was closed immediately for a 20-s period (RP1).
Intermediate period After the first reach period was over, the sliding door was closed and the horizontal sliding door was simultaneously opened so that the ''missing'' reward (in the unexpected trials) fell near the front corners of the box. Thus, creating the impression that the reward had always been there and that subjects had not found it before because of its difficult location, not because E had manipulated it. While E closed and opened the sliding door, simultaneously with the opening of the trap, she spoke loudly (i.e., ''Look at that!'') to prevent subjects from using auditory cues that may have emanated from the fall of the ''missing'' pellet (applied in both expected and unexpected trials). If subjects did not reach inside the box or if they failed to find the ''missing'' pellet in the unexpected trials, the trial was ended after 20 s.
Second reach period If subjects retrieved the ''missing'' reward, the door was immediately closed and re-opened again for a last 20-s reach period (RP2).
The order of presentation of each condition (i.e., color or shape) was counterbalanced within each group. Within each condition, the order of presentation of the first trial was randomized. Expected and unexpected trials did not occur twice in a row.
All videos were digitalized and an observer coded them using the Interact Ò software (version 7). Following the previous study (Mendes et al. 2008), there were two dependent measures: (i) frequency of reaches inside the box and (ii) duration of reaches inside the box. Regarding frequency and duration of reaches, the mean values over the two trials of the same type (expected and unexpected) in each condition (shape and color) were computed. Therefore, each subject had two mean frequency and two mean duration values (corresponding to RP1 and RP2) for expected and unexpected trials in the shape and color conditions. A second observer, blind to the hypotheses of the study, scored a random sample of 20% of the trials. Inter-observer reliability for frequency was high (Pearson correlation r = 0.985, P \ 0.001, N = 116 and weighted Kappa = 0.910) as well as for duration of reaches (Pearson correlation r = 0.997, P \ 0.001, N = 116).
As the data failed to fulfill the requirement for parametric testing, non-parametric tests were used in all analyses (Kruskal-Wallis, Wilcoxon signed rank test, exact Wilcoxon test for N B 15). We used one-tailed tests because we had clear predictions, i.e., more and longer reaching in unexpected than in expected trials and within unexpected trials more and longer reaching in RP1 than in RP2. Regarding the frequency and duration of reaches in RP1 and RP2 of expected trials, no differences were expected.
First, species and age effect were tested for by using the scores obtained from the difference between (i) expected and unexpected trials within the first reach period (RP1); (ii) expected and unexpected trials within the second reach period (RP2), both for shape and color conditions. No difference in performance between species or age classes (infant: 0-5 year old; juvenile: 5-8 year old; subadult: 8-11 year old; adult [ 11) was found for any of the scores aforementioned (Kruskal-Wallis test, all P [ 0.110). An exception was a species difference in the shape condition regarding the duration and frequency of reaches both between expected and unexpected trials during the first reach period (RP1) (Kruskal-Wallis test: duration, H = 6.38, P = 0.034; frequency, H = 5.03, P = 0.075). Because significant P values in each dependent variable might be spurious, Fisher's omnibus test was computed (Haccou and Meelis 1994). The difference in duration and frequency of reaches between expected and unexpected trials during RP1 revealed no significant difference between species (Fisher's omnibus test, duration: v 2 = 8.67, df = 6, P = 0.19; frequency: v 2 = 8.02, df = 6, P = 0.24), thus supporting the view of a spurious significance. Data were, therefore, collapsed across species and age classes.
The main analyses tested the effect of the priming phase on the mean frequency and duration of reaches during the first and second reach periods (RP1 and RP2).
If apes individuate object according to shape, two patterns would be expected: (a) in the first reaching period (RP1), subject should search more often and longer in the unexpected than in the expected trials. (b) In the unexpected trials, subjects search longer and more often in RP1 (searching for a missing pellet) than in RP2 (after having found that pellet). Regarding (a), an analysis on the whole sample revealed that during the first reach period (RP1), the priming group reached significantly more often and for longer time during unexpected compared with expected trials (Wilcoxon signed rank test: frequency, T ? = 8.5, N = 12, P = 0.005; duration, T ? = 11, N = 12, P = 0.013; Fig. 3). In contrast, during RP1, the naı¨ve group did not show any significant differences in frequency or duration of reaching in expected compared with unexpected trials (frequency, T ? = 17, N = 9, P = 0.287; duration, T ? = 25, N = 11, P = 0.260; Fig. 3).
Regarding (b), the corresponding analyses could only be run with subjects who participated in RP2. In the intermediate period, in each group, not all subjects retrieved the ''missing'' pellet in at least one of the two unexpected trials (N priming group = 8 and N naı¨ve group = 8). Thus, differences in performance across both reaching periods (RP1 and RP2) was analyzed for this sub-sample only (the same subsample analysis has been previously described elsewhere; see Mendes et al. 2008). During unexpected trials, the priming group reached significantly more often and longer during the first (RP1) compared with the second reach periods (RP2) (frequency: T ? = 3, N = 8, P = 0.023; duration: T ? = 4, N = 8, P = 0.027; Fig. 4). However, during the expected trials, the group reached equally often and equally long during RP1 compared with RP2 (frequency: T ? = 10, N = 7, P = 0.625, two-tailed; duration: T ? = 10, N = 8, P = 0.313, two-tailed; Fig. 4).
In contrast to the performance of the priming group, during unexpected trials, the naı¨ve group reached equally often during the first (RP1) compared with the second Fig. 3 Mean average (?SE) frequency a and duration b of reaches during the first 20-s reach period (RP1) in both priming and naı¨ve groups during the shape condition reach periods (RP2) (T ? = 8, N = 7, P = 0.180). However, the naı ¨ve group reached significantly longer during RP1 compared with RP2 (T ? = 5, N = 8, P = 0.039). During expected trials, no significant differences were found between RP1 and RP2 both for frequency and for duration of reaches (frequency: T ? = 10, N = 6, P = 1.0, two-tailed; duration: T ? = 7, N = 8, P = 0.148, twotailed; Fig. 4).
In contrast to the performance of the ''priming'' group in the shape condition during the first reach period (RP1), in the color condition the same group did not show any differences in frequency or duration of reaching in unexpected compared with expected trials (Wilcoxon signed rank test: frequency, T ? = 23.5, N = 11, P = 0.214; duration, T ? = 25, N = 12, P = 0.151, Fig. 5). In contrast, during RP1, the ''naı ¨ve'' group reached more often and for longer time during unexpected compared with expected trials (frequency, T ? = 10, N = 9, P = 0.080; duration, T ? = 17, N = 11, P = 0.087, Fig. 5).
As for the shape condition, also here only some subjects, from both groups, retrieved the ''missing'' pellet in at least one of the two unexpected trials (N priming group = 10 and N naı¨ve group = 8). Thus, as previously conducted, we will focus on those subjects while analyzing their performances. During unexpected trials, the priming group reached significantly more often and for longer time during RP1 than during RP2 (frequency: T ? = 3, N = 8, P = 0.02; duration: T ? = 9, N = 10, P = 0.032, Fig. 6). In contrast, during expected trials, the group reached significantly more often and for longer time during RP2 than during RP1 (frequency: T ? = 3, N = 8, P = 0.031, two-tailed; duration: T ? = 6, N = 10, P = 0.027, two-tailed, Fig. 6).
Similar to the performance of the priming group, during unexpected trials, the naı¨ve group also reached significantly more often and for longer time during RP1 compared with RP2 (frequency: T ? = 0, N = 7, P = 0.008; duration: T ? = 4, N = 8, P = 0.027, Fig. 6). However, during expected trials, no significant differences were found between RP1 and RP2 both for frequency and for duration of reaches (frequency: T ? = 9, N = 6, P = 0.813, two-tailed; duration: T ? = 9, N = 8, P = 0.250, two-tailed, Fig. 6).
Replicating previous work, the present study found that great apes can spontaneously use color properties to individuate objects. The different previous experience (with or without priming) of the two groups did not make much difference to their use of color: both groups showed increased search behavior when finding a food item of a different color compared to the one they originally saw and showed decrease in search only after finding the original item. With regard to shape properties, in contrast, the present study showed that great apes can use shape to individuate objects, but only after some previous experience: the priming group, unlike the naı¨ve group, also showed the above-mentioned search patterns in shapebased individuation tasks. What this clearly suggests is that apes' failure to spontaneously use shape for the individuation of food items that was found in previous studies (and replicated here) does not reflect any deep competence problem. This lack of spontaneous shape-based food object individuation can be alleviated with some previous experience-and with very little and very shallow experience (a couple of encounters with the food items) indeed.
These findings have at least two wider implications. First, there are implications regarding the nature of cognition about food. In comparative and developmental psychology, there is currently some debate about the question of whether and to which degree food constitutes a special cognitive domain with dedicated domain-specific, hard-wired machinery (on a par with naı ¨ve physics, number, space etc.; see e.g., Hauser and Spelke 2004;Santos et al. 2001;Shutts et al. 2009;Spelke and Kinzler 2007). The accounts range from strong domain-specific nativism (there are innate domain-specific beliefs-for example, in the domain of food that color is a reliable indicator of identity, whereas in the domain of tools, color is irrelevant but shape matters.) via intermediate positions (e.g., there are domain-specific learning mechanisms that gradually lead to differential sensitivity to different properties in different domains) to purely domain-general accounts (e.g., that there is only one kind of general purpose learning mechanism that inductively picks up on different diagnostic values of different properties in different domains; for an excellent exposition of this logical space of possible accounts, see Shutts et al. 2009). The fact that some experience-in fact, very little experience-can make properties available for food object individuation that are not spontaneously used (i.e., shape), speaks against any strong domain-specific nativist position. In this respect, the present findings are in line with developmental data that differential attention to color as diagnostic property for food individuation and induction is a relatively late developing phenomenon that only arises after infancy (Shutts et al. 2009). Both of these lines of research taken together thus narrow down the logical space of accounts to such construals that either posit weaker domain-specific learning mechanisms (rather than strongly innate domainspecific beliefs) or that posit domain-general learning mechanisms leading to domain-specific predispositions (e.g., Karmiloff-Smith 1992). Needless to say, more comparative and developmental research is needed to decide between these and further narrow down the hypothesis space.
Second, the present findings have some implications regarding object individuation in non-human animals. Against the background of the debate about potentially uniquely human and linguistically constituted property/ kind-based object individuation, the present findings corroborate previous findings that property-based object individuation is possible in the absence of language. Moreover, the findings suggest that property/kind-based object individuation in primates is not just a very limited phenomenon in some very restricted domain, say for just one kind of property, but seems to be a more general and reliable ability extending to different types of properties. What the present findings do not tell us, however, is what underlies primates' competence in the kinds of tasks used here. In particular, do these tasks tap true kind-based object individuation, or do they just measure sophisticated tracking of features (see Mendes et al. 2008;Xu 2002). What we need in future research to decide between these different possibilities are tasks that tease apart property and kind information. This could be done, for example, by introducing different kinds of object transformations (property transformations pitted against kind transformations; see, e.g., Feigenson and Carey 2003;Xu et al. 2004, Fig. 6 Mean average (?SE) frequency a and duration b of reaches during both 20-s reach periods (RP1 and RP2) of those subjects (N ''priming'' group = 10 and N ''naı¨ve'' group = 8) who found the ''missing'' pellet in at least one of the unexpected trials of the color condition for some attempts in these directions in infancy work). The use of match-to-sample tasks (based on color and shape properties) might also help to clarify the aforementioned question.
Acknowledgments This research was partially supported by a Ph.D. grant from the ''
Open Access This article is distributed under the terms of the Creative Commons Attribution Noncommercial License which permits any noncommercial use, distribution, and reproduction in any medium, provided the original author(s) and source are credited.
The domestic dog mitochondrial DNA (mtDNA)-gene pool consists of a homogenous mix of haplogroups shared among all populations worldwide, indicating that the dog originated at a single time and place. However, one small haplogroup, subclade d1, found among North Scandinavian/Finnish spitz breeds at frequencies above 30%, has a clearly separate origin. We studied the genetic and geographical diversity for this phylogenetic group to investigate where and when it originated and whether through independent domestication of wolf or dog-wolf crossbreeding. We analysed 582 bp of the mtDNA control region for 514 dogs of breeds earlier shown to harbour d1 and possibly related northern spitz breeds. Subclade d1 occurred almost exclusively among Swedish/Finnish Sami reindeer-herding spitzes and some Swedish/ Norwegian hunting spitzes, at a frequency of mostly 60-100%. Genetic diversity was low, with only four haplotypes: a central, most frequent, one surrounded by two haplotypes differing by an indel and one differing by a substitution. The substitution was found in a single lineage, as a heteroplasmic mix with the central haplotype. The data indicate that subclade d1 originated in northern Scandinavia, at most 480-3000 years ago and through dog-wolf crossbreeding rather than a separate domestication event. The high frequency of d1 suggests that the dog-wolf hybrid phenotype had a selective advantage.
Studies of the global diversity for dog mitochondrial DNA (mtDNA) have shown that a homogenous gene pool is universally shared, indicating a single geographical origin for all dogs. Among 1543 dogs from across the Old World (Pang et al. 2009), the mtDNA haplotypes [based on 582 bp of the control region (CR)] were distributed in six phylogenetic groups, called clades A-F.
Clades A, B and C were homogenously distributed at high frequencies, normally 100%, among all populations and therefore probably formed in a single domestication event (Pang et al. 2009). Diversity was distinctively higher in southern East Asia than in other regions, indicating that the place of origin was East Asia (Savolainen et al. 2002) specifically South China or Southeast Asia (Pang et al. 2009). The other clades (D, E and F) had limited geographical distributions and low total frequency. Clades E and F were found exclusively in East Asia and possibly formed together with clades A, B and C. By contrast, clade D was geographically restricted to North Europe, Siberia, Southwest Asia and the Mediterranean (Angleby & Savolainen 2005;Pang et al. 2009), and is therefore the only mtDNA haplogroup that clearly did not originate in East Asia. Analysis of complete mtDNA genomes showed clade D to consist of two subclades (d1 and d2), which separated at least 50 000 years ago (Pang et al. 2009), well before the origins of dogs approximately 10 000-15 000 years ago (Clutton-Brock 1995;Pang et al. 2009; see also Appendix S1). The two subclades had separate geographical distributions (d1 restricted to North Eurasia and d2 to Southwest Asia and the Mediterranean) and are therefore likely have separate origins from wolf.
Subclade d2 and clades E and F were found in only a few percent of dogs within their distribution ranges. By contrast, subclade d1 had a frequency above 30% in native breeds in its core distribution area in Northern Scandinavia (Angleby & Savolainen 2005;Pang et al. 2009), making it the only mtDNA haplogroup that both is geographically restricted to a limited region outside East Asia and has a high frequency. Therefore, it represents the only clear sign of a major separate influx of Ôwolf genesÕ into the dog gene pool.
Importantly, a haplogroup introduced through crossbreeding of wolf into an established dog population would normally, starting from a low initial frequency, remain in the minority. By contrast, haplogroups introduced through independent domestication of wolf, by a human population without dogs, would have an initial frequency of 100%. The high frequency of subclade d1 is therefore the only clear indication in the mtDNA data of a possible independent domestication of wolf, separate from that which formed clades A, B and C. Detailed knowledge about the origins of subclade d1 is, consequently, of great importance for elucidating the earliest history of the domestic dog. Therefore, we investigated the origin of mtDNA subclade d1 to find out where and when it formed and whether this was through independent domestication of wolf or through hybridization between male dog and female wolf.
We performed an extensive survey of the breeds previously known to harbour subclade d1 haplotypes [analysing 68.8% of all known female lineages, according to the pedigree data bases of the Swedish (http://kennet.skk.se/hund data/) and Finnish (http://jalostus.kennelliitto.fi) Kennel Clubs] and of possibly related breeds and populations, based on similar morphology and/or close geographical distribution (Table S1). Hereby, we can present a comprehensive picture of the genetic and geographical diversity for mtDNA subclade d1.
We studied 582 bp of the mtDNA CR for 514 individuals (280 DNA sequences generated in this study), representing at least 328 female lineages of Scandinavian and Arctic Spitz breeds (Tables 1, S2 & S3; see Appendix S1 for exact definition of ÔlineageÕ).
Subclade d1 haplotypes were found only in breeds from Scandinavia, except for one lineage each in East Siberian and Russo-European Laika (Table 1; Table S2). A high proportion of individuals having d1 was found for some common Scandinavian breeds: Lapponian Herder (75% of investigated lineages), Ja ¨mthund (74%), Finnish Lapphund (65%), and Norwegian Elkhound (grey) (46%) and some breeds supposedly founded by few lineages had exclusively d1 haplotypes (100%). Notably, several common Scandinavian/Nordic breeds, e.g. Swedish Vallhund and Norwegian Buhund, did not have d1 haplotypes. Thus, subclade d1 was found at high frequency in all Sami-related breeds (Finnish and Swedish Lapphund, and
Lapponian Herder) and in some North Scandinavian hunting dog breeds. The high frequency of d1, above 50%, in lineages (Table 1, Fig. 1a) of a number of Scandinavian breeds, is remarkable. It is the only example of a mtDNA haplogroup found only in a specific type of dogs from a restricted geographical area and in the majority of the individuals in this area.
The fact that the d1 haplotypes were almost exclusively found among breeds from Northern Scandinavia and Finland strongly indicates that this haplotype originated in this region. Other datasets do not give further clues: Neolithic dog samples from southern Sweden (Malmstro ¨m et al. 2008) carried only haplogroups A and C, but the samples were from outside the historical distribution of the breeds carrying d1, and among wolves no haplotypes similar to subclade d1 were found among extant populations across Eurasia, including Scandinavia (Aggarwal et al. 2007) or historical Scandinavian samples (Flagstad et al. 2003).
Despite the exhaustive sampling, which increased the number of lineages shown to carry d1 haplotypes from 20 to 63, only four previously identified haplotypes (Savolainen et al. 2002; see also Appendix S1) were found (Fig. 1b). The haplotypes have a star-like distribution: a central haplotype (D1, found in 45/63 (71.4%) of d1 lineages and in all breeds having subclade d1, except Finnish Spitz) surrounded by three less frequent haplotypes, two differing by a single-base indel and only one haplotype (D4), found in a single lineage, differing by a substitution. Importantly, the lineage having the D4 haplotype was heteroplasmic for the substitution separating it from D1, thus having a mixture of D1 and D4.
At inspection of nine related individuals, two had non-heteroplasmic D1 (as interpreted from the Sanger chromatograms), two non-heteroplasmic D4 and five had a mixture of D1 and D4 in varying proportions. In addition, some individuals having haplotype D3 were heteroplasmic.
The central position and high frequency for haplotype D1 and the non-fixed state of D3 and D4 suggests that haplotype D1 was the founder haplotype for subclade d1, introduced from wolf to dog, and that D2-D4 have been derived by mutations in the dog population.
The low genetic diversity, with a single lineage of 63 carrying a non-fixed substitution, indicates a recent origin for subclade d1. The mean number of substitutions compared with haplotype D1 among the 63 lineages, the statistic q (Forster et al. 1996), is 0.016 (SEM = 0.0020) and the substitution rate for the 582 bp region has been estimated to one substitution per 40 000-155 000 years (Pang et al. 2009). From this, we estimate the time since the origin of dog haplogroup d1 to be 480-3000 years. Considering the non-fixed state of haplotype D4, this is most likely an overestimation (See further discussion in Appendix S1).
There is archaeological evidence for dogs in both Sweden and Siberia by 8000 years ago (Arnesson Westerdahl 1983;Lepiksaar 1984;Morey 2006). Therefore, an origin of subclade d1 less than 3000 years ago indicates that it derives from crossbreeding of wolf into an already established dog population carrying haplogroups A-C, and not from independent domestication of wolves. The high frequency of d1 haplotypes among Scandinavian dogs, therefore, warrants an alternative explanation. One possibility is that the Ó 2010 The Authors, Animal Genetics Ó 2010 Stichting International Foundation for Animal Genetics, 42, 100-103 Klu ¨tsch et al.
offspring from the crossbreeding had a successful phenotype selected for by humans. If females descending from the crossbred litter were selected for during the first few generations, an increased frequency of d1 would be the result.
The sharing of the d1 haplotypes between the Lapphund breeds associated with the non-Indo-European speaking and nomadic Sami and some hunting breeds connected to Indo-European speaking farmers (Table S1) is notable. Possibly, efficient hunting and herding dogs were items of trade between the two populations. The direction of this trade is not clear, but an origin of d1 among the Sami related breeds is indicated, as all these breeds have d1 haplotypes, while only some breeds linked to the Indo-Europeans have this haplotype (Table 1).
As haplogroup d1 probably derives from dog-wolf crossbreeding, there are no clear signs in the mtDNA data that dogs were domesticated more than once (Pang et al. 2009). Together with haplogroups d2, E and F, haplogroup d1 represents one of only four indications of crossbreeding between female wolf and male dog through history. Whether female dog-male wolf crossbreeding has been equally rare is unclear because of the lack of comprehensive studies of paternally-inherited markers.
17 Arctic spitz breeds, 17 of the most common breeds and types of spitz dogs in the Arctics (e.g. Samoyed, Siberian Husky, Inuit sled dog and seven varieties of Laika; see TableS2
for a complete list); n, number of samples; No. of lineages, the minimum number of female lineages among the samples; Prop. (%) of known lineages, proportion (in percent) of the known female lineages in the Swedish and Finnish pedigree data bases; d1 (%), number of lineages (percent of the analysed lineages within parenthesis) having a subclade d1 haplotype; D1 (%) through D4 (%), number of lineages (percent of the lineages having a d1 haplotype within parenthesis) having haplotype D1 through D4.
Ó 2010 The Authors, Animal Genetics Ó 2010 Stichting International Foundation for Animal Genetics, 42, 100-103
Ó 2010 The Authors, Animal Genetics Ó 2010 Stichting International Foundation for Animal Genetics, 42, 100-103 Dog-wolf hybridization in Scandinavia
The authors are grateful to numerous dog owners for the collection of samples, and especially
Additional supporting information may be found in the online version of this article.
Table S1 Summary of (often anecdotal) information in the literature about breed history of Scandinavian and Arctic Spitz-type breeds. Table S2 Representation of subclade d1 haplotypes among Scandinavian and Arctic spitz breeds (see also Note to Table 1). In addition, the number of individuals carrying haplotypes from clades A-C, E and F is given. Table S3 List of samples giving mtDNA haplotype, breed, geographical origin and pedigree information (name and registration number, motherÕs and fatherÕs name and earliest female ancestor).
As a service to our authors and readers, this journal provides supporting information supplied by the authors. Such materials are peer-reviewed and may be re-organized for online delivery, but are not copy-edited or typeset. Technical support issues arising from supporting information (other than missing files) should be addressed to the authors.
Anthropogenic pollutants comprise a wide range of synthetic organic compounds and heavy metals, which are dispersed throughout the environment, usually at low concentrations. Exposure of ruminants, as for all other animals, is unavoidable and while the levels of exposure to most chemicals are usually too low to induce any physiological effects, combinations of pollutants can act additively or synergistically to perturb multiple physiological systems at all ages but particularly in the developing foetus. In sheep, organs affected by pollutant exposure include the ovary, testis, hypothalamus and pituitary gland and bone. Reported effects of exposure include changes in organ weight and gross structure, histology and gene and protein expression but these changes are not reflected in changes in reproductive performance under the conditions tested. These results illustrate the complexity of the effects of endocrine disrupting compounds on the reproductive axis, which make it difficult to extrapolate between, or even within, species. Effects of pollutant exposure on the thyroid gland, immune, cardiovascular and obesogenic systems have not been shown explicitly, in ruminants, but work on other species suggests that these systems can also be perturbed. It is concluded that exposure to a mixture of anthropogenic pollutants has significant effects on a wide variety of physiological systems, including the reproductive system. Although this physiological insult has not yet been shown to lead to a reduction in ruminant gross performance, there are already reports indicating that anthropogenic pollutant exposure can compromise several physiological systems and may pose a significant threat to both reproductive performance and welfare in the longer term. At present, many potential mechanisms of action for individual chemicals have been identified but knowledge of factors affecting the rate of tissue exposure and of the effects of combinations of chemicals on physiological systems is poor. Nevertheless, both are vital for the identification of risks to animal productivity and welfare.
Historically, ruminant animals production systems were of relatively low intensity; the inputs of energy, food and fertiliser were small and outputs of meat, milk and by-products -E-mail: s.rhind@macaulay.ac.uk were low. Accordingly, fertiliser inputs comprised, primarily, animal and human manure with more unusual products such as seaweed being used only where available. Significant accumulation of pollutants in soils, and potential exposure of animals to elevated concentrations of pollutants was rare and occurred only in certain highly specialised circumstances, for example, where soils were repeatedly fertilised, with ash and bird carcases (Meharg et al., 2006). On the other hand, the potential for pollutant exposure in modern production systems is greatly increased, for several reasons.
First, modern animal production systems involve the widespread use of pesticides and herbicides and traditional organic fertiliser has been replaced to a significant extent with synthetic nitrate fertilisers. However, the economical and environmental costs of inorganic fertiliser production, together with anthropogenic waste generation, have led, in many countries, to a return to the use of processed sewage sludge instead (Commission of the European Communities, 1994;Swanson et al., 2004). The chemical profile of sludge (Smith, 1995) reflects the mix of thousands of different environmental chemicals to which we are exposed, and its use gives rise to potentially increased exposure to, and bioaccumulation of, chemicals. Other potentially polluted materials such as composted green waste, derived from domestic and municipal sources are also being applied to land. Thus, domestic ruminants may be exposed to concentrations of pollutants that are higher than those occurring 'naturally' in the environment, either through the application of chemicals or because they are exposed to modern waste through recycling of waste to land. However, at this time, it is logistically impossible to characterise the magnitude of potential exposure owing to the complexity of the mixture and the high cost of analysis.
Second, while aerial input of pollutants to the environment as a result of human actions would originally have been relatively trivial, comprising small amounts of potentially toxic metals (PTMs) derived from the smelting of ores (Hong et al., 1996) and polycyclic aromatic hydrocarbons (PAHs) as a result of natural and man-made fires (Bostrom et al., 2002;Yunker et al., 2002), in recent decades aerial inputs have increased enormously. Thousands of new, partially volatile, organic chemicals manufactured for a multitude of industrial, agricultural and domestic purposes are being released into the environment. Occasionally, air pollution has been sufficiently severe to cause damage to plants (e.g. acid rain) but effects on animals, and ruminants in particular, have seldom been proven, although often suspected (Kelly, 1995).
Environmental pollutants -what are they and from where do they originate?
Most, but not all, pollutant types are derived from human activities. From around the time of the Second World War, the nature of the pollutant burden within the environment, its distribution and effects have changed significantly, largely as a result of human activities and inventions. As indicated above, while PTMs have been used for thousands of years, their use has increased enormously and so they have become ubiquitous in the environment and now appear at low concentrations in soils and at higher concentrations in products such as sewage sludge (Smith, 1995). The potential for lowenvironmental concentrations of metals to perturb animal physiology is now being recognised (Spurgeon et al., 1994).
Nitrates are another category of inorganic pollutant that can interfere with animal physiology and in particular with reproductive function; specifically, they have been implicated in altered thyroid function and the disruption of gonadal steroidogenesis (Guillette and Edwards, 2005;Edwards et al., 2006). Thus, their presence in drinking water is of concern to animal (and human) health and productivity.
Human activity and industrialisation have resulted in an increase in the manufacture, use and release into the environment, of a wide variety of organic substances (Table 1).
Although not designed to be physiologically active, many of these can bind to steroidal and other cellular receptors or otherwise interfere with endocrine signalling or enzyme systems and thus affect physiological processes in species as diverse as bacteria (Fox, 2004) and mammals (Toppari et al., 1996;Sweeney et al., 2000). Environmental concentrations of these pollutants are seldom high enough to exert toxic effects, in the conventional sense, being similar to background levels (Rhind, 2009), but through subtle effects on physiological systems they can interfere with normal function. Rhind et al.
Chemicals that induce effects by perturbing endocrine systems or mimicking endocrine mediators are collectively described as endocrine disrupting compounds (EDCs). Although pollutants can be very different, chemically and mechanistically, for the purpose of this review, it is appropriate to consider all of these organic and inorganic pollutant classes together and to loosely define them as EDCs because all are known to have disruptive capabilities and they have the potential to interact, additively (Bemis and Seegal, 1999).
Investigations of the effects of EDC exposure on animals have frequently involved the application of pharmacological doses of individual pollutants to laboratory rodents (to elucidate effects and mechanisms of action and to identify risks associated with exposure to individual chemicals) and less controlled studies of wild animals in which effects had been observed following known pollution incidents (Colborn et al., 1993; Institute for Environment and Health (IEH), 1999). Such studies have identified several characteristics of EDCs that make them biologically significant in relation to all animals, including domestic ruminants. They (i) are highly persistent -for example, many EDCs, including PAHs, polychlorinated biphenyls (PCBs) and polybrominated diphenyl ethers (PBDEs), have half lives of about 10 years or more (Smith, 1995) or in the case of heavy metals are never degraded, (ii) are ubiquitous, that is, although production and use can be localised, EDCs are distributed throughout the environment, (iii) frequently accumulate in animal tissue because they are hydrophobic and lipophilic (Nimrod and Benson, 1996), (iv) exert effects on physiological systems at very low concentrations, orders of magnitude lower than those known to have acute toxic effects (Brevini et al., 2005;Fowler et al., 2007a), (v) have unpredictable effects as they can act additively or synergistically (Payne et al., 2000; Rajapakse et al., 2002; Crofton et al., 2005; Hauser et al., 2005), depending on circumstances; mixtures of compounds can induce biological responses, even when each chemical is present at concentrations too low to induce a biological response by itself (Rajapakse et al., 2002;Kortenkamp, 2007) and (vi) can induce changes in organ structure or function in subsequent, unexposed, generations (Bøgh et al., 2001;Anway and Skinner, 2006;Edwards and Myers, 2007;Steinberg et al., 2008).
Collectively, these properties mean that low-level exposure to pollutants has the potential to affect ruminant productivity and, for example, through changes in the immune system, health and welfare. However, the magnitude of the responses, which is likely to depend on many factors, including the rate, timing and duration of exposure, is ill defined and poorly understood for all animal species.
Ruminants, like other terrestrial species, can be exposed to pollutants, at least theoretically, through ingestion of food and water, through inhalation and by absorption through the skin. It is widely accepted that the primary route of exposure in such species is via the diet (Fries, 1995;Norstrom, 2002) but the significance of other routes of exposure has not been extensively investigated and they may yet prove to be important. Exposure of target organs also depends on the chemical class and associated properties, the age and stage of development of the exposed animal (i.e. foetal, neonatal and adult), the rate of pollutant uptake and rates of subsequent degradation, excretion and/or metabolism (Meador et al., 2008;Rhind, 2008). None of these determinants has been well characterised in ruminants, or indeed for any species, or for any of the classes of EDCs. However, as exposure rates are normally very low and as the health and productivity of the majority of ruminant populations, like that of humans, appears to be unaltered by such levels of exposure, at least superficially, it might be concluded that the environment is almost always entirely benign.
In certain production systems, ruminants can be exposed to slightly higher levels of thousands of different pollutants, relative to those seen in the wider environment, for example when animals are grazed on pastures fertilised with sewage sludge (Rhind, 2005) or drink water contaminated with sewage (Meijer et al., 1999). The practice of recycling human waste is ancient and 'night soil' was collected and returned to land both before and after industrialisation. Processed sewage sludge, as generated in the 21st century, however, is a very different product to that, which has been used as fertiliser in the past as it contains variable combinations of anthropogenic pollutants including organic pollutants and PTMs from domestic, agricultural and industrial sources, at much higher concentrations than those found in the rest of the natural environment (Smith, 1995). However, the actual pattern of exposure to individual pollutants is unknown because it would require analysis of thousands of different chemicals. Rates of tissue accumulation of selected chemicals, and associated effects on the physiology of grazing animals on pastures fertilised with sewage sludge have been addressed both theoretically and in practice. Theoretical estimation of tissue accumulation of individual pollutants would suggest that increases in tissue levels would be small and of no physiological consequence (Wild and Jones, 1992;Duarte-Davidson and Jones, 1996). These conclusions have been largely supported by empirical studies that involved exposure of animals to sewage sludge or to specific compounds (Fries and Marrow, 1977;Fries et al., 1978;Fries, 1996;Rhind et al., 2005aRhind et al., , 2005bRhind et al., , 2007Rhind et al., and 2009)). None of these studies reported the patterns of reproductive performance associated with exposure to pollutants but related studies, discussed below, have shown that even such small increases in tissue concentrations of the pollutants measured, following sludge exposure, are associated with physiological changes. It should be noted that a relatively limited range of chemical types has been measured in tissue and it cannot be assumed that those measured are the ones responsible for inducing the observed effects; at best, the reported values represent an index of the total pollutant 'insult', which is likely to include several thousand chemicals.
The fundamental mechanisms of action of pollutants have been investigated and reviewed previously (Sikka and Naz, 1999;Rhind, 2002). The majority of experiments have been based on laboratory rodents and studies in ruminants are relatively rare; a small number of studies, mostly in sheep, have addressed the physiological effects of EDC exposure, using the classical model whereby relatively high concentrations of selected pollutants are administered for short periods of time (Beard et al., 1999; Sweeney et al., 2000; Wright et al., 2002). However, additional results are now emerging from investigation of the effects of prolonged, low-level exposure to EDC mixtures in sheep that have been maintained on sewage sludge-treated pastures (Erhard and Rhind, 2004; Paul et al., 2005; Fowler et al., 2008; Bellingham et al., 2009; Lind et al., 2009). In these studies, the concentrations of chemicals to which animals were exposed, and the specific mixture of chemicals involved, were not comprehensively defined because this would have been logistically impossible. However, effects of exposure were demonstrated by comparing animals reared on sludgetreated pastures with others reared on comparable pastures treated with inorganic fertiliser, which contains minimal amounts of pollutants.
The majority of reported ruminant trials have concentrated on the effects of EDC exposure on the reproductive axis and the results can be classed according to organ/function.
The endogenous activity of the gonads is driven by the actions of the hypothalamus and pituitary gland, both of which are steroid-sensitive and, thus, potential targets for EDCs. Studies in sheep and goats have reported significant effects of exposure to chemicals with known endocrine disrupting effects, including octylphenol (Sweeney et al., 2000; Wright et al., 2002), bisphenol A (Evans et al., 2004; Savabieasfahani et al., 2006), methoxyclor (Savabieasfahani et al., 2006), PCB153 (Lyche et al., 2004;Oskam et al., 2005) and valporate (Krogenaes et al., 2008), on the hypothalmicpituitary (HP) gland axis. The majority of these studies have investigated the effects of EDCs on the HP axis during development, that is, during gestation and/or lactation, the most sensitive 'windows', when the endocrine system can be permanently altered (IPCS, 2002). Comparison of the results obtained in these studies show that although the chemicals applied and/or the purported mechanisms of action were similar, the observed effects were often chemical and species specific. For example, the timing of puberty, which reflects activation of the HP axis, was advanced in female lambs exposed to octylphenol (Wright et al., 2002), delayed in both male (Oskam et al., 2005) and female (Lyche et al., 2004) goats exposed to PCB153, and remained unchanged in male (Oskam et al., 2005) and female (Lyche et al., 2004) goats exposed to PCB126, during development. These results illustrate the complexity of the effects of EDCs on the HP axis and the difficulty of extrapolating between, or even within, species.
In line with the hypothesis that exposure, during critical windows of development can be more disruptive to physiological systems, the results of studies that exposed sheep to the same dose of octylphenol, at different development stages, indicated different physiological consequences. Exposure to pharmacological doses during development (in utero and early post natal period), the period during which HP axis differentiation and sexual dimorphism occurs (Rhind et al., 2001; Robinson, 2006), was found to be potentially more detrimental than later exposure and was seen to induce effects on reproductive function/physiology which were manifested only in later life. For example, gestational exposure to octylphenol at 1 mg/kg per day for 2 weeks resulted in reduced foetal FSH secretion, which compromised testis development (Sweeney et al., 2000) and lactational exposure was associated with altered semen quality (Sweeney et al., 2007). Similarly, octylphenol exposure of female lambs, in utero, altered FSH secretion during the late follicular phase, and changed the timing of puberty (Wright et al., 2002). However, exposure to similar doses of octylphenol during the pre-pubertal period had no significant effect on either LH or FSH secretion in female lambs (Evans et al., 2004).
Bisphenol A exposure, also at pharmacological doses (5 mg/kg per day for 2 months) has also been shown to suppress LH secretion in female sheep either when exposure occurs during development (Savabieasfahani et al., 2006) or during the pre-pubertal period (Evans et al., 2004). Although it was not possible to determine the exact nature and location of action in these studies, Katoh et al. (2004) showed that bisphenol exposure affected growth hormone through an effect exerted at the level of the pituitary gonadotrophes, whereas other studies such as those by Wright et al. (2002) with octylphenol, and Lyche et al. (2004) with PCB153, have indicated that EDC exposure can have effects at the level of the hypothalamus as puberty is driven by maturational changes in the hypothalamus and the timing of puberty was affected in both studies. A criticism of the above studies that have investigated the effects of EDCs on the HP axis is that they addressed effects of exposure to pharmacological doses of single chemicals, at concentrations hundreds or thousands of times higher than the levels present in the environment, where effects are likely to be exerted through the actions of many chemicals, at low concentrations, in combination. Studies using the sewage sludge model described earlier provide a means to address this real world exposure. Preliminary results indicate that sludge exposure alters the population of gonadotrophes in the pituitary glands of adult ewes that had been maintained on these pastures and changes the phenotype of pituitary cell populations. Changes in the activity of a number of neurotransmitter systems within the hypothalamus have been reported (Bellingham et al., 2009). Given the fundamental importance of the HP axis in the regulation of normal gonadal function, alterations to this system by EDCs, may have deleterious consequences for ovarian or testicular function and thus reproductive function and fertility. Testis As for the HP axis, most studies of EDC effects have involved laboratory animals and employed levels of chemical exposure that are probably not environmentally relevant (Hotchkiss et al., 2008). More recently, several well-designed studies have investigated the effect of mixtures of EDCs on the developing rodent testis and its functions, and have shown that combinations of, for example, anti-androgenic EDCs, exert major effects at doses at which the individual EDCs have no significant effect (Christiansen et al., 2008; Rider et al., 2009). Such observations indicate that it is likely that in domestic animals, exposure to the thousands of EDCs in the environment will exert effects on the developing testis, although the large numbers of chemicals involved and the potential complexities of their interactions make it difficult to predict the incidence or severity of such effects. Nevertheless, the high incidence of reproductive abnormalities in human males at birth (cryptorchidism and hypospadias) and in adulthood (e.g. low-sperm counts), together with the evidence of temporal changes in incidence of these disorders, especially of falling/low-sperm counts (Swan et al., 2000), is at least consistent with environmental impacts (Skakkebaek et al., 2001; Sharpe and Skakkebaek, 2003).
The issue of falling sperm counts in human males remains controversial, largely because of difficulties in proving/disproving that it is happening and of identifying potential causes (Swan et al., 2000). One argument against sperm counts having fallen is the absence of evidence of any similar decline in domestic ruminants, over the same time period; it is assumed that their exposures should be broadly similar to that of humans (Setchell, 1997). However, this comparison is invalid for two reasons. First, and most important, domestic ruminants store sperm and are thus able to maintain a high and uniform sperm count over many frequent ejaculations, whereas humans do not store sperm and therefore sperm counts are greatly affected by ejaculatory frequency (Sharpe, 1994). Consequently, sperm counts in the ejaculates of domestic ruminants do not provide an accurate insight into the level of sperm production, whereas in humans it does (Sharpe, 1994;Sharpe and Skakkebaek, 2003). Second, domestic animals are constantly selected for high fertility/ sperm production and normal husbandry practice would have resulted in the culling of less fertile animals within the timescale that sperm counts have 'fallen', whereas reproductive technologies act to perpetuate and potentially exacerbate defects in human sperm production.
The rodent EDC mixture studies referred to above have not so far addressed effects on sperm counts/sperm production, although the reported adverse effects on foetal testis development including suppression of androgen production/ action would be expected to reduce Sertoli cell proliferation in foetal life (Scott et al., 2007 and2008). Similar results have also been found in the foetal sheep after pregnant ewes were reared on pasture fertilised with sewage sludge (Paul et al., 2005); specifically, foetal blood testosterone levels were reduced alongside of reductions in Leydig, Sertoli and germ cell numbers. Although these effects were considerable, they did not allow dissection of mechanisms and of cause and effect relationships; on the basis of rodent studies, it would be expected that reduced intra-testicular testosterone concentrations during development would result in reduced Sertoli cell number in the adult (Scott et al., 2007). However, studies in the rat have also shown that even when Sertoli cell number is reduced at birth by 40% to 50%, compensation occurs rapidly after birth so that normal Sertoli cell numbers are restored by puberty and maintained into adulthood (Hutchison et al., 2008;Scott et al., 2008). However, this recovery occurred following cessation of the causal treatment (in this instance, dibutyl phthalate) and so it could be argued that continued EDC exposure through foetal and postnatal life, which is more akin to 'real world' exposures, might interfere with compensatory Sertoli cell proliferation; this possibility remains to be tested.
Rat studies have led to the identification of 'a male programming window' (Welsh et al., 2008), during which androgens (produced by the foetal testis) act to programme later development of the reproductive tract and genitalia. Deficient androgen action during this time window leads to permanent reductions in size of the penis, prostate and testis and increased risk of malformations such as hypospadias and cryptorchidism (Welsh et al., 2008). Exactly, the same time window applies to programming of the male-female difference in anogenital distance and so the latter, at any age after birth, can provide an index of overall androgen exposure/action within the male programming window (Scott et al., 2008; Welsh et al., 2008). A similar time window applies to domestic ruminants, the best characterised being in the sheep (Wood and Foster, 1998), and is thought, by analogy to the rat and human, to start when testosterone production by the foetal testis first commences (Welsh et al., 2008). With regard to the effects of EDC exposure, the most important implication of these findings is that only exposure within the male programming window is likely to affect androgen-dependent reproductive development.
However, development of the normally formed penis (Welsh et al., 2008) and the increase in Sertoli cell number/determination of adult testis size (Scott et al., 2008) also depend upon androgen action after the male programming window and therefore may be susceptible to EDC effects for a greater period. Ovary Ovarian follicle formation in ruminants occurs during foetal life (Ru ¨sse, 1983) and involves the assembly of meiotically arrested oocytes and somatic pre-granulosa cells into primordial follicles (Hirshfield, 1991). EDCs could, potentially, perturb the function of each of the cell types involved. Female germ cells begin meiosis during early development, and EDC-exposure at this time can potentially reduce a female's lifetime reserve of oocytes, which cannot be renewed, unlike males in which continuous spermatogenesis may quench transient EDC effects. On the other hand oogonia may be less sensitive than male germ cells to some gonotoxic insults (e.g. Guerquin et al., 2009).
The developmental stage at which damage occurs determines the impact that exposure to chemicals will have on reproduction (Hoyer, 2005;Uzumcu and Zachow, 2007). Chemicals selectively damaging large growing or antral follicles only temporarily interrupt reproductive function, unlike when damage to the primordial follicle population occurs, because these are replaced by recruitment from the primordial follicle pool. Two key developmental processes occur subsequent to oocyte meiotic arrest: (i) primordial follicle assembly (at 75-day gestation in sheep: Sawyer et al., 2002) and (ii) primordial follicle recruitment. Both are coordinated by paracrine and autocrine growth factors and occur in the later stages of ruminant gestation (McNatty et al., 1999). Primordial follicles are then activated and recruited into the growing cohort of primary follicles, a vital determinant of reproductive life span (Skinner, 2005). Thus it is possible to investigate the effects of EDC exposure in relation to the known developmental changes that occur during maturation of the oocycte, its release and fertilisation.
As multiple ovarian systems are sensitive to EDCs, associated effects are diverse and dependent on the specific chemical involved (Table 2). As for other organs, exposure to environmental concentrations of pollutant mixtures has been shown to perturb ovarian development in the sheep (Fowler et al., 2008; Mandon-Pepin et al., 2009) and follicle health in the adult offspring (Figure 1; Amezaga et al., 2009). Disruption of the first stages of gametogenesis and gonadal differentiation, in utero, ultimately control reproductive viability in mature offspring and can have transgenerational consequences (Anway and Skinner, 2006).
With regard to future investigations of these effects, it should be noted that much of the literature concerning reproductive effects of EDCs is based on rodent models but ruminants exhibit significant differences in physiology compared with, say, laboratory rodents such as mice; for example, steroidogenesis occurs in ruminants during foetal, not postnatal, ovarian development. Thus, the use of the rodent model alone to understand risks posed to domestic ruminants is unwise.
Oocyte maturation and early embryo development During maturation, the oocyte and early embryo are particularly susceptible to pollutants, even at the background doses to which they might be exposed, in vivo, when their mothers are not exposed to elevated environmental levels of pollutants (Pocar et al., 2001a and2001b). For example, incubation with PCBs (0.0001 to 1 mg/ml of a mixture of PCBs (Aroclor); exposed for 24 to 48 h) significantly reduced the percentage of bovine oocytes reaching metaphase II, and lower doses that do not impair meiotic processes reduced Rhind et al.
fertilisation rates and embryonic development (Pocar et al., 2001a and2001b). Some of the underlying mechanisms responsible for this, and similar effects, have been identified (Table 2) but many potential mechanisms remain to be investigated. The processes of endometrial transformation, and development and implantation of the embryo depend on an exchange of hormonal signals between the embryo and mother. During early embryo development and the implantation window, specific amounts of oestradiol (E2) and progesterone (P) are required for endometrial maintenance and chorionic gonadotropin for maintenance of hormone production by the corpus luteum (Makrigiannakis et al., 2006). These vital endocrine dialogues can be disturbed by EDCs. In addition to indirect effects of EDCs on embryo development, exerted via changes in hormone profiles, direct embryotoxic effects are possible through actions of EDCs on hormone receptors (Agras et al., 2007;Davey et al., 2007). It has been shown that the hormone receptor expression in blastocysts differs between embryoblast and trophoblast, which represent different cell lineages, and this results in different developmental impacts (Navarrete Santos et al., 2004a, 2004band 2008). Along with steroid hormone receptors, many of the effects of EDCs on pre-implantation embryos, and on implantation, are mediated through the aryl hydrocarbon receptor (AhR) and peroxisome proliferatoractivated receptor (PPAR) signalling pathways. The AhR, a binding partner of dioxins and coplanar PCBs (McMillan and Bradfield, 2007), is essential for fertility, being involved in folliculogenesis, oestrogen biosynthesis and signalling, progesterone biosynthesis and corpus luteum function (e.g. Pocar et al., 2004, 2005a and 2005b, Li et al., 2006; Barnett et al., 2007, Ohtake et al., 2009). It may be necessary, also, for normal development of the pre-implantation embryo (Clausen et al., 2005) and embryo-maternal signalling during implantation. Both indirect and direct EDC effects were observed at environmentally relevant concentrations with exposure durations between 4 h and 7 days (short-term) or 6 to 12 weeks (long-term).
Pre-implantation exposure to EDCs of a variety of classes, at environmentally relevant concentrations, can perturb the reproductive and implantation systems of early embryos through some of the above mechanisms (Table 3). Effects include perturbation of energy metabolism (Tonack et al., 2007) and downregulation of relevant genes (Hanlon et al., 2005). As PPAR signalling is disrupted by EDCs (Huang, 2008;Nakanishi, 2008) and affects AhR expression (Hanlon et al., 2003;Lovekamp-Swan and Davis, 2003;Villard et al., 2007), it directly links metabolic and EDC pathways with pollutants.
In summary, embryo implantation is highly vulnerable to endocrine disruption, as EDCs, (plasticiser, PAHs, PCBs, PDBEs, dioxins, pesticides, organotins and heavy metals) at concentrations as low as 10 nM or 2 ng/kg and for an exposure period as short as 4 h, can interfere with the actions of many hormones and receptors essential for pre-and peri-implantation development of the embryo and endometrium.
Effects on other aspects of animal health and welfare Although direct effects on components of the reproductive system are often among the most noticeable adverse effects of pollutants, animal performance and welfare can also be compromised by sub-optimal function in other physiological systems. Milk production and associated success in rearing offspring are critical to healthy populations. Rodent studies have shown that mammary development, differentiation and gene expression can be perturbed by exposure to organic pollutants before or around puberty (Fenton, 2006;Moral et al., 2008) and during pregnancy (Fenton, 2006), potentially compromising neonatal nutrition and survival. Similarly, studies of the mammary tissue of sheep exposed to sewage sludge (Fowler et al., 2007b) demonstrated changes in tissue structure and associated protein expression.
The capacity of environmental pollutants to adversely affect the immune system and the importance of exposure during foetal development, are well known from studies of species exposed to heavy pollutant burdens (Martineau et al., 1988;De Swart et al., 1995;Dietert and Piepenbrink, 2006). The complexity and individual variability of the immune response makes it difficult to detect when the suppression is modest, as is likely to be the case in ruminants, but it remains highly likely that EDC exposure is having adverse effects on immune function in this taxon as well.
EDC exposure can affect thyroid function and therefore metabolism; for example, EDCs have been shown to inhibit the expression of nuclear thyroid hormone receptors or perturb the hypothalamic-pituitary-thyroid axis (Jekat et al., 1994;Hansen, 1998;Sugiyama et al., 2005). In addition, PCBs acting via AhRs, can induce multiple histological and physiological changes within the gland, affecting thyroid hormone production (Hansen, 1998). As the thyroid gland is involved in the regulation of many fundamental physiological processes, particularly in the developing animal (Erenberg et al., 1974), and in seasonal reproductive transitions (Shi and Barrell, 1992), its disruption could adversely affect animal health, welfare, reproduction and productivity.
Effects of exposure to environmental pollutants on bone structure have been identified previously in both wildlife (Lind et al., 2004a; Lundberg et al., 2008) and domestic species (Lundberg et al., 2006) and recent work involving the sewage sludge paradigm has shown that exposure to a mixture of pollutants at low concentrations can increase mineral content and reduce bone strength, at least in females, (Lind et al., 2009).
Studies of humans also indicate that environmental pollutants can affect adipogenesis (Stahlhut et al., 2007). Experiments involving various species indicate that phthalate exposure can induce insulin resistance and alter receptor activity, glucose transporters and transcription factors (Alonso-Magdalena et al., 2005; Fujiyoshi et al., 2006; Grun et al., 2006). Thus, it seems likely that subtle changes in nutrient partitioning and in feed efficiency will occur in EDCexposed ruminants although they are, as yet, undetected.
Exposure of rats to PCBs increased serum cholesterol concentrations and blood pressure, risk factors for heart disease (Lind et al., 2004b), and studies of humans suggest a relationship between dioxin exposure and risk of cardiovascular disease (Humblet et al., 2008). Theoretically, such subclinical disorders have the potential to compromise ruminant health and welfare, causing small reductions in productivity.
One measure of neuroendocrine development is offspring behaviour. Altered patterns of behaviour have been reported in children exposed to various pollutants during pre-natal and early post-natal development (Vreugdenhil et al., 2002; Lanphear et al., 2005; Korrick and Bellinger, 2007) and in lambs exposed to sewage sludge, via their dams (Erhard and Rhind, 2004). Altered social or sexual behaviour, learning ability or fearfulness all have the potential to reduce animals' capacity to obtain food, breed and compete successfully with others in the flock or herd.
In summary, environmental pollutants can adversely affect diverse physiological systems and processes in many species, including ruminants. These effects are not generally reflected in visible reductions in animal performance but sub-clinical effects may result in subtle reductions in animal performance, with associated economic consequences. Furthermore, such underlying physiological changes may become increasingly important as new chemicals are manufactured or as concentrations of others in the environment are increased. These add to the animal burden but readily observable effects may only become apparent if, or when, a critical level of the 'insult' is reached. For example, human sperm quality and fertility are reportedly declining over time (Nordstrom Joensen et al., 2009) but effects on conception rates may become obvious only when the number of fertile sperm declines to a critical level. The fact that some effects are known to be exerted on the developing foetus and are expressed in the adult animal and, critically, also in subsequent generations (Bøgh et al., 2001), means that exposures now may lead to animal production problems in the future.
In order to understand risks and effects, the input of pollutants into biological systems has to be defined. However, at present, measures of environmental concentrations are costly, often technically difficult, and generally limited to relatively few of the thousands of chemicals that may be important. Such measurements rarely take account of factors that can affect biological availability such as substrate binding or conversion into different forms with different chemical characteristics and/or biological effects.
Similarly, the processes regulating transfer of pollutants between the environment and the target organs, in animals at each stage of development, must be better characterised. The overall efficiency of this process depends on several components including the rate of ingestion or inhalation and pollutant availability, as determined by the strength of bonds between it and the substrate (food, soil, water and air). Transfer depends, also, on the efficiency of physiological processes such as uptake from the maternal gastrointestinal tract, lungs or skin and, following uptake by the dam, uptake from the maternal circulation via the placenta, rates of foetal metabolism, excretion and absorption into lipid stores (Rhind, 2008). These processes have been quantified for very few species or pollutant classes.
Understanding the effects of pollutants requires knowledge of actions at the cellular and molecular level. Different sets of genes are expressed while others are silenced, in a developmental stage and tissue-specific manner, by the coordinated action of a number of epigenetic mechanisms that involve chemical modifications to both DNA and chromatin (Li, 2002;Morgan et al., 2005). The precedent that EDCs can epigenetically modify at least one of these modifications (i.e. DNA methylation) in the germ line, thereby promoting transgenerational abnormalities including impaired male fertility, has been established in rats exposed to the agricultural Rhind et al. chemicals vinclozolin and methoxychlor (Anway et al., 2005). Although it remains to be determined if the effects of EDCs in ruminants operate by similarly dysregulating the normal pattern of epigenetic-mediated gene expression, a different 'insult', in the form of clinically relevant reductions in specific micronutrients during the peri-conceptional period, in sheep has been shown to lead to widespread epigenetic alterations to DNA methylation in offspring. Such alterations are associated with obesity, insulin resistance and high blood pressure (Sinclair et al., 2007). Consequently, the focus of future studies on the effects of EDC exposure in ruminants will need to consider such modes of action.
The process of identifying causal relationships between different classes of pollutant and their effects has to be extended using numerous approaches already proven in ruminants. These include measurements of changes in organ structure and/or function (Evans et al., 2004;Paul et al., 2005) and in gene or protein expression (Edwards and Myers, 2007;Fowler et al., 2008). However, current understanding of additive and synergistic effects of EDCs, especially in complex mixtures, is, at best, very limited and so elucidation of these effects will be a critical area of future research. As it is logistically impossible to address every possible combination, it will be appropriate to study combinations with different mechanisms of action in order to better comprehend the effects of mixtures (Kortenkamp, 2007). There is also likely to be a significant role for bioinformatics and computer modelling approaches (Suk et al., 2002). However, such studies may be partially constrained by lack of understanding of mechanisms of action and of data pertaining to the effects of each EDC, individually, let alone when part of a mixture.
The variability associated with each of the multiple processes regulating tissue concentrations of EDCs results in very large individual animal variation in tissue concentrations (Rhind, 2008). Investigation of the relationships between genotype and phenotypic responses might be expected to lead to improved predictability of effects but, owing to the complexity of mixtures and multiple genomic factors it is almost impossible for a single DNA variant site to be consistently associated with a particular trait (Nebert, 2005). This area of work is likely to be of major significance in the future, although frequently unpopular with funding agencies.
Although a number of physiological systems have already been shown to be affected by a wide range of chemicals, it is important to recognise not only that there may be chemicals, which are not yet widely recognised as endocrine disruptors, but also that there may be physiological responses to exposure that are not currently recognised as effects of EDC (Guillette, 2006). Thus, it is important that research into these phenomena is approached with an open mind and receptiveness to previously unidentified risks and mechanisms. The work of the European REACH (Registration, Evaluation and Authorisation of Chemicals) programme is designed to address the issue of chemical use and, by implication, environmental pollutants. Although it will undoubtedly help to identify and control the use of the most persistent, bioaccumulative and toxic chemicals, concerns remain that the legislation may fail to take account of effects of mixtures and of chemicals present at levels deemed to be below the 'no effect level' (Santillo and Johnston, 2006).
Environmental pollutants can adversely affect animal health and reproductive function, through either direct or indirect effects on numerous organs and systems. However, empirical evidence of the relationships between exposure and physiological effects is scarce, particularly for ruminants, reflecting the fact that levels of exposure to each individual chemical are generally very low and they do not act individually. At this time, effects of environmentally relevant levels of exposure to EDCs are not yet reflected in visibly reduced animal performance. Nevertheless, concerns remain that there may be subtle perturbations of reproductive function and since some of the observed changes in physiological function may be expressed in subsequent generations, even without further exposure to pollutants, there may be even greater cause for concern. Like ruminant productivity, human health is generally considered to be good/improving but, at the same time, the incidence of breast cancer in women in the United Kingdom, a disease considered to be related to EDCs, is increasing at a rate of 2% every year (Office for National Statistics, 2008) indicating that some trends in health/performance may only be observed on a population basis rather than an individual basis. It is postulated that comparable insidious effects on ruminants may also be present but until appropriate end points are recognised and measured, such potential threats may remain hidden.
PAH 5 polycyclic aromatic hydrocarbons; PCB 5 polychlorinated biphenyls; PBDE 5 polybrominated diphenyl ethers.
EDC 5 endocrine disrupting compound; BPA 5 bisphenol A; DES 5 diethylstilbestrol; EE2 5 ethinyl estradiol; PCB 5 polychlorinated biphenyls; MEHP 5 mono ethylhexyl phthalate.
Much of the work reported in this paper, and the preparation of the paper, was funded by the
Background Human immunodeficiency virus (HIV) treatment side effects have a deleterious impact on treatment adherence, which is necessary to optimize treatment outcomes including morbidity and mortality. Purpose To examine the effect of the Balance Project intervention, a five-session, individually delivered HIV treatment side effects coping skills intervention on antiretroviral medication adherence. Methods HIV+ men and women (N=249) on antiretroviral therapy (ART) with self-reported high levels of ART side effect distress were randomized to intervention or treatment as usual. The primary outcome was self-reported ART adherence as measured by a combined 3-day and 30-day adherence assessment.
Results Intent-to-treat analyses revealed a significant difference in rates of nonadherence between intervention and control participants across the follow-up time points such that those in the intervention condition were less likely to report nonadherence. Secondary analyses revealed that intervention participants were more likely to seek information about side effects and social support in efforts to cope with side effects. Conclusions Interventions focusing on skills related to ART side-effects management show promise for improving ART adherence among persons experiencing high levels of perceived ART side effects.
While the life-extending benefits of antiretroviral therapies (ART) for human immunodeficiency virus (HIV) are welldocumented, aversive side effects accompany drug benefit [1]. Side effects are predictable, undesirable, and dose-related pharmacologic effects that occur within therapeutic dose ranges. The most common side effects from ART are gastrointestinal problems such as diarrhea, nausea and vomiting, and dermatological problems such as rashes. Additional "unseen" negative effects that become apparent over time include cardiac and liver problems, bone loss, and increased triglyceride levels [2]. Side effects are often cited when evaluating the impact of ART on the HIV treatment arena [3][4][5]. While newer ART drugs have fewer side effects, the goal of a completely side effect-free, clinically effective regimen has yet to be realized. As such, HIV-positive patients will have to face the realities of side effects in the foreseeable future.
Perceived or anticipated side effects from ART have been linked to failure in timely initiation and maintenance of ART and are a threat to optimal adherence [6,7]. Poor ART adherence is related to virologic failure and resistance, hastened disease progression, increased morbidity and mortality, and elevated health care costs [1,[8][9][10][11][12][13][14][15]. Side effects from ART are consistently found to predict poor drug adherence [15,16] and affect the acceptance and maintenance of ART among those who may benefit from treatment [17,18]. In a sample of 2,765 persons with HIV in four US cities, patients' reports of several specific side effects were associated with an increased likelihood of poor adherence [19]. The link between side effects and nonadherence seems to be one of which HIV+ individuals are aware, with side effects consistently cited as a reason for nonadherence to ART [20].
Side effects create a unique coping challenge. Unlike disease-related symptoms, side effects may be coupled with a belief that the problems are necessary to stay healthy (i.e., they are inevitably tied to the medication). Side effects may also be viewed as ultimately controllable, that is, that one has the power to stop taking medication and consequently eliminate side effects [21][22][23]. Programs to help patients manage ART side effects so that adverse effects do not negatively impact adherence and treatment continuation offer promise. The purpose of the current study was to evaluate the impact of a one-on-one side effects coping intervention on rates of nonadherence among adults living with HIV.
This study was conducted in San Francisco, CA, USA and approved by the Institutional Review Board. Voluntary, written informed consent was obtained from all participants. The trial was registered at clinicaltrials.gov (NCT00643903).
Between February 2005 and March 2007, HIV-infected individuals were recruited from community agencies and medical clinics to participate in a brief in-person interview used to screen participants for eligibility in the randomized intervention trial. To be considered for screening, potential participants were required to be at least 18 years of age, provide written informed consent, be taking a recognized ART regimen (verified by documentation from pharmacy, letter from provider, or examination of prescription bottles) for at least the prior 30 days, and report not being currently involved in another behavioral intervention study related to HIV. Severe neuropsychological impairment and psychosis were assessed on a case-by-case basis by interviewers in consultation with senior project personnel, including the principal investigator, a licensed clinical psychologist. Participants were eligible for enrollment in the trial if they reported a level of side effect distress on a previously used symptom/side-effect checklist equivalent to the upper 40% of a prior sample [21,22]. Those who met criteria were scheduled for an enrollment visit and baseline interview approximately 1 week later.
A double-baseline randomized controlled design of the active intervention compared to treatment as usual was used in this study. The double baseline was employed to observe naturally occurring changes over time and regression to the mean of key variables prior to randomization. Interviews were conducted using laptop computers in private settings in research offices. Procedures involved a combination of audio computer-assisted self-interviewing (ACASI) and computerassisted personal interviewing (CAPI) using the Questionnaire Development System (Nova Research Company, Bethesda, MD, USA). ACASI has been shown to be an effective method of decreasing social desirability bias and thereby enhancing veracity of self-report of sensitive behaviors, including sexual and substance use risk acts [24,25]. The second baseline interview was scheduled 3 months after the initial baseline interview.
Simple randomization was implemented immediately following the second baseline interview using the SAS System's random number generator under the uniform distribution, aligning treatments in order with consecutive participant ID.
The Balance Project experimental intervention was designed on the basis of prior studies of ART side effects and adherence [19,21,22] and was based on elements of social problem solving training [26,27] and coping effectiveness training [28] rooted in Stress and Coping Theory [29]. It consisted of five 60-min individual counseling sessions with each session designed around topics relevant to ART side effects coping. See Table 1 for an outline of the five sessions. Intervention sessions followed a standard structure and set of activities, but were individually tailored to participants' specific life contexts, stressors, and goals. Participants received $30 at the 3-month assessment if they completed all five sessions prior to that assessment interview. Participants in the control condition received no active psychosocial interventions prior to the final trial assessment interview.
Facilitators were master's level clinicians with expertise in HIV-related issues, were trained using standard materials, and were "certified" if supervisors' observations and quality assurance ratings indicated skilled implementation. All intervention sessions were audio recorded and 10% were rated to ensure replication with fidelity.
Follow-up assessment interviews were scheduled at 3 (second baseline interview), 6, 9, and 15 months for both the intervention and control groups. Participants received $25 for completing the screening/enrollment interview, $40 for each of the two baseline assessment interviews and the 6-and 9-month follow-up interviews, and $50 for the 15-month final interview.
Basic demographic, treatment history, and health care utilization data were collected by CAPI with trained interviewers. Depression was assessed with the Beck Depression Inventory II [30].
We assessed ART adherence using two well-validated selfreport measures. The first was the adherence measure designed for the Adult AIDS Clinical Trials Group [31], which assesses missed pills over the prior 3 days. This measure has been used widely with diverse samples and the short-term recall period has been associated with long-term clinical outcomes. The measure was computerized for ACASI administration to minimize the social desirability associated with adherence reporting [31][32][33]. Mean 3-day adherence was calculated by dividing the number of pills reported as being taken by the number of pills that were prescribed in the regimen. Second, we administered the visual analog scale developed by Walsh [34] that assesses 30-day adherence, reporting separately for each drug along a continuum anchored by "0%" to "100%." This measure has shown to be correlated with other measures of adherence, such as medication event monitoring systems [35,36] and a 30-day timeframe has recently been supported as preferable to other approaches of self-report [37]. For the visual analog scale, the mean percent adherence was calculated across all drugs in the participant's regimen. Because there is a tendency for people to over-report adherence on these measures, we defined nonadherence as less than perfect adherence on either the AIDS Clinical Trial Group Measure or the visual analog scale, thus establishing a conservative definition of self-reported nonadherence.
Coping with Side Effects To monitor changes in ways of coping with treatment side effects associated with the Balance Project intervention, we administered the SECope at each assessment timepoint [38]. This 20-item measure assesses strategies for coping with HIV-treatment side effects, and includes scales of Positive Emotion Focused Coping, Social Support Seeking, Nonadherence, Information Seeking, and Taking Side Effect Medications, all with evidence of reliability and validity from prior studies [38].
One-way and cross-tabular frequency tables were generated for categorical variables; means and standard deviations were generated as measures of central tendency for continuous variables. Primary inferential analyses consisted of random coefficient multilevel (i.e., HLM) models that contained random intercepts or random intercepts and slopes to model within-participant correlations of responses over time, providing a unified method to model both binary (i.e., nonadherence) and continuous outcomes (i.e., the coping measures). Each model contained fixed main effect terms for group assignment (intervention vs. control), time of measurement (treated as a continuous variable with measurement points of baseline, 3, 6, 9, and 15 months), and their interaction. Random effects consisted of participant-specific intercepts or participant-specific random intercepts and slopes. Models involving the continuous SECope outcomes also added a single level 1 residual variance estimate, which is customary in random coefficient modeling. Additional within-group slope coefficients and confidence intervals were produced for models that exhibited a statistically significant group-by-time interaction effect. For each outcome, a model with random intercepts only was compared to a model with random intercepts, slopes, and their covariance using the Bayesian Information Criterion [39].
Random effects models were fitted using SAS PROC GLIMMIX version 9.2. Parameter estimation was obtained through maximum likelihood estimation with integral approximation via adaptive Gaussian quadrature with 15 integration points. To guard against possible misspecification of the model covariance structures influencing inferences, variance estimation was performed using Morel's robust variance estimator [40]. Linearity of the continuous time effect was assessed using the cumulative-sums-of-residuals method [41]. Histograms and predicted value-by-residual scatterplots were used to evaluate normality and homoscedasticity of the model residuals, respectively.
Table 2 provides details of the participant characteristics for each randomization group. The sample was predominately male, with 55% identifying as White, 18% as African American or Black, and 15% as Hispanic. The mean age was 46 years, the mean time since testing HIV positive was almost 14 years, and the mean time since starting a first ART regimen was 10 years. The most frequently endorsed symptoms that were attributed to ART and the proportion of respondents reporting them were as follows: fatigue or loss of energy (96.4%), feelings of sadness and depression (82.3%), sleep problems (78.6%), muscle aches or joint pain (75.5%), and stomach bloating, pain, or gas (74.7%). The sample was balanced across randomization conditions with regard to all demographic, treatment, side effect, and depression variables.
Of the 249 individuals enrolled in the trial and randomized, 95% completed the 6-month assessment, 93% completed the 9-month assessment, and 93% completed the final 15month assessment. Of the 128 individuals randomized to the intervention condition, 116 (91%) completed the first session and 112 (88%) completed all five sessions. See Fig. 1 for the trial participant flow details. Attrition rates across the two conditions were not different, suggesting that the extra incentive payment for intervention completion did not differentially affect study retention.
Assessments of linearity via cumulative sums of residuals supported the null hypothesis of linear time effects for all outcomes (p>0.10 for all assessments). Examination of residuals' histograms and predicted-value-by-residual-value scatterplots indicated approximate normality and constant variance of model residuals. For all outcomes studied, the Bayesian Information Criterion preferred the more parsimonious random intercepts model to the more complex random intercepts-plus-slopes model. Consequently, all results reported below originate from random intercept-only models.
A statistically significant group-by-time interaction effect was obtained, such that control group and intervention group participants differed in non-adherence change over time (see Table 3 and Fig. 2). Specifically, within-group slopes analysis revealed that the odds of non-adherence
544 Assessed for Eligibility 295 Excluded 210 Not Meeting Inclusion Criteria 85 Declined 128 Assigned to Intervention 116 (91%) Received Intervention as Assigned 249 Randomized 121 Assigned to Wait List Control 10 Lost to Follow-up 4 Discontinued Intervention 8 Lost to Follow-up 128 Included in Analysis 0 Excluded From Analysis 121 Included in Analysis 0 Excluded From Analysis Coping with Side Effects No statistically significant group, time, or group-by-time effects were observed for the SECope Positive Emotion Focused Coping, Taking Side Effect Medications, or Non-Adherence subscales (Table 3). A significant overall difference between control and intervention group participants was found, however, such that control participants reported using side-effects management strategies less than their intervention group counterparts. Statistically significant group-by-time interaction effects were obtained for the Social Support and Information Seeking subscales, such that control group participants reported using these strategies less often than intervention participants did during the study (Table 4). Within-groups slopes analyses showed no change in intervention participants' use of social support (B=0.001; 95% CI=-0.01, 0.01; p=0.83) and information seeking (B = 0.002; 95% CI = -0.01, 0.01; p = 0.73). However, control participants' seeking of both social support (B=-0.02; 95% CI=-0.03, -0.01; p=0.001) and information (B=-0.02; 95% CI=-0.03, -0.01; p=0.007) decreased over time (Tables 3 and 4). N=249 for all analyses. Intercept represents the expected value of the intervention group. Group represents the difference between the control group and intervention group. Time is measured in months (0.25, 3, 6, 9, and 15) and represents the intervention group's change in the outcome over time. Group×time represents the difference in slopes between the control and intervention groups. Non-Adh non-adherence. SECope P positive emotion-focused coping, SECope N non-adherence strategies, SECope T taking side-effects medications, SECope S social support seeking, SECope I information seeking *p<0.05, **p<0.01; ***p<0.001
Note: Intervention was delivered between months 3 and 6. Overall group x time interaction was significant (p<.05)
The findings from the current trial support the five-session Balance Project intervention to promote ART adherence among HIV-positive adults experiencing high levels of perceived ART side effects. Follow-up analyses suggest that the intervention may have been particularly effective in influencing individuals' efforts to access information and social support for coping with HIV treatment side effects. This is the first published trial to demonstrate beneficial effects on ART adherence by focusing on side-effects management. Consistent with the evidence that side effects are associated with nonadherence, these results demonstrate that efforts to improve patients' side-effects management skills have the potential to reverse the negative impact of perceived side effects on ART adherence.
The current intervention, although individually delivered, was relatively low dose compared to other behavioral interventions in health care contexts, which sometimes prescribe 15 or more sessions [42,43]. A low-dose intervention may be readily implemented in clinics and agencies that provide health care and support to persons living with HIV. There is also the potential for elements of the intervention to be delivered prior to the initiation of ART. Such interventions may offset the harmful effect of anticipated side effects on future rates on ART uptake by giving sideeffects management skills to people prior to initiating therapy. Preemptive intervention may be considered in the context of building patients' readiness for ART and may result in greater ART uptake and subsequent adherence and maintenance of ART in the face of side effects that may develop. This intervention approach may also be particularly useful if there is a need to change ART regimens following treatment failure, a context in which there may be a higher perception or actual increased probability of significant side effects during the initial period of a new regimen.
Although designed and implemented in the context of HIV disease, the intervention is rooted in broader theories of health promotion and behavior change, including Stress and Coping Theory [29] and Social Problem Solving Theory [26,44]. Consequently this intervention approach may be appropriate for other illness contexts in which side effects from treatment represent a substantial barrier to adherence and optimal outcomes.
There are several noteworthy limitations in the current study. First, the experimental design used a treatment as usual comparison rather than a matched attention control condition, so it is impossible to determine the potential confounding effects of increased attention by trial staff in the experimental condition rather than the control condition. Further, the assessment did not capture comprehensive data on access and use of adherence resources outside of the trial, a construct that is gaining attention in adherence research [45]. Second, the use of self-reported adherence data, while supported by validity studies [46] has been questioned in HIV research as being inflated because of recall, social desirability, and other biases. To minimize the biases inherent in self-reported data, we employed several techniques. We used conservative cut-offs of validated measures of adherence that have demonstrated meaningful relationships with important outcomes, such as viral load, in other studies. ACASI interviewing for the adherence portion of the interview was used, thereby removing the interviewer's presence and minimizing social desirability bias. Such approaches to computerized adherence assessment have shown favorable effects in other studies [47]. Despite these strategies, it should be noted that the use of self-reported adherence outcomes remains a limitation of the study, and results may have differed if other adherence assessment approaches had been used. Third, because the study used a convenience sample rather than a probabilitybased sample, there are limits of the degree to which findings are generalizable to other populations. Related, the sample was predominately male, and thus caution should be exercised when generalizing the findings to women. Finally, because the eligibility criteria did not select based on levels of adherence at baseline, there was a limit of the degree to which patients could improve on their adherence, as some reported perfect adherence at study entry. However, even with the potential ceiling effect of adherence scores in the sample, the intervention was shown to significantly protect against the likelihood of subsequent nonadherence. In summary, this study supported the hypothesis that improving coping skills related to HIV treatment side effects can have a protective effect on treatment adherence. While the side-effect burden associated with available HIV treatments has lessened over recent years with the development of new drugs, a truly side effect-free ART regimen has not yet been developed. Therefore, treatment side effects are likely to remain a substantive threat to adherence. Interventions aimed at mitigating the impact of side effects on treatment adherence offer promise to help optimize treatment outcomes for the growing numbers of people living with HIV.
Review of prior session and progress toward goal Emotion vs. problem focused coping in HIV treatment Social support skills Distinguish tangible, emotional and informational support Identify positive vs. negative social support Explore/Diagram current social support network Problem solve social support building around side effect stressors
a All p values comparing control to intervention were >0.
ann. behav. med. (2011) 41:83-91
ann. behav. med. (2011) 41:83-91
Acknowledgments This work was funded by grant
The authors have no conflict of interest to disclose.
Open Access This article is distributed under the terms of the Creative Commons Attribution Noncommercial License which permits any noncommercial use, distribution, and reproduction in any medium, provided the original author(s) and source are credited.
Impaired balance control during gait can be detected by local dynamic stability measures. For clinical applications, the use of a treadmill may be limiting. Therefore, the aim of this study was to test sensitivity of these stability measures collected during short episodes of over-ground walking by comparing normal to impaired balance control. Galvanic vestibular stimulation (GVS) was used to impair balance control in 12 healthy adults, while walking up and down a 10 m hallway. Trunk kinematics, collected by an inertial sensor, were divided into episodes of one stroll along the hallway. Local dynamic stability was quantified using short-term Lyapunov exponents (k s ), and subjected to a bootstrap analysis to determine the effects of number of episodes analysed on precision and sensitivity of the measure. k s increased from 0.50 ± 0.06 to 0.56 ± 0.08 (p = 0.0045) when walking with GVS. With increasing number of episodes, coefficients of variation decreased from 10 ± 1.3% to 5 ± 0.7% and the number of p values >0.05 from 42 to 3.5%, indicating that both precision of estimates of k s and sensitivity to the effect of GVS increased. k s calculated over multiple episodes of over-ground walking appears to be a suitable measure to calculate local dynamic stability on group level.
Recently, there has been a growing interest in quantifying gait stability using a method based on nonlinear time series analysis, i.e., local dynamic stability. [11][12][13] Local dynamic stability is defined by ''maximum finite-time Lyapunov exponents'', which describe how the system's states respond to very small perturbations continuously in real time. 10,13,28 It has been suggested that dynamic stability of gait can be of clinical use, 30 specifically in identifying elderly at risk of falling, since it was shown to differentiate between older adults with and without a history of falls. 24 Precise estimation of Lyapunov exponents theoretically requires at least 10 n data points, with n being the dimension of the attractor, and improves with the number of ''cycles'' captured by those data points. 28 In a recent study, we showed that for human gait data, limited increases in precision were obtained when using data series longer than 150 strides (sampled at 50 samples/s). 6 A convenient method of continuously collecting data over many successive strides is to use a treadmill. Unfortunately, such large number of strides may not be feasible for frail elderly or patient populations and treadmills are relatively expensive instruments. Recently it was shown that inertial sensor data can be used as an alternative to optoelectronic data for calculation of Lyapunov exponents of trunk kinematics during gait. 5 This would allow for data collection in the typical clinical setting where gait is assessed on a short walkway. Using inertial sensor data collected during over-ground walking, where speed can not be controlled exactly and where the Lyapunov exponent is calculated over multiple short episodes, would make this measure more applicable, e.g., to diagnose elderly at risk of falling. However, it is unknown if this would compromise the precision of the measure and its ability to detect impaired balance control (sensitivity). In addition, although statistical precision of estimates will by definition increase monotonically with sample size (i.e., more episodes), the magnitude of the increase in Address correspondence to Jaap H. van Diee¨n, Research Institute MOVE, Faculty of Human Movement Sciences, VU University Amsterdam, Van der Boechorststraat 9, 1081 BT Amsterdam, The Netherlands. Electronic mail: j.vandieen@fbw.vu.nl precision cannot be predicted in a straightforward manner, such as in sampling from a normally distributed population, because consecutive episodes may not represent independent samples. Sensitivity of local dynamic stability measures estimated from multiple short episodes of over-ground walking at preferred walking speed can be investigated by comparing normal to impaired balance control. Balance control can be impaired in young adults by randomly varying galvanic vestibular stimulation (GVS). Specifically, the firing rates of the vestibular afferents are decreased or increased by the anodal and cathodal currents of GVS, 2 resulting in an observable, adjustable sway in the lateral direction during walking. 3,16 It is shown that in subjects facing forward, bipolar binaural stochastic GVS leads to coherent stochastic mediolateral postural sway. 26 Therefore, stochastic GVS is suggested to quantitatively and qualitatively model instability of vestibular origin. 25,26,29 Recently, it was shown that local dynamic stability measures can detect the impairment of balance control induced by GVS during treadmill walking at several walking speeds, including preferred walking speed, in young adults. 33 The present study used bootstrap analyses to quantify the increase in precision of estimates of local dynamic stability as a function of number of episodes used for analysis. In addition, we tested the hypothesis that GVS would cause a decrease in local dynamic stability in over-ground walking. Finally, we quantified the effect of number of episodes analyzed on sensitivity to the GVS-induced balance impairment at the group and individual level.
Ten subjects (7 males and 3 females; mean age 23 ± 2.6 years; height 1.8 ± 0.1 m; and mass 76.3 ± 12.5 kg) participated in the study. Exclusion criteria were any orthopedic or neurological disorders, or other injuries that could interfere with gait. Alcohol consumption was prohibited for 24 h prior to testing. Subjects filled in a medical intake form and signed informed consent. The protocol was approved by the Local Ethical Committee.
Portable wireless inertial sensor nodes (Xsens, Xsens Technology, Enschede, The Netherlands), which include 3D gyroscopes and accelerometers, were attached at the back of the trunk over the spine at the level of T6. Data-acquisition (50 samples/s) was done with MT Manager (version 1.5.0, Xsens Technology, Enschede, The Netherlands), and analyses were performed using MATLAB (version 7.5, The MathWorks BV, Natick, MA, USA).
To apply GVS, flexible, carbon electrodes were attached with electroconductive adhesive gel (Tac Gel TM ) over the mastoid bones. GVS was applied binaurally and bipolarly to each subject by a computer controlled galvanic stimulator (IDEE, Maastricht University, Maastricht, The Netherlands). To prevent adaptation, the galvanic stimulus was composed of a linear summation of sinusoids with five different frequencies (0.02, 0.07, 0.11, 0.30, and 0.50 Hz), all starting at a phase of 0 degrees and each having a maximum amplitude of 0.6 mA. This pseudo-random stimulus had a maximum amplitude of about 2.2 mA, irrespective of the electrical conductivity of the skin and temporal bone.
Prior to the measurement, the reaction of subjects to GVS was tested, although no negative psychological or mental side effects are known from literature, 2,19 following which the subjects were familiarized with the sensation of GVS for about 20 s. Some subjects reported a slightly dizzy feeling or a weak stinging under the electrodes. The dizziness disappeared after a few seconds and the stinging was remedied by reattaching the electrode.
Measurements were performed during over-ground walking at preferred walking speed, walking up and down an approximately 3 m wide and 10 m long hallway. Each trial, both with and without GVS, lasted 3 min and included on average 20 turns. Participants were instructed to look straight ahead, since GVS induces a body sway lateral to the orientation of the head, regardless of the body orientation. 34
For each trial, time series of 9000 samples (3 min) were obtained. These data were analyzed without filtering, because of the complications associated with applying linear filters to nonlinear signals. 11,12 For each subject, the data were divided into episodes, each episode representing a one-way stroll of 7 strides along the hallway without a turn. The walking speed was calculated for each episode, based on the known walking distance and time elapsed between turns. To exclude the influence of data series length on local dynamic stability measures, 6,15 for each episode a fixed number of 7 continuous strides was time normalized by shape-preserving piecewise cubic interpolation to 350 total samples, i.e., an average of 50 points per cycle.
Note that this implied almost no creation or reduction of data points given a sampling rate of 50 samples/s and an average stride time of 1.08 s.
The calculation of maximum Lyapunov exponents to qualify local dynamic stability requires constructing a states space from measured data. Although several studies used five-dimensional state spaces, 6,10,15 based on embedding reconstruction methods 31 we preferred a biomechanical state space consisting of 12 dimensions. This 12D state space was reconstructed from the 3D angular velocities along with the 3D linear accelerations of the trunk and their time-delayed copies. 5,18 A standard time delay of 12 samples was used for the time-delayed copies, which roughly corresponds to the derivatives. This approach was first proposed by Kang and Dingwell, 20 who reasoned that a 3D free body has a total of 12 state variables (i.e., 3D positions, velocities, orientations, and angular velocities). Recently, it was shown that such an approach yielded results closest to the expected value for a known system, and was most sensitive to differences between conditions. 17 Maximum Lyapunov exponents represent the exponential rate of divergence from nearest neighbor points on adjacent trajectories in the reconstructed state space, and thus the response to small local perturbations of the system. 11 Positive exponents indicate local dynamic instability, with larger exponents demonstrating an increased sensitivity to local perturbations, 11 while values <0 indicate local dynamic stability.
To calculate maximum Lyapunov exponents, Euclidian distances between neighboring trajectories in state space were computed as a function of time and averaged over all original pairs of initial nearest neighbors. The short-term Lyapunov exponent (k s ) was calculated from the slopes of linear fits to the first 0-0.5 stride of the log of these divergence curves, rather than to 0-1 stride, since the divergence curve has been found to be nonlinear after about 75 samples (Fig. 1). 7 In addition, a recent study has demonstrated that k s based on a single step was related to the probability of falling in a passive dynamic walking model and could therefore be successfully used to predict stability. 8 It should be noted that the region where the slope was fitted may not be fully linear, and the exponent found may thus not be a ''true'' Lyapunov exponent. However, it still provides a well-defined metric for estimating the sensitivity of human walking to small intrinsic perturbations. 10,21 In addition, this k s (i.e., calculated from 0 to 0.5 strides) has been shown to discriminate fall prone elderly from elderly without a history of falling. 24 Bootstrapping A bootstrapping procedure 4 was used to assess the statistical precision of k s as a function of number of episodes. Bootstrap analysis can be used for statistical inferences without requiring assumptions on distribution of the data. For introductory texts see Zhu. 35 In addition, bootstrap analysis can be used to support decision making on sample size (Efron and Tibshirani 14 ), as has been done for biomechanical variables previously (e.g., van Diee¨n et al. 32 ; Bruijn et al. 6 ). These previous applications usually considered nested bootstrap analyses which allow inferences on effects of for example number of subjects measured as well as number of measurements per subject. In the present study bootstrap analyses were performed 'within subjects', since our main focus was on clinical application. First, k s was calculated for each subject and for each of the 16 episodes of 7 strides per condition. Next, for every subject, a sample of a number (n) of randomly selected episodes was drawn (with n ranging from 3 to 12) for both conditions. A maximum sample size of n = 12 out of 16 episodes was used, to avoid too much overlap between the bootstrap samples and thus an overestimate of the statistical precision. This procedure was repeated 1000 times for every n.
To express statistical precision of the estimate of k s , the coefficient of variation (COV), defined as the standard deviation over the 1000 samples normalized to the mean was calculated for each sample of n episodes.
To examine the sensitivity of k s to GVS-induced balance impairment as a function of number of episodes at group level, a paired t test was performed over sample means of all subjects between the two conditions. In addition, for each individual subject, an unpaired t test between the two samples (one normal walking selection of n episodes, and one GVS selection of n episodes) was performed, to quantify sensitivity at the individual level. Finally, both for the individual and group level comparisons, the percentage of the 1000 t test results with p values above 0.05 was calculated.
All subjects produced the required 16 episodes of 7 strides. The mean walking speed was 1.31 m/s (SD 0.25) when walking with GVS and 1.35 m/s (SD 0.26) without GVS. In addition, the mean standard deviations between episodes within subjects were 0.080 (SD 0.024) with GVS and 0.083 (SD 0.020) without GVS. No significant difference between the two conditions was found for either the mean walking speed (p = 0.149) or the within-subject standard deviation (p = 0.768).
Averaged over all 16 episodes, k s was higher for walking with GVS compared to walking without GVS for most of the subjects (Fig. 2). At the group level, k s significantly increased from a mean of 0.50 (SD 0.06) during normal walking to a mean of 0.56 (SD 0.08) when walking with GVS (p = 0.0045), indicating a decreased local dynamic stability when walking with GVS.
Although GVS affected k s at the group level, this effect was less apparent at the individual level. Whereas walking with GVS showed higher values for most subjects when averaged over all 16 episodes, a substantial variation of k s over episodes existed within each subject for both conditions (see for an example Fig. 3).
Therefore, the effect of averaging k s over a varying number of episodes for each subject on the precision was examined by a bootstrapping procedure. As expected, the COVs decreased monotonically with an increasing number of episodes (Fig. 4). For walking with GVS, the COV decreased from 10% (SD 1.3) when using 3 episodes to 5% (SD 0.7) for 12 episodes. Without GVS, the decrease was from 11% (SD 3.1) to 5% (SD 1.4). The increase in precision above 11 episodes was limited.
In addition, the effect of averaging k s over a varying number of episodes on the sensitivity was studied. At the group level, a clear decrease in number of p values >0.05, indicating an increase in sensitivity, with increasing number of episodes was found (Fig. 5) with percentages of p values >0.05 decreasing from 42.0% when using 3 episodes to 3.5% for 12 episodes.
However, on the individual level, the effect of increasing number of episodes on sensitivity varied strongly (Fig. 6). Five subjects showed less than 20% improvement in sensitivity to GVS when using 12 episodes compared to 3 episodes, whereas the other five subjects showed improvements of 27% up to 78%.
In search for methods that would allow the use of the maximum Lyapunov exponent as a measure of balance impairments in clinical situations, we studied the effects of different numbers of short episodes of over-ground walking data on the precision and sensitivity of k s . We found a significant effect of GVS on k s at group level, with higher values when GVS was applied, indicating that k s was able to detect the imposed balance impairment during over-ground walking. However, this effect was less apparent at the individual level. This might be due to the fact that our healthy young subjects responded differently to GVS, but our statistical design was not aimed to evaluate individual responses to GVS. Nevertheless, we showed that averaging over a larger numbers of episodes led to a considerable increase in precision, leading to sufficient sensitivity at the group level but not at the individual level.
Walking speed influences local dynamic stability of walking. 7,12,15 However, no significant differences in mean walking speed or within-subject variance of speed were found in this study between the two conditions. Therefore, the differences in local dynamic stability found between the two conditions are most likely not the result of variation in walking speed between trials, although speed fluctuations within episodes could not be determined.
Various studies have described the influence of time series length and the number of strides on maximum Lyapunov exponents and have suggested that trial lengths of 150 strides or more are needed for precise estimates of the Lyapunov exponent. 6,21 However, these large numbers of strides are required when one single trial is measured, and both Bruijn et al. 6 and Kang and Dingwell 21 have suggested that measuring several trials may increase precision. The decrease of COV and p values >0.05 with an increasing number of episodes found in the current study indicate that precision and sensitivity can be improved by using multiple episodes, instead of using large numbers of consecutive strides. It should be noted that perhaps not the ''true'' Lyapunov exponent is found, since the region where the slope was fitted may not have been fully linear. However, the used k s proved sensitive to small intrinsic perturbations 10 and able to discriminate fall prone elderly from healthy elderly. 24 A large range of mean k s values (k S = 0.06-3) has been reported in literature. 6,7,10,13 The mean values found in the current study were approximately k s = 0.60 and fall within the same range, but are small compared with most of the reported values of k s previously mentioned. The relative low values may be explained by the small number of strides used in the calculations: nearest neighbors that lay far apart cannot diverge far and in longer data series the nearest neighbor will tend to be closer. 6 In addition, different definitions of the state space, timescales used for the calculation of k s , or number of embedding dimensions may lead to different values of k s . 17
The effect of GVS on k s found in the current study is in agreement with expectations, since GVS is known to cause instabilities in the mediolateral direction. 16,34 The same direction also appears affected by aging, 27 suggesting a similarity between the effects of GVS and of aging on balance control. In addition, the effects of both GVS and aging include higher values of k s . 24,33 According to the literature, the estimated differences in k s between older and young adults are within the range of Dk s = 10-50%. 9,[22][23][24] When comparing fall prone and healthy elderly, the estimated difference is Dk s = 10-20%. 18,24 In the current study and the previously mentioned study of van Schooten et al., 33 the mean difference due to the effect of GVS was Dk s = 10%. Although walking speed and conditions (treadmill and over-ground walking) varied in the above-mentioned studies, 1,7,13,15 this suggests that instabilities induced by GVS are probably smaller than those induced by aging and that age-related effects are even more likely to be detected by maximum Lyapunov exponents.
The purpose of the study was to find a clinically applicable method to detect balance impairments using local dynamic stability measures. The results showed that an increase in precision of k s may be gained by measuring multiple short trials during normal overground walking. Therefore, no additional equipment is needed to study walking stability in a clinical setting and limitations caused by low endurance of patients can partially be overcome by using multiple short walking trials. Moreover, the measure can be assessed with a single inertial sensor on the trunk, which is portable, small, reasonably cheap, and straightforward to utilize, since it does not have to be aligned with a global coordinate system. 5 However, the resulting precision and sensitivity of k s at 12 episodes was still insufficient to detect the effect of GVS in each individual subject. Nevertheless, since the effect of GVS on gait stability appeared relatively small compared to reported effects of aging and pathology, sensitivity to age-related or pathological balance impairments merits further study.
Overall, the presented effect of using multiple episodes on the precision and sensitivity of k s suggests that this measure is suitable for scientific purposes involving group level analysis. However, future studies should address the sensitivity of local dynamic stability of gait to differences in fall risk among older adults.
0090-6964/11/0500-1563/0 Ó 2011 The Author(s). This article is published with open access at Springerlink.com
The authors acknowledge Warner ten Kate for the support and trial versions of the inertial sensors and
This article is distributed under the terms of the Creative Commons Attribution Noncommercial License which permits any noncommercial use, distribution, and reproduction in any medium, provided the original author(s) and source are credited.
Background and Aims Changes in size inequality in tree populations are often attributed to changes in the mode of competition over time. The mode of competition may also fluctuate annually in response to variation in growing conditions. Factors causing growth rate to vary can also influence competition processes, and thus influence how size hierarchies develop. † Methods Detailed data obtained by tree-ring reconstruction were used to study annual changes in size and size increment inequality in several even-aged, fire-origin jack pine (Pinus banksiana) stands in the boreal shield and boreal plains ecozones in Saskatchewan and Manitoba, Canada, by using the Gini and Lorenz asymmetry coefficients.
The inequality of size was related to variables reflecting long-term stand dynamics (e.g. stand density, mean tree size and average competition, as quantified using a distance-weighted absolute size index). The inequality of size increment was greater and more variable than the inequality of size. Inequality of size increment was significantly related to annual growth rate at the stand level, and was higher when growth rate was low. Inequality of size increment was usually due primarily to large numbers of trees with low growth rates, except during years with low growth rate when it was often due to small numbers of trees with high growth rates. The amount of competition to which individual trees were subject was not strongly related to the inequality of size increment. † Conclusions Differences in growth rate among trees during years of poor growth may form the basis for development of size hierarchies on which asymmetric competition can act. A complete understanding of the dynamics of these forests requires further evaluation of the way in which factors that influence variation in annual growth rate also affect the mode of competition and the development of size hierarchies.
Size variability in plant populations may be due to differences in competitive status, genetics, the differential effects of herbivores and pathogens (Weiner and Thomas, 1986), or to spatial and temporal environmental heterogeneity (Schwinning and Weiner, 1998;Wichmann, 2001). In tree populations, size variability also contributes to the structural diversity of a forest stand, which is important for many ecological functions (Brassard and Chen, 2006). Studies of size variability in tree populations have focused mainly on fitting various probability distribution functions to size (diameter) distributions [as noted in Garcia (2006), there are hundreds of papers concerning diameter distribution models in forestry literature databases]. However, size variability in tree populations can also be described as a size hierarchy, and so can be described by other characteristics such as its degree of size inequality (Weiner and Solbrig, 1984). Changes in size inequality are often attributed to changes in the mode of competition during different stages of stand development (e.g. Gates et al., 1983;Weiner and Thomas, 1986;Newton and Smith, 1988;Kenkel et al., 1997). Previous studies have observed that inequality is greater at higher densities (Brand and Magnussen, 1988;Knox et al., 1989), increases prior to self-thinning and decreases as self-thinning progresses (Mohler et al., 1978;Knox et al., 1989). In addition, although previous studies have examined the relationship between size and size increment [also known as the distribution modifying function (Westoby, 1982;Weiner, 1990;Weiner and Damgaard, 2006)], there has been little attention placed on examining the inequality of size increment itself.
Inter-tree competition is considered to be either a resource pre-emption (size asymmetric) process or a resource depletion (size symmetric) process. Immediately after stand initiation, individual trees are small in comparison with their relative density, so if competition exists at all, its mode is symmetric. Over time, as trees grow larger and a size hierarchy begins to develop, the mode of competition is thought to become asymmetric as larger trees pre-empt light from smaller trees. Tree size at any given point contains a 'memory' of the processes that influenced that individual as it grew from a smaller size, and therefore changes in the inequality of tree size should be best explained by long-term changes in stand characteristics, such as density, mean tree size and the average amount of competition. In contrast, the size increment of individual trees varies greatly from year to year in response to transient factors such as annual variation in weather and insect defoliation (e.g. Larsen and MacDonald, 1995;Brooks et al., 1998;Hofgaard et al., 1999;Hogg et al., 2005;Hogg and Wein, 2005). These factors may have different effects on large and small individuals (Orwig and Abrams, 1997;Piutti and Cescatti, 1997;Wichmann, 2001). Therefore, the inequality of size increment may be more strongly related to annual variation in stand-level growth rate, which can be considered a surrogate variable accounting for the transient environmental factors affecting a stand. The influence of these transient factors on the mode of competition between plants is, at present, not well understood (Schwinning and Weiner, 1998).
The present study uses detailed growth data obtained from tree-ring reconstruction to investigate annual changes in the inequality of size and size increment in four even-aged fire origin jack pine (Pinus banksiana) stands. When combined with cross-dating of recent and historical mortality, tree-ring reconstruction can give annual data on the size and growth rate of individual trees (e.g. Henry and Swan, 1974;Oliver and Stephens, 1978;Johnson and Fryer, 1989;Stoll et al., 1994;Carrer and Urbinati, 2001). Although intensive to collect, these data can be advantageous because they follow the growth and mortality of individuals over time (Weiner, 1995) and are at an annual resolution. We hypothesize that at any given time, tree sizes are more equal than tree size increments, and that annual trends in inequality will be more variable for size increment than for size. We also hypothesize that, because sizes change slowly, the inequality of tree sizes will be best predicted by long-term changes in stand population parameters such as stand density, mean tree size and the average amount of competition. In contrast, because the inequality of tree size increment is more variable, we hypothesize that it will be best predicted by short-term population parameters such as stand-level annual growth increment.
Plots were established in fire-origin jack pine (Pinus banksiana Lamb.) stands on sandy soils located in the boreal forest of western Canada (Fig. 1). The study sites were located near (1) Candle Lake, Saskatchewan (53 . 98N, 104 . 78W) and (2) Thompson, Manitoba (55 . 98N, 98 . 68W) (Fig. 1). The Candle Lake sites were in the Boreal Plains ecozone, the Thompson sites in the Boreal Shield ecozone (Ecological Stratification Working Group, 1996). Based on spatially interpolated climate normals for the period 1970period -2000period (McKenney et al., 2006)), the mean annual temperature was 20 . 3 8C at Candle Lake and 22 . 8 8C at Thompson. The mean annual precipitation was 466 mm at Candle Lake and 516 mm at Thompson.
In each region, a plot was sampled in a region on a mesic (relatively nutrient-rich) site and a xeric (relatively nutrientpoor) site, determined on the basis of ecological classification and indicator species. The plots were sampled in the summer of 2005, so the last complete year of growth observed was 2004. The polar coordinates of each living tree, standing dead tree and lying log in a 900-m 2 (30 Â 30-m) area were mapped using a surveying transit and tape measure. Height was measured for living trees, and breast height diameter for all trees. Two randomly orientated increment cores were extracted at breast height from living trees and a cross-sectional disc was cut from dead trees. Some (n ¼ 25) trees at each site were also cored near ground level to estimate stand ages. The samples were a complete census of all living and dead trees recognizable at the time of sampling. Study plot characteristics are summarized in Table 1.
The samples were air-dried, the cores were mounted on grooved boards and cross-sectional discs were cut into 1 -2-cm-thick slices. These were polished with up to 600 grit sandpaper, scanned as 1600-dpi greyscale images and imported into WinDendro (Regent Instruments, Quebec, Canada) for ring width measurement. When suppressed, jack pine can form light rings (Volney and Mallett, 1992) that were not always visible on the scanned images. Simultaneously, suppressed samples were examined with a microscope and rings not visible on the scanned images were added to the WinDendro file. Trees were considered to be functionally dead when radial growth ceased at breast height (Mast and Veblen, 1994), which may have underestimated year of death in some cases of extreme suppression (,5 % of samples). Year of death was determined by cross-dating against a master chronology developed from a sample (n ¼ 25) of the largest trees at each site. Samples were cross-dated visually by reference to narrow marker years (Yamaguchi, 1991) induced by periodic jack pine budworm (Choristoneura pinus Freeman) defoliation (Volney, 1988). Dating accuracy was checked by calculating the correlation between the raw ring widths on a sample and the raw ring widths on the site master chronology, as well as shifting sample dates +1 -5 years. This was done iteratively until most (82 % living, 76 % dead) samples had the highest correlation at the final assigned date (91 % + 1 year for living trees, 90 % + 1 year for dead trees). The average correlation between a dead sample and the master chronology at the final assigned date was R 2 ¼ 0 . 88 (s.d. ¼ 0 . 09, range ¼ 0 . 76 -0 . 99, n ¼ 429). For living trees, it was R 2 ¼ 0 . 81 (s.d. ¼ 0 . 19, range ¼ 0 . 58 -0 . 99, n ¼ 536). The study sites were even-aged, so visual cross-dating was sufficient to date samples confidently, and the correlation tests were used only to identify gross errors. Some samples were too decomposed to measure ring widths. In these cases, the mean year of death of the three largest and three smallest trees nearest in diameter, of the same class (snag or lying log), and at the same plot as an excessively decomposed tree was used as an estimate of its year of death, and the mean ring widths of these same trees were used as an estimate of its growth. In a test, this method was unbiased with a mean absolute difference of +3 . 3 years between the true and estimated year of death (Metsaranta et al., 2007). Jack pine snags remain standing long enough and lying logs decompose slowly enough that these techniques can reliably reconstruct growth in these forests for up to 50 years into the past (Metsaranta et al., 2007).
Stemwood volume and volume increment were used to describe tree size and size increment. The diameter (inside bark D ib ) of each tree was determined annually from the ring-width measurements and this was used to estimate cumulative volume and volume increment using three equations. First, diameter inside bark (D ib ) was converted to diameter outside-bark (D ob ) using Husch et al. (2003).
From data in Halliwell and Apps (1997), k was estimated to be 0 . 964 (n ¼ 221, r 2 ¼ 0 . 99). Second, heights were predicted using the Chapman -Richards function:
where H is tree height (m) and D is outside bark diameter (cm). From data in Halliwell and Apps (1997) and the plots in the present study, the parameters were estimated to be a ¼ 18 . 87, b ¼ 0 . 11 and c ¼ 1 . 48 (r 2 ¼ 0 . 98) for Candle Lake, and a ¼ 19 . 65, b ¼ 0 . 08 and c ¼ 1 . 56 (r 2 ¼ 0 . 98) for Thompson. Third, volume was determined from H and D using the taper equation of Kozak (1988):
where the components of the equation are as defined in Kozak (1988). Parameters for Candle Lake were obtained from Ga ´l and Bella (1994), and for Thompson from Klos (2004). The total volume of each tree was determined using numerical integration, and all individual tree values were summed to obtain total stand volume. Individual tree volume increment was obtained by subtracting volume in year (y 2 1) from volume in year y. Wholestand volume increment was obtained by summing individual tree values and was expressed in units of m 3 ha 21 year 21 . The accuracy of these scaling methods was tested for a variety of species, height and volume estimation methods and it was found that they usually predict volume increment with a mean error of less than 5 % and always predict volume increment with a mean error of less than 10 %, relative to volume increment obtained by full stem analysis.
The level of competition to which each tree was subject over time was found by annually determining a distanceweighted absolute size index of competition for each tree. The index is similar to Hegyi's (1974) relative size index, but uses the absolute size of competitors rather than weighting them by the size of the subject tree. The absolute size of competitors may be a better measure of competition than relative size (Ramseier and Weiner, 2006). The index was calculated as where C i is the index for subject tree i, D j is the diameter of competitor j, d ij is the distance between subject tree i and competitor j, and N i is the number of competitors for subject tree i. As trees grow, the definition of which trees compete with each other changes, so the competitor search radius for each tree was made to be temporally variable based upon an estimate of its crown width. Crown width was estimated from diameter using
where W is crown width (m) and D diameter (cm). From the data in Halliwell and Apps (1997), the parameters were estimated to be a ¼ 0 . 353 and b ¼ 0 . 682 (n ¼ 235, r 2 ¼ 0 . 69). The search radius for competitors was defined as 3 . 5 times the crown width in a given year (Lorimer, 1983), and the index was calculated only for those trees where the search radius did not fall outside of the plot. At each site, the average amount of competition that trees in each plot were subject to in each year was calculated, and this was used as a predictor of the trends in size and size increment inequality. To help interpret competition effects, trends in the variability of competition to which trees at each site were subject were also determined by calculating the coefficient of variation (CV %) of the competition index.
The Gini coefficient (Weiner and Solbrig, 1984) used to describe annual changes in the inequality of size and size increment. Annual changes in the Lorenz asymmetry coefficient to were also calculated to establish whether the observed trends in inequality were primarily due to large or small trees (Damgaard and Weiner, 2000). The Gini coefficient is the difference between the sample Lorenz curve and the line of perfect equality, where the Lorenz curve is a plot of the cumulative number of individuals (x-axis) against the cumulative proportion of their total size ( y-axis). The Gini coefficient ranges from 0 to 1, where 0 indicates perfect equality (the size or growth of all individuals is the same, or the amount of competition each tree is subject to is the same) and 1 indicates perfect inequality (one tree contains all of the size or growth, or faces all the competition). It can be calculated from data ordered by increasing size as (Dixon et al., 1987)
The Lorenz asymmetry coefficient (S) summarizes the degree of asymmetry in a Lorenz curve. This is important because populations with different Lorenz curves can have the same Gini coefficient, depending on whether most of the inequality is due to large or small individuals (Weiner and Solbrig, 1984). The Lorenz asymmetry coefficient is defined as the point at which the slope of the Lorenz curve is parallel to the line of equality. It is defined as
and is calculated using the following three equations from Damgaard and Weiner (2000):
When S . 1, the inequality present is due mostly to a small number of very large individuals. When S , 1, the inequality present is due mostly to a large number of very small individuals. Coefficients and 95 % confidence intervals were calculated from 1000 bootstrap samples (Dixon et al., 1987) for each set of annual data on size and size increment at each plot.
We examined how changes in the inequality of size and size increment over time are affected by long-term population parameters (age, stand density and competition) and short term population parameters (stand-level annual volume increment). The temporal development of the Gini coefficients for volume and volume increment at each site was modelled using stand density (DENS), mean tree volume (SIZE), the average competition index (COMP), and stand-level annual volume increment in the current year (AVI) and one year previously (AVI1) as predictor variables. Multiple linear regression was used to estimate the parameters and to determine the significance of each of these variables as predictors of annual changes in size and size increment inequality. The LM function in the STATS package for the R Statistical System (R Development Core Team, 2007) was used to perform these calculations. As the data were time series, the potential confounding effects of serial autocorrelation were examined by using generalized least squares to estimate the parameters with the GLS function in the NLME package (Pinheiro et al., 2007) for R, assuming that the residuals followed a first-order autoregressive (AR1) error structure. The parameter estimates obtained by GLS were not substantially different from those obtained by ordinary least squares, so only the results obtained by ordinary least squares are presented. In general, stand density, mean tree size and average competition were expected to be significant predictors of changes in both size and size increment inequality. It was also expected that stand-level annual volume increment would be a significant predictor of inequality of size increment, but would not be a significant predictor of inequality of size.
From 1950 to 2004, the Gini coefficient for size was nearly always less than 0 . 5, meaning that tree sizes could be characterized as equal (Fig. 2). During the same time period, the Gini coefficient for size increment was also generally less than 0 . 5, but there were periods at all sites when it was greater than 0 . 5, indicating that size increment could often be considered more unequal than equal (Fig. 2). Size inequality generally declined over time. Based upon the 95 % confidence intervals, the nutrient-poor sites had more unequal tree sizes. Size increment was more unequal than size, and its inequality was also more variable from year to year than inequality in size, showing both increasing and decreasing trends from year to year, depending upon the site. Inequality in size increment was not different at rich and poor sites.
The Lorenz asymmetry coefficient for size increment was also more variable from year to year than the Lorenz asymmetry coefficient for size (Fig. 3). For the vast majority of the time, the Lorenz asymmetry coefficient for size was not significantly different from 1, indicating that the observed inequality in tree size was not due to either large or small trees. The Lorenz asymmetry coefficient for size increment, however, had many periods of time when it was significantly less than 1 at all sites, indicating that the observed inequality in size increments was often due to larger numbers of trees with small size increments. There was one clear exception to this trend. At the nutrient-rich site at Candle Lake, the two years (1966 and 1967) with Lorenz asymmetry coefficients greater than 1 correspond to the years with the lowest growth rate at that site, and also to two years during which the historical records of the Canadian Forest Insect and Disease Survey indicate that this area was subject to a jack pine budworm defoliation event. In this specific case, the observed inequality in growth rates at this stand was due to a small number of trees that had high growth rates, most likely because they were not defoliated and continued to grow at a normal rate.
The average amount of competition (standardized by the maximum average annual competition observed at a given site so that each plot could be compared on the same scale) showed both increasing and decreasing trends over time at each plot (Fig. 4). Initially, average competition increased at each site up to about 1970 (Fig. 4). After that point, it stayed relatively constant at the nutrient-rich sites, while the nutrient-poor sites showed a second increase in average competition that started about 15 years later (about 1985, Fig. 4). Overall, the increase in average competition from its minimum value was higher at nutrient-poor sites (where the minimum value was 0 . 4-0 . 6 times the maximum) than at nutrient-rich sites (where the minimum value was 0 . 7-0 . 8 times the maximum). The CV % for competition ranged from 20 to 50 % at all sites (Fig. 4), indicating that even though the average amount of competition that trees were subject to changed over time, the amount that each tree was subject to in a given year tended be similar. In addition, although there were periods of time at each site where the CV % for competition had small increasing or decreasing trends, the value of the CV % at any given site ranged only in the order of +10 % over the whole study period, indicating that the variability in the amount of competition to which trees were subject to did not change substantially over time.
Size inequality was well described by long-term changes in stand dynamics. With the exception of the nutrient-rich site at Thompson, density, mean tree size and average competition were significant predictors of size inequality at all four sites (Table 2). At three of the four sites, size inequality was positively associated with density and mean tree size, and negatively associated with average competition. At the nutrient-poor site at Thompson, size inequality was negatively associated with all three predictors. Only at the rich site at Candle Lake was stand-level annual volume increment (in this case, lagged by 1 year) a significant predictor of changes in size inequality. Variability in the significance and sign of the coefficients associated with the predictor variables indicates that to some extent the specific relationships between these predictors and changes in inequality were site specific.
Some combinations of density, mean tree size and average competition were also significantly associated with changes in the inequality of size increment at all but the rich site at Candle Lake, where only stand-level annual volume increment (in this case lagged by 1 year) was a significant predictor (Table 2). Again, the sign and significance of the coefficients for these predictors varied, indicating that the specific relationship between them and changes in the inequality of growth rate were also site specific. Stand-level annual volume increment was a significant predictor of changes in the inequality of size increment at all four sites. At Candle Lake, the significant predictor was lagged by 1 year, while at Thompson it was not. In all cases, the sign of the coefficients for this predictor (significant or not) were negative, indicating increasing inequality in volume increment when stand growth rates were low. Figure 5 plots the relationship between the inequality of size increment and variation in the stand-level annual volume increment to demonstrate this relationship graphically.
Overall, these results showed that factors influencing the annual growth rate of the stand in a given year were also influencing the inequality in growth rates for individual trees in that year. The inequality in size increment was higher in years with poor growth, indicating that it was the years with poor growth that contributed most to the generation of the size hierarchy in these populations. The average amount of competition to which each tree was subject was a predictor of the inequality of size at all four sites, and of the inequality of size increment at two of the four sites. However, changes in competition, or at least in the way that it was quantified here, were not generally sufficient to explain the observed inter-annual variation in the inequality of size increment. This contention is supported by the fact that the annual growth rate was a significant predictor of inequality in size increment at all four sites. Secondly, there was a high degree of variability in inequality in size increment from year to year, even though the CV % of the competition index showed that each tree was subject to a relatively similar amount of competition in a given year, and that the overall variability in the competition index stayed relatively constant from year to year.
In jack pine, poor growth is probably related to drought or defoliation, the dominant agents of selection on the sandy, nutrient-poor sites on which this species is dominant in this region. In most years, inequality in size increment was primarily due to large numbers of trees with low growth rates (Fig. 3), but was also due to small numbers of trees with high growth rates during some years of low growth rate, at some sites. For example, the period with high values for the Lorenz coefficient of asymmetry for the nutrient-rich site at Candle Lake during the 1960s (Fig. 3) was coincident with a period of defoliation (Volney, 1988), suggesting that defoliation caused inequality in size increment to be due to a small number of trees with large growth rates. The trees with large growth rates during these years probably escaped defoliation, and would be in a position of relative competitive advantage. Defoliation in jack pine causes greater mortality in suppressed than dominant trees (Gross, 1992), and escaping defoliation may be one of the factors that allowed the surviving trees to become dominant.
Previous studies have shown that competitive status of trees affects their response to variation in precipitation. For example, Orwig and Abrams (1997) and Wichmann (2001) have noted that increased water availability benefits large trees more than small trees. In a study of European beech, Piutti and Cescatti (1997) showed that growth was negatively correlated with precipitation under water deficit conditions for small trees, but that growth in large trees showed no relationship with water deficit conditions. Similarly, Orwig and Abrams (1997) showed that, in general, small trees were more severely affected by drought than large trees for a wide variety of species and site types. In contrast to the periods of defoliation, the general trend in this study was that inequality in size increment was mostly due to a large number of trees with small growth rates (Fig. 3). The trees that performed poorly may have been poorly adapted to drought, possibly due to inappropriate genetics, a poor micro-site or shallow rooting depth. During low precipitation periods, better-adapted trees could maintain some growth and become relatively larger compared with poorly adapted trees during these drought years. Similar to the situation for trees that escape defoliation, this advantage would improve their relative competitive status and allow them to become dominant in future years. Size was more equal than size increment, and size inequality was also less variable from year to year. In a given region, size inequality was higher at nutrient-poor than nutrient-rich sites. This conforms to expectation because self-thinning, which acts to reduce inequality by removing the smallest individuals, typically occurs more slowly at nutrient-poor sites. However, this observation may have been confounded by density also being higher at poor sites. The observed greater inequality in size increment also conformed to expectation. In some years, the growth rate for individual trees can be close to zero, which would result in high inequality. By contrast, it is not possible for tree size to be close to zero, so there is less potential for inequality in size. Overall, some combination of stand density, mean tree size and average competition index were significant predictors of size inequality at all sites. These variables all change over time in a highly interactive manner, which was reflected in the variation in the sign and significance of the coefficients for these predictors.
The observed relationships were site-specific and not always consistent with the expectation of increasing inequality of size at high stand densities or high levels of competition. Self-thinning mortality generally decreases size inequality by removing the smallest individuals. However, small changes in the relative position of dead trees in the overall size distribution of the stand can result in either increases or decreases in the inequality of the size distribution in the years following a mortality event, and these changes can also be influenced by the growth rate of the surviving trees (Kenkel et al., 1997). Factors causing years of high and low growth rate may also concurrently influence the probability of mortality for trees of slightly different size classes. For example, years of high growth rate may increase the probability of mortality for only the smallest trees, as high growth rates may increase the asymmetry of competition and cause 'regular' or autogenic mortality associated with stand dynamics (Oliver and Larsson, 1996). On the other hand, years of low growth rate may increase the probability of mortality for all size classes and cause 'irregular' or allogenic mortality (Oliver and Larsson, 1996), which is not necessarily exclusively in the smallest size classes, particularly if the causes of low growth are environmental.
Mean tree size, stand density and average competition were also significant predictors for the inequality of size increment at all but the nutrient-rich site at Candle Lake, where stand-level annual volume increment (lagged by 1 year) was the only significant predictor. This indicates that inequality in size increment was also somewhat related to stand dynamics. Again, however, the sign and significance of the coefficients were not consistent, indicating that observed relationships were also site-specific. These inconsistencies are suggestive of changes in the relative importance of one-sided (where small trees have little effect on large trees) and two-sided (where small trees also have an important effect on large trees) competition over time, both of which are observed to occur in even-aged tree populations (Brand and Magnussen, 1988). The sitespecific nature of the relationships between variables associated with stand dynamics and size and size-increment inequality suggests that the relative importance of these two modes of competition over time is also site-specific.
The results of this study suggest that factors influencing the annual growth rate are also influencing the development of size hierarchy in these forests, and that it is primarily the years with low growth rate that influence this development. Studies quantifying competition effects on tree growth usually measure subject trees and their competitors at a single point in time only, resulting in static estimates of competition indices for only single points in time (Burton, 1993). Spatial or aspatial indices of competition (e.g. Lorimer, 1983;Tome and Burkhardt, 1989;Holmes and Reed, 1991;Biging andDobbertin, 1992, 1995) are the dominant mechanism for generation of size hierarchy in many tree growth models. Using these indices usually results in moderately increased correlations between observed and predicted growth rates over time. However, there is clearly much residual variability in the growth response in these models that is not explained by competition. Schwinning and Weiner (1998) indicated that the effects of transient factors such as weather variation and FIG. 4. Annual trajectories of the average amount of competition to which each tree is subject (solid line) and the coefficient variation (CV %) of competition to which each tree is subject (dashed line) at each study plot. Competition was quantified using a distance-weighted absolute size index (eqn 4), with a variable search radius defined as 3 . 5 times each tree's crown width.
defoliation on development of size hierarchies are poorly understood, but the data presented in the present study indicate consistent and significant effects of yearly growing conditions. The sensitivity of jack pine to variation in weather (Larsen and MacDonald, 1995;Brooks et al., 1998;Hofgaard et al., 1999) and periodic defoliation by jack pine budworm (Volney, 1988;Gross, 1992) is clearly important to the differentiation of growth rates of trees within populations as their effects are likely to be different from those induced by density-dependent effects (Weiner and Thomas, 1986). Differences in growth rate among trees during years of poor growth may form the basis for development of size hierarchies on which asymmetric competition can act. This suggests that a complete understanding of the process of competition in these forests requires further evaluation of how factors that influence variation in the annual growth rate also affect how size hierarchies are generated in these populations.
*
ACKNOWLEDGEMENTS We thank
The variables in the table are stand density (DENS), mean tree volume (SIZE), average amount of competition to which each tree is subject (COMP), where competition is calculated using a distance-weighted absolute size index, and stand-level annual volume increment in the current year (AVI) and 1 year previously (AVI1).
Parameters in bold type were significant at P ,0 . 05.
† Background and Aims Current understanding of stomatal development in Arabidopsis thaliana is based on mutations producing aberrant, often lethal phenotypes. The aim was to discover if naturally occurring viable phenotypes would be useful for studying stomatal development in a species that enables further molecular analysis. † Methods Natural variation in stomatal abundance of A. thaliana was explored in two collections comprising 62 wild accessions by surveying adaxial epidermal cell-type proportion (stomatal index) and density (stomatal and pavement cell density) traits in cotyledons and first leaves. Organ size variation was studied in a subset of accessions. For all traits, maternal effects derived from different laboratory environments were evaluated. In four selected accessions, distinct stomatal initiation processes were quantitatively analysed. † Key Results and Conclusions Substantial genetic variation was found for all six stomatal abundance-related traits, which were weakly or not affected by laboratory maternal environments. Correlation analyses revealed overall relationships among all traits. Within each organ, stomatal density highly correlated with the other traits, suggesting common genetic bases. Each trait correlated between organs, supporting supra-organ control of stomatal abundance. Clustering analyses identified accessions with uncommon phenotypic patterns, suggesting differences among genetic programmes controlling the various traits. Variation was also found in organ size, which negatively correlated with cell densities in both organs and with stomatal index in the cotyledon. Relative proportions of primary and satellite lineages varied among the accessions analysed, indicating that distinct developmental components contribute to natural diversity in stomatal abundance. Accessions with similar stomatal indices showed different lineage class ratios, revealing hidden developmental phenotypes and showing that genetic determinants of primary and satellite lineage initiation combine in several ways. This first systematic, comprehensive natural variation survey for stomatal abundance in A. thaliana reveals cryptic developmental genetic variation, and provides relevant relationships amongst stomatal traits and extreme or uncommon accessions as resources for the genetic dissection of stomatal development.
The potential surface available for regulated gas exchange between plants and the atmosphere is set by stomatal number and distribution in the aerial epidermis. In Arabidopsis thaliana, stomata differentiate gradually during organ development, through a series of stereotyped yet flexible cell division and fate acquisition events. Stomatal abundance in different plant surfaces and environments is regulated, resulting in variable stomatal numbers and distribution patterns in mature organs (Bergmann and Sack, 2007;Casson and Hetherington, 2010, and references therein). This suggests that overlapping, partly redundant developmental pathways involving many genes must operate to produce a diversity of stomatal patterns and numbers while guaranteeing their functionality. Dissecting such gene circuits has only just begun, and several positive and negative regulators of stomata differentiation have been identified genetically and molecularly (reviewed by Bergmann and Sack, 2007;Nadeau, 2009;Dong and Bergmann, 2010).
The first recognizable stomata developmental event in A. thaliana Col-0 (see Fig. S1 in Supplementary Data, available online) is the asymmetric division of a protodermal cell [meristemoid mother cell (MMC)], termed entry division, that initiates a stomatal cell lineage; the smaller product, the meristemoid (M), undergoes up to three sequential asymmetric amplifying divisions, oriented in an inward spiral that places the M in the centre of a recognizable structure made by the larger division products (Bergmann and Sack, 2007). While the central M differentiates into a guard mother cell, which divides symmetrically to make a stoma, each of the larger cells can differentiate into a pavement cell or can become a MMC, experience an asymmetric entry division and start a satellite stomatal lineage (Bergmann and Sack, 2007). This is termed a spacing division because it puts at least one non-stomatal cell between the primary (also termed planet; Lucas et al., 2006) and the satellite stomata. Spacing divisions must involve cell -cell signalling events that hinder the development of stomata in contact, while amplifying divisions may also rely on unequal distribution of stomatal fate determinants and, thus, be regarded as executing a lineage-based programme. Therefore, stomatal development is an iterative process involving cell lineage and cell interaction-based processes.
The genetic and molecular dissection of stomatal development is mostly based on severe, aberrant phenotypes produced by induced (often loss-of-function) mutations. Such alleles, which impede stomata formation or produce stomatal clustering, have identified a large suite of stomata developmental regulators. Positive regulators are needed for stomata lineage initiation and development, while negative regulators are determinant for enforcing correct spacing (reviewed by Dong and Bergmann, 2010;Rowe and Bergmann, 2010). A few studies found some genes whose loss-of-function influence stomatal abundance without altering normal stomata development (Zhang et al., 2008;Dong and Bergmann, 2010, and references therein), but little is known about how diversity in functional, non-aberrant stomatal patterns could be set or how primary and satellite lineages contribute to final stomatal abundance.
The quantitative traits currently used to score stomatal abundance are stomatal index (SI), which measures the proportion of epidermal cells that are stomata, and stomatal density (SD) or number of stomata per area unit. SI and SD are the result of cell division patterns and of cell differentiation and expansion during organ growth (Geisler et al., 1998;Geisler and Sack, 2002). SD depends on stomatal number and on the size and number of non-stomatal epidermal cells (mostly pavement cells), while SI depends solely on cell-type proportion, regardless of cell size, and therefore both traits provide complementary information on final stomatal abundance and pattern. In developmental terms, the more widely examined organs for SI and SD are the cotyledon and first true leaf. Though both contribute little to total transpiration and photosynthesis in the adult plant, they are crucial in the first stages of plant development, when transpiration is needed for cell expansion-mediated plant growth while seedlings still have a shallow root system for soil water uptake. Arabidopsis thaliana seed resources are mostly depleted 3 days after germination (Penfield et al., 2005;Graham, 2008), and seedlings rely on photosynthesis as the sole source of nutrients and energy for cell division and dry matter increase. Therefore, in early postembryonic development, plant survival also depends on the compromise between photosynthesis and transpiration, and seedling stomatal abundance and distribution are probably under selective pressure in natural environments.
Naturally occurring variation among wild genotypes is an alternative genetic resource for studying stomatal development. Although many genes involved in stomatal development are remarkably conserved during plant evolution (Peterson et al., 2010), natural selection under diverse conditions may have resulted in natural variants that incorporate adaptive responses of stomata developmental gene networks to environmental cues. Indeed, natural variants in poplar and rice are being used to address the genetic basis of stomatal abundance in these taxa (Ferris et al., 2002;Laza et al., 2009). In A. thaliana, natural variation of complex biological processes has been successfully studied at the molecular level, and natural alleles with a broad range of quantitative effects have been isolated (Koornneef et al., 2004;Alonso-Blanco et al., 2009;Lefebvre et al., 2009). Variation in the stomatal density response to CO 2 doubling has also been reported for a number of A. thaliana accessions (Woodward et al., 2002). However, a detailed record of natural stomatal-related trait variants with traceable accessions is not yet available in this model species.
The Arabidopsis thaliana wild genotypes studied (see Table S1 in Supplementary Data, available online) included 47 accessions of the Versailles 48 nested Core Collection (McKhann et al., 2004) and 18 accessions analysed by Clark et al. (2007) for sequence polymorphisms. Ler was excluded from the original Clark's collection because it carries non-natural alleles derived from a fast-neutron mutagenesis (Re ´dei, 1962). These collections were built to maximize global genetic diversity in A. thaliana. Col-0 (N1092) was used as a laboratory reference accession for both collections. Three accessions were common to the two collections; thence, this study includes 62 different wild accessions plus Col-0. The INRA set seeds were obtained from the National Institute of Versailles's Agronomic Research (INRA, France), and the remaining accessions were provided by the Notthingham A. thaliana Stock Centre (NASC). Seeds were stratified at 4 8C in darkness for 3 d, then sown on pots containing a 2 : 1 : 1 mixture of soil (Prohumin, Klasmann-Deilmann 50-50), perlite and vermiculite, and grown in controlled chambers (Conviron MTR30) under a 16-h photoperiod at 21 + 1 8C, 60-70 % relative humidity and 150 + 20 mmol m 22 s 21 irradiance. Data were obtained from four or five plants for each accession (INRA and Clark sets, respectively), simultaneously grown in two random blocks. Accessions of INRA and Clark sets were grown separately.
The mature adaxial epidermis of cotyledons and first leaves of accessions were scored for stomata and pavement cell abundance. The adaxial epidermis was selected because it has simpler cellular patterns and fewer cell numbers than the abaxial epidermis. Cotyledons and first leaves of the same individuals were analysed. To ensure that organ growth was completed, cotyledons and first leaves were harvested 2 and 7 d after bolting, respectively. Sixteen accessions (see Table S1 in Supplementary Data) did not flower in the present conditions, and their organs were collected 1 week later than those of the latest bolting accession in the assay. Surface replicas were obtained with dental resin (Geisler et al., 2000) and micrographed with a Leica w IRB microscope equipped with a Leica w DC300F camera. Images were digitally processed with Adobe Photoshop CS3 (Adobe Systems Inc.). Epidermal cell counts of each individual were an average from two 0 . 327-mm 2 areas centred along the organ apical-to-basal and the median-to-margin axes. Cells having at least 25 % of their surface inside the sampling area were scored. In leaves, this area excluded trichomes and their surrounding socket cells. Stomatal and pavement cell densities (SD and PD) were calculated as number of stomata or pavement cells per area unit (cell number mm 22 ), respectively, and stomatal index (SI) as percentage of epidermal cells that were stomata.
Organ areas were measured in the same individuals scored for epidermal traits. Organ replicas were photographed with a Leicaw Stereomicroscope equipped with a Leica w -DC300F camera and areas determined with Leica w IM50 image manager v. 4.0 software.
The maternal environment (environmental growth conditions of the mother plants of a seed progeny; ME) can have effects on offspring phenotypes (Donohue, 2009). Since ME are often non-uniform across different experiments, to find out if the traits under study were affected by differences in laboratory ME, tests were carried out. Several separate seed harvests were obtained for a number of accessions. Seven accessions representing the range of stomatal abundance variation in the Clark set (Bur-0, Col-0, Cvi-0, Nfa-8, Shakdara, Ts-1 and Van-0) were grown and selfed from the seed lot used for trait description, in two independent experiments (referred as ME1 and ME2) conducted in the same chamber at different times. Additionally, the four accessions present in both the Clark and the INRA collections (Bur-0, Col-0, Cvi-0 and Shakdara) were grown and selfed from the INRA collection seed stocks in a chamber different from that used for ME1 and ME2 (ME3). Both chambers were Conviron MTR30 and plants in the three ME were grown using the same environmental setting described above. Subsequently, seeds derived from a single plant for each accession and ME were simultaneously grown, and six individuals per accession and ME were scored for the various traits. Data were analysed in two sets: one set contained four accessions [genotypes (G)] in three maternal environments (4G × 3ME), and the other had seven accessions in two ME (7G × 2ME).
Seeds harvested from plants grown simultaneously (i.e. under the same ME) were grown as described above. Developmental homogeneity was established by monitoring germination every 6 h. Radicle emergence set germination time point at 0, and ten synchronous seedlings per accession were collected 48 h later. Specimens were fixed in ethanol : acetic acid 9 : 1 (v/v), dehydrated through ethanol : water series, rehydrated, mounted in Hoyer's medium (Liu and Meinke, 1998), and inspected with differential interference contrast (DIC) (Eclipse 90i microscope and Nikon DXM1200C camera). Epidermal cell types were recorded in two fields per cotyledon, avoiding the central zone and edges, and represented about 70 % of the organ surface. Cell types were identified accordingly to Zhao and Sack (1999) and Geisler and Sack (2002). Primary and satellite stomatal lineages were scored as previously defined (Berger and Altmann, 2000;Kutter et al., 2007;Zhang et al., 2008). No distinction was made between secondary and higher-order satellite lineages. Ten additional individuals per accession were grown to maturity and their stomatal index determined.
Cell density values were log e transformed, while cell index values were arcsin-root transformed to improve normality of distributions and homogeneity of variances. None of the outcomes and conclusions changed when using the original data, and therefore most descriptions are based on untransformed data to simplify interpretation. INRA and Clark sets are described separately because two-way ANOVA and t-test analyses showed that some trait values in the four accessions common to both sets (Bur-0, Col-0, Cvi-0 and Shakdara) differed significantly between assays, probably due to slight environmental differences resulting from their independent growth. For each trait, the amount of variation was calculated by the coefficient of variation (CV ¼ standard deviation/ average) and by the fold change (maximum accession value/ minimum accession value). Broad-sense heritability (h 2 ) was calculated as:
where V G (genetic variance) is the among-genotype (accession) variance component, and V E (environmental variance) is the residual error variance component estimated by restricted maximum likelihood (REML) analysis (Lynch and Walsh, 1998). Genetic correlations between traits were estimated by Pearson's and Spearman's tests using accession mean trait values (family mean of the selfing offspring, which carry practically the same genotype; Lynch and Walsh, 1998). Similar relationships and conclusions were achieved from both analyses (unless indicated) and, hence, only Pearson coefficients are reported. Fisher's z-test (Zar, 1996) was used to identify correlations that differed significantly between the two accessions sets. Differences between mean stomatal lineage indices (arcsin-root transformed) of selected accessions were tested by Student's t-tests. Environmental maternal effects on traits were evaluated by a two-factorial analysis of variance (ANOVA), with genotype (accession) and maternal environment as fixed factors. All statistical analyses were performed with the SPSS 11 . 0 package (SPSS Inc., Chicago, IL, USA). To identify and classify phenotypic diversity within the accession sets, hierarchical cluster analysis was carried out with MeV v4.2 (TM4 Microarray Software Suite, Boston, MA, USA) using standardized data to perform an average linkage clustering based on uncentred Pearson's correlation as the distance metric.
Stomatal abundance traits in 62 wild genotypes of A. thaliana, grouped in two sets (namely, the INRA and the Clark sets), which were grown independently and therefore described separately, were analysed (see Materials and methods, and Table S1 in Supplementary Data, available online). Col-0 was included in both sets as a reference genotype. Accessions were examined for six traits: stomatal index, stomatal density and pavement cell density in cotyledons (C) and first leaves (L) in the adaxial epidermis (namely SIC, SDC, PDC SIL, SDL and PDL). Pavement cell density was included to address its impact on stomatal density variation. Representative sampled epidermes for accessions with distinct stomatal abundance are shown in Fig. 1. No aberrant phenotypes were found in any of the natural accessions studied (not shown).
All traits displayed considerable and continuous variation among the genotypes scored in each set (Fig. 2, and Table S2 in Supplementary Data). Stomatal indices showed the lowest variation both in cotyledons and first leaves (1 . 3-to 1 . 7-fold). In cotyledons, SI varied 1 . 5-fold, ranging from the lowest values in Can-0 and Ts-1 to the highest values in Bur-0 and Van-0 in the INRA and Clark sets, respectively (see Table S1 in Supplementary Data). Bur-0, represented in both sets, also had the second maximum SIC within the Clark group. Remarkably, Col-0 cotyledons, commonly used to study stomatal development, showed one of the highest stomatal index values, ranking right after Bur-0 in the two data sets. In leaves, minimum stomatal index values occurred for Can-0 and Nfa-8, while maximums corresponded to Ishikawa (followed by Bur-0) and Rrs-7 (followed by Bur-0 and Van-0) in the INRA and Clark sets, respectively.
The largest variation was observed in stomatal densities, which varied between 2 . 1-and 3 . 5-fold (see Table S2 in Supplementary Data). Sakata and Got-7 had the lowest SD values in cotyledons, and Can-0 and Tsu-1 in leaves (INRA and Clark sets, respectively). The highest SD corresponded to Jm-0 and Van-0 cotyledons, and Bur-0 and Van-0 leaves (INRA and Clark sets, respectively).
Pavement cell density showed an intermediate amount of variation in both organs (ranging between 1 . 9-and 2 . 3-fold). Accessions setting the upper and lower limits of PDC were the same that flanked the SDC variation range (Sakata and Got-7 had the lowest PDC, and Jm-0 and Van-0 the highest one). In leaves, Ishikawa and Tsu-1 showed the lowest PDL, while Jm-0 and Bor-4 had the highest PDL values (INRA and Clark sets, respectively).
Mean, minimum and maximum values for all traits were always higher in leaves than in cotyledons (Table S2 in Supplementary Data), suggesting a supra-organ regulation of stomatal abundance and pavement cell density, likely related to heteroblasty (Tsukaya et al., 2000). Ratios between equivalent traits in leaves and cotyledons were calculated for each accession (Table S1 in Supplementary Data). Few exceptions to the general behaviour were found, ratios varying from 1 . 3 to 3 . 2 for SD, from 1 . 1 to 2 . 3 for PD and from 1 . 0 to 1 . 6 for SI. Remarkably, Col-0 is one of six accessions with nearly identical SI in leaves and cotyledons.
The six traits showed high to moderate broad-sense heritabilities (h 2 ), which ranged from 0 . 33 to 0 . 71 (Table S2 in Supplementary Data). Overall, stomatal and pavement cell densities had similar heritability in cotyledons and leaves, with higher h 2 values than for stomatal index. Therefore, our data show that there is substantial genetic variation among accessions for all six traits.
To test for co-ordinated regulation of stomatal abundance in cotyledons and leaves and for relationships among epidermal traits within each organ, correlation coefficients between all pair combinations of the six traits were estimated (Fig. 3A-E, and Table S3 in Supplementary Data). Stomatal and pavement cell densities showed the strongest correlations, both in cotyledons and leaves (r ¼ 0 . 73-0 . 93; P , 10 28 ), indicating that variation in pavement cell size is a major cause for SD variation among accessions. In addition, SI and SD positively correlated in both organs, with larger correlations in cotyledons (cotyledons: r ¼ 0 . 86, P , 10 25 ; leaves: r ¼ 0 . 78, P , 10 24 ). The relationship between SI and PD differed between organs; in cotyledons it was moderate in the two accession sets (r ¼ 0 . 49 -0 . 61; P ≤ 5 × 10 23 ), and in leaves it was weak in the Clark set (r ¼ 0 . 5, P ≤ 3 × 10 22 ) or absent in the INRA set. In addition, correlations between cotyledon and leaf values for SI, SD and PD traits were moderately positive (Fig. 3A, and Table S3 in Supplementary Data).
To identify common and distinct phenotypic patterns underlying the observed correlations between trait pairs, clustering analyses were carried out using an uncentred Pearson correlation with average linkage metric (Fig. 3F, G). Dendrograms show, in a heat-map format, the standardized trait values (z-scores) for each accession, represented as a rank of divergence for the trait mean in the set. In both collections, accessions were assigned to two main clusters: one with prevalently low-to-medium values (cluster type I) and another containing accessions with mostly high-to-medium values (cluster type II). Both cluster types included a group with all trait values ranking similarly, either low-moderate (type Ia; 36 % of accessions) or high-moderate (type IIa; 13 % of accessions). These two main clusters also included groups with trait values differing among organs, either standing higher (Ib and IIb; 6 % and 10 %, respectively) or lower (Ic; 4 %) in cotyledons than in leaves. In the INRA collection, cluster II included two additional groups with lower (IIc; 19 % of accessions) or higher (IId; 10 % of accessions) SI values than other traits. Therefore, about 49 % of accessions showed a phenotypic pattern with strong relationships among all trait values (Ia and IIa), in agreement with the correlations found among traits (Fig. 3A). The remaining groups contained phenotypes that deviated from these general correlations and showed, for instance, leaf or cotyledon stomatal index values differing from cell density values (see Discussion). Thus, clustering analyses revealed natural prevalent phenotypic patterns, and also accessions with distinct uncommon phenotypes.
Organ size regulation is expected to integrate mechanisms controlling stomatal lineage initiation and proliferation. To explore relationships between organ size and stomatal abundance, cotyledon and first leaf areas were measured in the INRA set. Correlations between cell and organ traits were estimated (see Fig. 4, and Tables S1, S4 and S5 in Supplementary Data). Substantial variation was found for cotyledon and leaf area (Fig. 4A, B). Notably, Sakata displayed extremely large organs, appearing as an outlier of the observed continuous distribution. Hence, organ area data are presented with and without Sakata (Fig. 4C -F and Tables S4 and S5 in Supplementary Data). Broad-sense heritability estimates were high in both cases (ranging from 0 . 59 to 0 . 76; Table S4 in Supplementary Data), indicating that there is substantial genetic variation for these traits among natural accessions. Relationships among cellular and organ size traits were analysed by Spearman correlation (Table S5 in Supplementary Data), since organ area data did not meet parametric test assumptions required for Pearson correlations. Cotyledon and first leaf areas were positively correlated (r ¼ 0 . 48 and 0 . 44; P , 0 . 05; with and without Sakata, respectively). Both showed negative correlations with their respective SD and PD (r between -0 . 29 and -0 . 59; P , 0 . 05), particularly in cotyledons. Organ area and SI were negatively correlated in cotyledons (r ¼ -0 . 3; P , 0 . 05), but not in leaves. Total pavement cell number for all accessions was estimated from their PD and organ area values (not shown). Cell number and organ area showed a strong positive correlation, which in leaves was higher than the relationship with any other trait (Table S5 in Supplementary Data), as previously reported (Granier et al., 2000;Cookson et al., 2005Cookson et al., , 2007)). Hence, in this sample of accessions, larger organs often had more and larger pavement cells (leading to lower stomatal densities) and vice versa, showing a relationship between organ and pavement cell number and size. All together, the present results show that variation in cell number and size contributes to genetic diversity in cotyledon and first leaf size. However, accessions with larger cotyledons tended to have a lower proportion of stomata (SI), suggesting a distinct genetic link between cotyledon size and stomatal lineage development.
Two complementary data sets were obtained to evaluate whether small differences in laboratory maternal environment (ME), expected to occur among separate plant growth experiments, would affect the traits surveyed in this study. One set assessed effects of three MEs on four accessions [genotypes (G)], while the other used two MEs on seven accessions (4G × 3ME and 7G × 2ME, respectively; see Materials and methods). Each ME refers to a separate experiment where plants of the various accessions were grown simultaneously from seed germination to seed production. Plants from all maternal sources were simultaneously grown and the traits scored. ME effects were estimated by two-factorial ANOVA, with genotype (G; accession) and ME as factors (Table S6 in Supplementary Data). A few significant effects of ME (P , 0 . 05) were only detected in the 4G × 3ME data set. MEs had no effect on stomatal abundance traits, except for a marginal influence on leaf SD (R 2 ¼ 2 . 9). Minor effects were also found on leaf PD (R 2 ¼ 7 . 09). However, MEs and genotype × ME interactions show significant effects on organ sizes, being larger in cotyledons (ME, R 2 ¼ 10 . 4; G × ME, R 2 ¼ 20 . 3) than in leaves (ME, R 2 ¼ 3 . 9; G × ME, R 2 ¼ 13 . 24).
To address if natural variants differ in the contribution of distinct developmental processes affecting stomatal index in mature organs, two quantitative processes were analysed: the proportion of primary lineages stemming from protodermal cells, and the proportion of satellite lineages that arise from nonstomatal cells in primary lineages (Fig. S1 in Supplementary Data). The analyses were restricted to cotyledons in four accessions with extreme stomatal index values (Fig. 3F, G): Ts-1 and Sp-0, with very low SI, and Col-0 and Bur-0 as accessions with the highest SI values in the two collections. Adaxial cotyledon epidermes were examined at two time points: 48 h postgermination (hpg), the shortest time amenable to analyses, and 21 d post-germination (dpg), when satellite lineages have appeared extensively. Lineage initiation was monitored by the primary lineage index (PLI) or proportion of primary stomata plus primary stoma precursors to total epidermal cells, and the satellite lineage index (SLI), or proportion of satellite stomata plus satellite stomata precursors (see Fig. 5 and Materials and methods). At 48 hpg (Table 1) the two high SI accessions Bur-0 and Col-0 had initiated a similar number of lineages ( primary plus satellite), which was significantly higher than that of the low SI accessions Sp-0 and Ts-1. The total lineage index (TLI ¼ PLI + SLI) at 48 hpg showed highly significant correlation with proportion of stomata (SI) in mature cotyledons at 21 dpg (TLI vs. SI: r ¼ 0 . 99, P ¼ 0 . 01), indicating that stomatal phenotypes observed at 48 hpg are predictive of mature organ phenotypes. In spite of this correlation, the low SI but not the high SI accessions showed a higher TLI at 48 h than SI at 21 dpg (Table 1). This could be due to an increased pavement cell production through extended stomatal lineage amplification divisions and/or to more pavement cell symmetric divisions.
Satellite lineage contribution to total stomatal lineages showed considerable variation among accessions, ranging from 10 % in Ts-1 to 33 % in Col-0 (Table 1). Although Bur-0 and Col-0 had different SLIs, they had a similar stomatal precursors index, which presented larger individual variability (as expected for a parameter very sensitive to small variations in temporal development; Fig. 5A). Regarding primary lineage index, Bur-0 and Col-0 showed the largest difference, while Bur-0 and Ts-1 displayed similar PLI values (P . 0 . 05). Therefore, two accessions with similar stomatal index showed a different primary lineage index (Col-0 and Bur-0), while accessions with a different stomatal index presented an equivalent primary lineage relative abundance (Bur-0 and Ts-1). Thus, primary and satellite lineage initiation frequencies can combine in different ways: low production of primary stomata with high production of satellite lineages (Col-0); high production of primary stomata and low satellite initiation (Ts-0); or moderate (Sp-0) or high (Bur-0) production of both primary and satellite lineages. These results reveal substantial natural variation in the initiation of both primary and satellite stomatal lineages, suggesting that the two processes are partially under independent genetic control.
Wild genotypes of A. thaliana show substantial genetic variation in epidermal cell-type abundance Analyses of quantitative traits related to cotyledon and first leaf stomatal abundance in 62 wild A. thaliana genotypes, evaluated in two sets, have revealed considerable natural genetic variation. Stomatal density is the most variable trait, while stomatal index showed the least variation. SI relates to cell division and differentiation, and its variation may result from relatively narrow limits that ensure a functional epidermis co-ordinated with the underlying mesophyll. The moderate pavement cell density variation found in the present study would mostly arise from cell-size phenotypes. Combinations of low SI and moderate PD variations most likely account for the higher SD diversity. The genetic variation found in this work suggests that phenotypes of the genotypes examined are not deleterious under natural environments. In agreement, aberrant stomatal phenotypes were not found, indicating that incorrect patterns reduce plant fitness. To determine if the genetic variation observed may be involved in environmental adaptation, correlations were inspected among stomatal traits and geographic and historical climate data from the accession collection sites. Marginal positive correlations were detected between leaf SI and mean monthly precipitation from October to April (Spearman r ¼ 0 . 29-0 . 35; P ¼ 0 . 014 -0 . 044), which suggests adaptive population differentiation to water availability. However, the broad geographic distribution of the accessions studied and imprecise information on their original habitats may influence these findings. In addition, given the high plasticity of the traits studied (Casson and Gray, 2008;Lampard, 2010, and references therein), phenotypes may differ between laboratory conditions and natural environments, fading a putative adaptive basis in the underlying genetic variation.
Maternal environments may influence offspring traits, especially those expressed early in the life cycle (reviewed by Donohue, 2009). It was found that cotyledon and first leaf cell-abundance traits were marginally or not affected by differences in the ME occurring in laboratory experiments. In agreement, A. thaliana stomatal density showed no ME effect during 15 generations grown at different CO 2 concentrations (Teng et al., 2009). As demonstrated for other species (Roach and Wulff, 1987), ME had a moderate but significant effect on cotyledon size, although genetic differences accounted for most of the phenotypic variation. Minor ME effects were also detected on first leaf size. Therefore, the phenotypic differences found among accessions for cellular traits are mainly determined by genotypic variation and not by the maternal environments used.
Correlation analyses uncovered an overall pattern of positive genetic relationships among stomatal and pavement cell-abundance traits within and between organs. The correlations we found most likely have a common genetic basis, because A. thaliana mutations in several stomata developmental genes simultaneously affect all or most traits examined here [notably GPA1 and ERECTA (Zhang et al., 2008;Nilson and Assmann, 2010;van Zanten et al., 2010); other genes are reviewed by Dong and Bergmann (2010)]. Nonetheless, such trait relationships may also originate from trait co-evolution (Armbruster and Schwaegerle, 1996).
Relationships between all traits for each organ suggest that shared genetic networks control cell-type proportion and density at the organ level, although they appear more closely related in cotyledons than in leaves. These results are consistent with the proposed mechanism co-ordinating cell proliferation and expansion at the organ level (Tsukaya, 2006;Fujikura et al., 2007Fujikura et al., , 2009;;Tsukaya, 2008;Micol, 2009). Correlations were also found for each trait between organs. In agreement with these correlation patterns, cluster analyses identified two consistent groups of accessions where all six cell-abundance traits were either low (group Ia) or high (group IIa), indicating that similar genetic networks control cell-type abundance in both organs. These two groups account for nearly half of the accessions, suggesting a mechanism for supra-organ co-ordination of stomatal abundance in A. thaliana. Interestingly, some accessions deviated from the prevalent trait correlations. Clustering analysis singled out groups with uncommon phenotypic patterns of opposite SI and cell densities values (e.g. Ishikawa and Jm-0) or having weak or absent relationships for all traits between organs (e.g. C24 or Shakdara). These accessions that deviate from the general tendencies suggest that the genetic programmes controlling the various traits and their co-ordination also involve unshared factors.
Cotyledon and leaf areas, which also showed significant genetic variation, negatively correlated with most cell-abundance traits in the 48 accessions examined (INRA set). In most accessions, large organs are built not only by more cells but also by larger ones (as inferred by their low pavement cell densities) and have lower stomatal densities; conversely, high stomatal and pavement cell densities (small cells) and lower cell numbers are the norm in small organs. In cotyledon, stomatal index and organ area were also negatively correlated. Overall, the present results show that wild accessions comply with the above-mentioned theories of co-ordinated cell behaviour in leaf growth (reviewed by Tsukaya, 2008), and suggest that, at least in cotyledons, The index of stomatal primary lineages, satellite lineages and total lineages (primary + satellite lineages), and the percentage of satellite lineages over the total stomatal lineages were determined in the adaxial epidermis of cotyledons 48 hpg in Ts-1, Sp-0, Col-0 and Bur-0. The stomatal index at maturity ( 21 dpg stomatal lineage development is partially linked to an organ level growth control.
Some accessions carry allele combinations determining extremely low or high stomata proportions and cell densities, like Ts-1, Sp-0, Van-0 or Bur-0. The highest SD was recorded for Van-0, an erecta mutant (van Zanten et al., 2010). Loss-of-function ERECTA mutations in Col and Ler increase SD (Shpak et al., 2005) and total epidermal cell densities (Tisne ´et al., 2008) but do not affect SI in adult leaves (Masle et al., 2005), or lower it in the abaxial cotyledon epidermes (Shpak et al., 2005). It was found that Van-0 also has a very high pavement cell density and SI, suggesting differences in erecta phenotypes between epidermes or the presence of natural erecta modifiers in Van-0. Thus, the present analyses discriminated a loss-of-function allele of a gene involved in epidermal cell density determination, further revealing additional features associated with this gene.
The wide natural genetic variation in stomatal abundance traits that are described here can be readily dissected by QTL mapping through available recombinant inbred line populations (Simon et al., 2008). Its combination with the complete genome sequences of a fast growing number of A. thaliana accessions (Weigel and Mott, 2009) will provide essential information on the mechanisms involved in this developmental process. Given the impact of stomatal abundance on transpiration and photosynthesis, these studies might provide new alleles of agronomic interest for genetic improvement of crop performance.
Primary and satellite stomatal pathways have been extensively described (Bergmann and Sack, 2007; see Fig. S1 in Supplementary Data). However, mechanisms regulating their proportions are poorly understood. The present analysis has found genetic variation in the relative frequency of primary and satellite lineages. A similar stomatal index could also be achieved with different relative frequencies. Therefore, primary to satellite stomata proportion is a distinct trait that reveals different phenotypes hidden under the same SI. This range of previously unknown genetic diversity suggests an even higher complexity of genetic factors involved in stomatal abundance control.
Primary stomata abundance directly relates to entry divisions, while satellite stomata abundance will be affected directly by spacing divisions, and indirectly by amplifying and entry divisions (Bergmann and Sack, 2007). Several genes are known to control these three asymmetric divisions (Dong and Bergmann, 2010), but very few seem to affect satellite lineages specifically. One is AGL16, whose microRNA-mediated regulation limits satellite lineage production with no apparent effect on primary stomata (Kutter et al., 2007). Also, two G-protein subunits, AGB1 and GPA1, function as mutually antagonistic modulators of satellite lineages (Zhang et al., 2008). GPA1 also controls epidermal cell size (Nilson and Assmann, 2010). The satellite stomata fraction may largely depend on the cell-proliferation time window in Col-0 adaxial cotyledon epidermes (Geisler and Sack, 2002). The earlier cell division arrest in adaxial versus abaxial epidermes concurs with lower satellite stomata production. Therefore, changes in the proliferation-window length may alter satellite stomata abundance. Division rate differences would have a similar influence, and both factors may also impact on higher-order satellization.
Primary and satellite stomatal pathways are assumed to contribute to the phenotypic plasticity of A. thaliana stomatal development in response to internal and environmental cues (Bergmann and Sack, 2007;Casson and Gray, 2008;Lampard, 2010), but whether each pathway displays specific responses remains to be established. Unravelling the molecular genetic basis of the observed variation would shed light on a basic developmental programme, and may also provide tools to understand plant productivity under different environments.
A systematic data collection is provided on cotyledon and first leaf traits related to stomatal abundance for 62 wild Arabidopsis thaliana accessions analysed in two sets. This survey reveals substantial genetic variation in SI, SD and PD in both organs, and shows that these traits are largely unaffected by the laboratory maternal environments used. SD shows the highest correlation with the other traits within each organ, strongly suggesting that these traits share genetic bases. Inter-organ correlations were also found, supporting the operation of supra-organ mechanisms for the control of cell-type abundance in the accessions surveyed. Clustering analyses showed that while half of the accessions had these strong relationships among all traits, accessions with uncommon phenotypic patterns, suggestive of differences among genetic programmes controlling the various traits, do also occur in nature. The 47 wild accessions surveyed had considerable variation in cotyledon and organ size, which were negatively correlated with cell densities. Notably, variation was also identified in two distinct processes during stomatal differentiation ( primary and satellite lineage initiation), which can be hidden under stomatal abundance trait values, because accessions with a similar SI could have very different lineage class ratios. Thus, genetic determinants of primary and satellite lineage initiation can combine in several ways and produce a similar final stomatal abundance phenotype. In addition to revealing this cryptic diversity in a developmental process, this first systematic, comprehensive natural-variation survey for stomatal abundance in A. thaliana provides relevant relationships amongst stomatal traits at the organ and supra-organ levels. This survey also identifies accessions with extreme values or uncommon trait correlations, which are valuable tools for further understanding of stomatal development in natural germplasm. SUPPLEMENTARY DATA Supplementary data are available online at www.aob.oxford- journals.org and consist of the following. Figure S1: images illustrating the main events and cell types during stomatal development in A. thaliana. Table S1: accession values of traits analysed. Table S2: summary statistics of stomata and pavement cell-abundance traits in adaxial cotyledon and first leaf epidermes. Table S3: Correlation analysis of epidermal cell-abundance traits. Table S4: summary statistics of cotyledon and first leaf areas analysed in the INRA set of accessions. Table S5: correlation analysis of organ area and cell-abundance traits in the INRA set of accessions. Table S6: Summary statistics of maternal environmental effects on epidermal cell abundance and organ size traits.
This work was supported by the
Primary extracranial meningiomas are rare neoplasms, frequently misdiagnosed, resulting in inappropriate clinical management. To date, a large clinicopathologic study has not been reported. One hundred and forty-six cases diagnosed between 1970 and 1999 were retrieved from the files of the Armed Forces Institute of Pathology. Histologic features were reviewed, immunohistochemistry analysis was performed (n = 85), and patient follow-up was obtained (n = 110). The patients included 74 (50.7%) females and 72 (49.3%) males. Tumors of the skin were much more common in males than females (1.7:1). There was an overall mean age at presentation of 42.4 years, with a range of 0.3-88 years. The overall mean age at presentation was significantly younger for skin primaries (36.2 years) than for ear (50.1 years) and nasal cavity (47.1 years) primaries. Symptoms were in general non-specific and reflected the anatomic site of involvement, affecting the following areas in order of frequency: scalp skin (40.4%), ear and temporal bone (26%), and sinonasal tract (24%). The tumors ranged in size from 0.5 up to 8 cm, with a mean size of 2.3 cm. Histologically, the majority of tumors were meningothelial (77.4%), followed by atypical (7.5%), psammomatous (4.1%) and anaplastic (2.7%). Psammoma bodies were present in 45 tumors (30.8%), and bone invasion in 31 (21.2%) of tumors. The vast majority were WHO Grade I tumors (87.7%), followed by Grade II (9.6%) and Grade III (2.7%) tumors. Immunohistochemically, the tumor cells labeled for EMA (76%; 61/80), S-100 protein (19%; 15/78), CK 7 (22%; 12/55), and while there was ki-67 labeling in 27% (21/78), \3% of cells were positive. The differential diagnosis included a number of mesenchymal and epithelial tumors (paraganglioma, schwannoma, carcinoma, melanoma, neuroendocrine adenoma of the middle ear), depending on the anatomic site of involvement. Treatment and follow-up was available in 110 patients: Biopsy, local excision, or wide excision was employed. Follow-up time ranged from 1 month to 32 years, with an average of 14.5 years. Recurrences were noted in 26 (23.6%) patients, who were further managed by additional surgery. At last follow-up, recurrent disease was persistent in 15 patients (mean, 7.7 years): 13 patients were dead (died with disease) and two were alive; the remaining patients were disease free (alive 60, mean 19.0 years, dead 35, mean 9.6 years). There is no statistically significant difference in 5-year survival rates by site: ear and temporal bone: 83.3%; nasal cavity: 81.8%; scalp skin: 78.5%; other sites: 65.5% (P = 0.155). Meningiomas can present in a wide variety of sites, especially within the head and neck region. They behave as slow-growing neoplasms with a good prognosis, with longest survival associated with younger age, and complete resection. Awareness of this diagnosis in an unexpected location will help to avoid potential difficulties associated with the diagnosis and management of these tumors.
Meningioma is a well-recognized tumor of the central nervous system (CNS) that typically arises in proximity to the meninges. These neoplasms are more common in females during the middle decades of life and account for 24-30% of primary intracranial tumors [1]. Most commonly extra-neuraxial, meningiomas are found overlying the surface of the brain or at the skull base [2]. Uncommonly, meningiomas occur in intraventricular [3,4], intraparenchymal [2,5], or intraosseous locations. In rare instances (\2%), they appear as an extracranial tumor, most often in the head and neck region [2][3][4][5][6][7][8][9][10][11][12][13][14][15][16][17][18][19][20], and specifically in the sinonasal tract [20], ear and temporal bone [19], and scalp [21]. The histopathologic diagnosis is usually straightforward; however, the diagnosis may pose challenges in these unexpected locations where the differential diagnosis includes paraganglioma, carcinoma, melanoma, schwannoma, and olfactory neuroblastoma, among others. Previous reports on meningioma in uncommon locations are largely limited to individual cases or small series [2,[4][5][6][7][8][9][10][12][13][14][15][16]. The objective of this retrospective study is to present our experience, the largest series to date, with extracranial meningiomas, highlighting clinicopathologic features and outcome associated with the surgical treatment of these tumors.
The records of 205 patients with primary extracranial meningiomas were retrieved from the registries of the Otorhinolaryngic-Head and Neck Tumor Registry, Soft Tissue Registry and the Neuropathology Registry of the Armed Forces Institute of Pathology (AFIP) between the years 1970 and 1999. The anatomic sites included ear, temporal bone, sinonasal tract (including paranasal sinuses), scalp, chest and thorax, pelvis and extremities. However, 59 patients were excluded from further consideration because of at least one of the following reasons: (1) paraffin blocks were unavailable for additional sections; (2) the original submitted case did not have sufficient demographic information supplied to warrant inclusion or from which to obtain adequate follow-up information; (3) patients with primary tumors of the intracranial cavity, who later developed extra-cranial manifestations were excluded; (4) the cases were diagnosed indefinitely, using terms such as ''consistent with,'' ''suggestive of,'' or ''suspicious for.'' The remaining 146 patients with meningioma compose the subject of this study based upon adequate clinical information, demographic findings, and materials for follow-up and hematoxylin and eosin-stained slides to make a definitive diagnosis.
Materials within the files of the AFIP were supplemented by a review of the patient demographics, medical history, surgical pathology and operative reports, oncology data services, cancer registry records, and by written questionnaires or oral communication with the treating physician(s). Follow-up data tabulated included information regarding the location of the primary site, the specific treatment modalities used, and the current status of the disease and patient. Extent of resection was defined as partial (P) or gross total resection (GTR) based on information coded at the treating institution. This clinical investigation was conducted in accordance and compliance with all statutes, directives, and guidelines of the Code of Federal Regulations, Title 45, Part 46, and the Department of Defense Directive 3216.2 relating to human subjects in research.
The macroscopic pathology observations noted within this study were gathered from the individual gross descriptions of the neoplasms given by the contributing pathologists. In all cases, tumor tissue consisted of formalin fixed, routinely processed, and paraffin embedded surgical specimens. Hematoxylin and eosin-stained slides from all cases were reviewed to confirm the diagnosis of meningioma by all four neuropathologists (EJR, JPB, GDS, and HM) and an otorhinolaryngic pathologist (LDRT), independently initially, and followed by a consensus conference for any case that raised a diagnostic differential consideration or was a Grade II or Grade III tumor. Meningiomas were assessed and classified according the World Health Organization 2000 and 2007 [1,22]. Mitoses were counted by examining 10 high power fields (HPF) for each tumor in the region of highest cellularity and determining an average number per 10 HPF. Lesions were considered atypical if they possessed a mitotic rate greater than four per 10 high-power fields (2.5 mm 2 ) [23,24] and/or had at least three of the following histologic features: hypercellularity, growth of tumor cells in sheets, prominent nucleoli, necrosis, and small cell formation [1,22].
Immunophenotypic analysis was performed in 80 cases using the standardized avidin-biotin method, using 4 lm thick, formalin-fixed, paraffin-embedded sections. Table 1 documents the pertinent, commercially available immunohistochemical antibody panel used. The analysis was performed on a single representative block in each case. When required, proteolytic antigen retrieval was performed with predigestion for 3 min with 0.05% Protease VIII (Sigma Chemical Co, St. Louis, MO, USA) in a 0.1 M phosphate buffer at a pH of 7.8 at 37°C. Antigen enhancement (recovery) was performed, as required, using formalin-fixed, paraffin-embedded tissue treated with a buffered citric acid solution and heated for 20 min in a calibrated microwave oven. Following this, the sections were allowed to cool at room temperature in a citric acid buffer solution for 45 min before continuing the procedure. Standard positive and negative (serum) controls were used throughout. The antibody reactions were graded as weak (1?), moderate (2?), and strong (3?) staining, and the fraction of positive cells was determined by separating the percentage of positive cells into four groups: \10%, 10-50%, 51-90%, and [90%. To quantify the percentage of nuclei staining with Ki-67, at least 500 tumor cells were manually counted; three separate counts were performed in the region of tumor with highest nuclear staining, and an average percentage of positively stained nuclei was calculated.
A review of publications in English (MEDLINE 1966-2008) was performed, and all cases primarily involving the ear and temporal bone, skin, orbit, oral cavity, nasal cavity, nasopharynx, paranasal sinuses, soft tissues, and solid organs (lung, liver, spleen) and hollow organs (colon, bladder, uterus, stomach) were included in the review. Cases of cranial cavity or spinal canal meningiomas which metastasized to distant sites were not included in the review, unless they involved the sites described above. Non-English articles were not included.
Survival probabilities were calculated according to the Kaplan-Meier method and were measured from the date of diagnosis until the date of last follow-up or until death. Bivariate associations between survival and prognostic factors were tested using the log-rank test. Multivariate associations between survival and prognostic factors were tested using the Cox proportional hazards regression. Categorical variables were analyzed using Chi-square tests to compare observed and expected frequency distributions. Comparison of means between groups were made with unpaired t tests. Confidence intervals of 95% were generated for all positive findings. The alpha level was set at P \ 0.05. Statistical analysis was performed using SPSS version 12.0 for Windows (SPSS, San Diego, CA) and STATA version 8.1 for Windows (Stata Corporation, College Station, TX). Often the date of last follow-up was given as month and year, but no day was available. In those cases, follow-up was calculated to the 15th day of the month.
The patients included 74 females and 72 males (Table 2). There was a statistically significant gender difference based on anatomic site (P = 0.044), with a greater number of men among patients with scalp skin lesions (63%) and a greater number of women among ear lesions (66%). Patients' age at presentation ranged from 3 months to The average age at presentation for women was older than men, at 45.6 and 39.6 years, respectively, but this difference did not reach statistical significance (P = 0.09); likewise, there was no statistical difference between genders when stratified for specific site (even though the average age for patients with nasal cavity tumors in males versus females was 42.1 versus 51.4 years). Ear and temporal bone tumors had an older mean age at presentation when compared to scalp skin and other soft tissue lesions: 50.1 versus 36.2 years (P = 0.004) and 35.6 years (P = 0.042), respectively. A possible explanation may be entrapped meningeal tissue during embryologic development, which undergoes neoplastic transformation and comes to clinical attention before an intraosseous lesion. Patients symptoms were referable to the anatomic site of tumor involvement. Skin scalp lesions and neck lesions presented with a mass. Tumors from the ear and temporal bone presented with hearing changes, either sensorineural or conductive hearing loss, otitis, headaches, dizziness, unsteadiness, vertigo, disequilibrium, tinnitus, otalgia, and bleeding. Facial nerve or other cranial nerves were involved in 4 patients. Patients with tumors arising in the nasal cavity presented with a mass lesion, epistaxis, sinusitis, pain, and/or obstructive symptoms. Four patients in this group also had visual changes, including blindness, possibly related to pressure effect from the mass as it expanded in size, although the blindness did not resolve after the tumors were removed. None of the patients in this series were asymptomatic, although sometimes the ''skin nodule'' or ''tumor'' was not clinically worrisome to the patient and may have been an incidental finding during examination for a different reason. Although not available in all patients, when questioned, no patients reported being part of a kindred with von Recklinghausen's disease or any other phakomatosis.
The duration of symptoms ranged from 2 weeks to 240 months, with an average of 27.1 months. The overall long duration of symptoms is most likely related to the generally nonspecific nature of the initial symptoms, and patients were frequently managed without a specific diagnosis. There was no difference in average symptom duration between the genders (female 27.0 months, male 27.3 months). When contiguous adjacent anatomic sites were involved, the symptom duration was generally shorter than patients whose tumor involved a single site: ear and temporal bone sites: 7.6 months; nasal cavity and paranasal sinuses: 31.3 months versus external auditory canal: 47.3 months and nasal cavity alone: 36.5 months.
Radiographic procedures were performed in the majority of patients with ear and temporal bone and nasal cavity lesions, but generally not performed in the evaluation of skin and soft tissue cases. In general, by computed tomography scans, a mass was identified, sometimes interpreted to represent an infectious or inflammatory condition. Generally, bony destruction was not identified, although displacement, sclerosis or remodeling was seen, especially in the sinonasal cavity cases. Specifically, a central nervous system (CNS) connection was generally not identified in the majority of patients where tests had been performed (81.4%), especially in the patients who had radiographic studies interpreted as normal (40%). However, in a few cases (n = 9), bone erosion with extension of the tumor into the base of the skull was noted. Still, the tumor bulk was extraneuraxial. Magnetic resonance with a T1-weighted, gadolinium diethylene triamine pentoacetic acid (Gd-DTPA) enhancement helped underscore the nature of the lesion and the extent of the tumor.
The vast majority of tumors affected the skin of the scalp (Table 3). There was no particular predilection anatomically, with forehead, vertex, occipital, frontal, parietal and temporal locations stated. It would be difficult to extrapolate association with the cranial bone suture lines, but some of the tumors were probably overlying these landmarks.
The ear and temporal bone tumors occurred in the middle ear alone (n = 26), external auditory canal only (n = 4), temporal bone only (n = 2), and involving the temporal bone, middle ear, external auditory canal and eustachian tube combined (n = 6). Tumors were unilateral. The tumors ranged in size from 0.5 to 4.5 cm, with a mean size of 1.2 cm. There was a statistically significantly larger size for tumors which arose in the temporal bone (mean 2.8 cm), as opposed to the external auditory canal (mean 1.1 cm), middle ear (mean 1.1 cm), or mixed (1.2 cm) (P = 0.02).
The sinonasal tract tumors occurred in the nasal cavity alone (n = 17), nasopharynx alone (n = 5), frontal or ethmoid sinus alone (n = 3), sphenoid sinus alone (n = 2), and in the nasal cavity and the paranasal sinuses (n = 8), including ethmoid, frontal, sphenoid, and/or maxillary sinus (Table 3). Five patients presented with bilateral nasal cavity disease. The tumors ranged in size from 1.0 to 8.0 cm, with a mean size of 3.5 cm.
The other anatomic sites included the neck (n = 4), pelvic bones or intra-abdominal cavity (n = 3), soft tissues of the back and Achilles tendon (n = 3), orbit (n = 2), mandible (n = 1) and parotid gland (n = 1). Specific information about the size of the lesion was not available in these soft tissue locations.
The cut surface, when not submitted in multiple fragments, was composed of grayish white-tan to pink, gritty, firm to rubbery masses. The ear, temporal bone, and sinonasal tract tumors were often insinuated into bone, although the surface epithelium was intact (either squamous or respiratory, respectively). Calcifications were identified in the fragments of tissue, but may have included the bone rather than representing psammoma bodies.
The tumors were separated into a variety of histologic types and grades. Overall, there were 113 meningothelial (Fig. 1), 11 atypical, 6 psammomatous, 4 transitional, 4 anaplastic, 3 metaplastic, 2 clear cell, 2 fibrous, and 1 angiomatous meningioma(s). Overall, there were 128 WHO Grade I, 14 WHO Grade II, and 4 WHO Grade III tumors. By definition, atypical and clear cell types are WHO Grade II tumors, while anaplastic meningioma is WHO Grade III. These results are presented in tabular form based on anatomic site and as an overall percentage of each tumor type in each anatomic site (Table 4). In general, ear and temporal bone, scalp skin, and sinonasal tract tumors were meningothelial (77.4%), with almost all (95%) ear and temporal bone tumors within this category. Specific tumor types were as follows:
Atypical (7.5%, n = 5 nasal cavity, n = 4 soft tissues, n = 2 scalp); Psammomatous (4.1%, n = 2 ear, n = 2 scalp, n = 1 back skin, n = 1 neck); Transitional (2.7%, n = 2 nasal, n = 2 scalp); Anaplastic (2.7%, n = 3 scalp, n = 1 retroperitoneum); Metaplastic (2.1%, n = 2 nasal, n = 1 scalp); Clear cell (1.4%, n = 1 pelvis, n = 1 scalp); Fibrous (1.4%, n = 2 scalp); and Angiomatous (0.7%, n = 1 nasal) types.
In general, the tumors were composed of lobules and whorls of neoplastic epithelioid cells with indistinct borders. The nuclei were generally round to oval nuclei with delicate nuclear chromatin and occasional intranuclear pseudo-inclusions. Intranuclear inclusions were present in 71% of cases, although they were not always abundant or easy to find. Psammoma bodies were frequently identified (n = 45), although less frequently in the soft tissue and scalp tumors than in other tumor locations. The tumors often had an ''infiltrative'' appearance, extending into the adjacent tissues. These tissues may be soft tissue, skeletal muscle, or bone (n = 31). The bone was often remodeled, with islands of tumor infiltrating between the bony trabecular. However, the presence of ''invasion'' did not seem to correlate with patient outcome. For the non-skin primaries, the surface epithelium was intact, lacking ulceration or erosion, suggesting a slow growth rate. The other histologic types showed features unique to each entity: Atypical meningiomas have increased mitotic activity along with increased cellularity, small cells with high nuclear to cytoplasmic ratio, patternless or sheet-like Finally, four tumors showed profound pleomorphism, remarkably increased mitotic activity (mean, 24/10 HPF), and significant necrosis. These tumors developed in the retroperitoneum (n = 1), and scalp skin (n = 3). Two patients were lost to follow-up, one had died with local disease (3.4 years), and one patient was alive without evidence of disease (10.2 years).
A particularly remarkable finding was cholesteatoma in association with the meningioma in 9 ear and temporal bone cases. This finding was not associated with a high recurrence rate: when a cholesteatoma was present, 22.2% of patients developed a recurrence versus 29.6% of patients without a cholesteatoma developed recurrent tumor.
Immunohistochemical stains showed strong and diffuse vimentin immunoreactivity in all tumor cells in all cases tested (n = 78) (Table 5). Epithelial membrane antigen (EMA) was identified in 61 of 80 cases tested, but the immunoreactivity was focal and generally weak (Fig. 2), accenting only a few cells here and there, rather than yielding a strong and diffuse immunoreactivity. Other epithelial markers were also positive, including keratin (AE1/AE3, CK1 cocktail) in 24%, CK7 in 21.8%, CAM5.2 in 5.6%, and CK20 in 1.9%. The immunoreactivity was in selected cells and was a focal finding. Particularly noticeable, was a ''pre-psammoma'' like immunoreactivity with CK7. This stain highlighted the periphery of concretions in a tight, concentric whorl, suggesting a pre-psammoma body like deposition. S-100 protein stained the cytoplasm and nucleus of a few tumor cells in 19.2% of cases. Synaptophysin (4%) and GFAP (1.4%) were uncommon findings. Chromogranin was not reactive. Isolated tumor cells showed nuclear labelling with Ki-67, although only 14 cases showed increased labelling of [4% (of the 78 cases tested). These cases included all of the anaplastic meningiomas, 2 of the atypical meningiomas, and 8 meningothelial meningiomas. Four of the meningothelial meningiomas did not show mitotic figures on hematoxylin and eosin stained material, but the biopsy size was limited and so the proliferation index is probably more reliable. Synuclein was tested in isolated cases (n = 18), and was negative in all cases tested.
Treatment and Follow-up All patients were treated by partial or complete surgical excision; complete surgical removal of the tumor was frequently impossible to determine for the tumors of the sinonasal tract and ear and temporal bone, while scalp skin and soft tissue lesions were frequently completely removed. Comments about the partial or complete excision were supplied by the operating surgeons, since the fragmentation of the specimen precluded a histologic determination. A few patients were subsequently treated with chemotherapy (n = 1), and radiation therapy (n = 5), one for recurrent disease. Patient follow-up was available for 110 patients (of the 146 in the study). The following information is set in that context, with percentages based only on the 110 patients in whom follow-up was available.
Overall, the patients were followed for an average of 14.5 years, with a range of 0.1 up to 32.1 years (Table 6). The majority of patients (n = 97) were alive (n = 62) or had died (n = 35) of unrelated causes, yielding a raw disease-specific survival of 88.2%. Fifteen patients had evidence of disease at last follow-up: 2 were alive (mean 14.5 years), and 13 had died with disease (mean 6.6 years).
It is important to state that these patients died with disease, but not ''from'' their disease. After accounting for differential lengths of follow-up using the Kaplan-Meier method, median overall survival in this cohort was 28.4 years. Five-year and 10-year survival rates are 79.9% and 73.1%, respectively. Median disease-specific survival could not be estimated because of the low number of deaths due to disease. However, disease specific survival rates for specific time points are: 91.2% at 5 years, 90.1% at 10 years, 90.1% at 15 years, 86.2% at 20 years, 86.2% at 25 years, and 75.4% at 30 years. These findings show that while a few patients die of their disease soon after discovery, the majority of patients survive a long time with death resulting from other causes rather than from tumor. However, the continued progression of the disease for decades indicates that long term clinical follow-up is required to monitor disease progression or recurrence. Twenty-six patients developed a recurrence. Recurrences found within a few months of the initial surgery probably represent residual disease after incomplete excision of the primary rather than representing true recurrence. Recurrences develop in the same site and on the same side as the previous tumor, often yielding a slightly more ''infiltrative'' pattern than the initial specimen. Tumors which extend into the base of the skull (ear/temporal bone and sinonasal tract tumors) tend to be very difficult to eradicate surgically, and tend to remain indolent for many years. While overall patients with recurrences experienced a good survival (median, 15.6 years), 42.3% of these patients died of their disease, 32.6% within 5 years of initial diagnosis. This contrasts with only 1.2% of patients without recurrences dying of disease within 5 years of diagnosis (P \ 0.001). Overall survival was lower among patients with recurrences but this difference was not statistically significant (median 31.4 years among patients without recurrence versus 15.6 years among patients with recurrences, P = 0.174). Recurrence appears to increase the risk of dying from disease.
The specific anatomic site of tumor development suggests that 5-year disease-specific survival is slightly higher among patients with scalp skin based tumors (93.1%) and tumors of the sinonasal tract (93.8%) in comparison to tumors of other soft tissue sites (85.7%) and ear and temporal bone (88.4%). However, as there are only a few cases within each group, a statistically significant result could not be measured (P = 0.763).
The histologic type does seem to correlate with dying from disease (P = 0.0378), although this finding could not be controlled for recurrence specifically. There were no deaths from disease among the 6 patients with psammomatous or clear cell type tumors; ten among the 87 patients with meningothelial meningiomas (5-year survival 91.5%); and 1 among the ten patients with atypical tumors (5-year survival 87.5%). Although five-year survival was only 50% among patients with anaplastic tumors, this represents one death from disease among two patients.
As would be expected, a greater percentage of patients die from their disease as the grade of the tumor increases. Fiveyear disease specific survival decreases from 92.4% (Grade I) to 88.9% (Grade II) to only 50% (Grade III). Further, it can be seen that the overall duration of survival also decreases as the grade of tumor increases. However, because only 13 patients had Grade II or III tumors, these differences did not reach statistical significance (P = 0.052 for overall survival; P = 0.133 for disease-specific survival).
Data about the completeness of the excision is difficult to comment on, as this is not a parameter which can be easily assessed nor confirmed. Therefore, while it may seem that the survival is shorter for patients with a partial removal, with only a single death from disease in this group, it makes a meaningful interpretation difficult.
Based on the data in this study, the overall 5-year survival for males and females is not different: females, 79.0%; males, 80.8%, P = 0.974. However, 14.4% of women died from their disease within 5 years of diagnosis, while only 1.9% of men died of their disease. This information suggests that female patients have an overall worse prognosis in comparison to male patients (P = 0.017).
It is difficult to define outcome based on age, but using a 40 year age cut-off, there is a statistically significant difference in patient outcome and length of survival. Two of 41 patients diagnosed before age 40 died of disease (5-year survival 97.2%), compared with 11 of 69 patients diagnosed after age 40 (5-year survival 87.7%, P = 0.049). Overall survival was also greater among younger patients (P = 0.005), indicating that age may be an important prognostic factor. It goes without saying that patients with necrosis (n = 5), tend to survive for a shorter period. Four of the 5 patients with necrosis died during the study (5-year survival 40.0%), two from their disease (5-year diseasespecific survival 50%). These data indicate a poor prognosis for patients with necrosis when compared with other patients (overall 5-year survival 81.8%, P = 0.010; 5-year disease-specific survival 92.9%, P = 0.009).
Overall survival is similar for patients with and without an increased proliferation index during the first 5 years after diagnosis. Five year overall survival is 66.9% among those with a proliferation index \4, compared with 72.7% among those with a proliferation index C4. However, patients with an increased proliferation index ([4) have higher death rates overall, with median survival of only 10.2 years compared to 31.4 among those with proliferation index \4, and 10-year survival of 54.6% compared with 62.2% among those with proliferation index \4 (P = 0.047). Disease-specific survival was also lower among those with an increased proliferation index, but this difference did not reach statistical significance.
There are no significant differences between patients who have cholesteatoma and patients who do not have cholesteatoma with respect to all-cause survival (P = 0.125) or disease-specific survival (0.925).
A final comment regarding patient management and outcome: the anatomic sites of the head and neck, especially the ear and temporal bone, nasal cavity, paranasal sinuses, and even the scalp skin can result in disruption of the adjacent tissues. A number of patients developed mastoiditis, sinusitis, or scalp abscesses. These complications resulted in seeding to the peripheral blood to cause systemic sepsis, which resulted in the patient's death. Therefore, in these patients, even though they may not have residual tumor, the deaths are a direct result of the management for the tumor. Therefore, it is important to consider the inherent risks of surgeries in these vital structures.
In summary, older patients, patients with necrosis, and patients with increased proliferation index had significantly reduced overall survival. Women, patients older than 40 years at diagnosis, and patients with recurrence or necrosis were significantly more likely to die of disease (P \ 0.05 log rank test; Table 6). When all potential prognostic factors were evaluated jointly using Cox proportional hazards regression, the only independent predictor of overall survival was age group; when non-significant variables were removed sequentially from the model, the remaining significant prognostic factors were age group, anatomical site and presence of necrosis (Table 7). Mortality was more than three times higher among patients whose age was greater than 40 years than among those younger than 40 years at diagnosis (hazard ratio 3.6, 95% confidence interval 1.8-8.0). We had hypothesized that grade would be a stronger predictor of survival than either necrosis or proliferation index because both of these variables are part of the grading criteria. However, at least in this population, necrosis is a better predictor of all-cause mortality than either grade or proliferation index. The best predictor of disease-specific survival was recurrence. Adjusting for other prognostic factors, the estimated hazard ratio associated with recurrence is 19 (95% confidence interval 3.5-104.3). After adjusting for recurrence, gender, age and necrosis were no longer statistically significant, perhaps because women and patients with necrosis were twice as likely as other patients to experience recurrence. Patients with proliferation index of 4% or higher were three times more likely to die of disease than patients with proliferation index below 4%, although given the small number of patients this did not reach statistical significance. Patients with unknown proliferation index were 90% less likely to die of disease. When nonsignificant variables were removed from the regression model sequentially, differences among sites reached statistical significance. Compared with tumors of the ear, the estimated risk of dying of disease was more than 7 times higher for patients with nasal tumors, and more than 60% lower for patients with skin scalp tumors.
Arachnoid cells (arachnoid granulations, meningiocytes, meningothelial cells, pacchionian bodies) are thought to arise from neural crest. They normally line the inner aspect of the arachnoid membrane, and fill the cores of the arachnoid villi that project into the lumens of dural veins and venous sinuses. Increasing evidence supports the development of meningiomas from arachnoid cap cells, with different mechanisms to suggest how extracranial meningiomas arise:
(1) Arachnoidal cells are present in the sheaths of nerves or vessels where they emerge through the skull foramina. (2) Displaced pacchionian bodies become detached, pinched off, or entrapped during embryologic development in an extracranial location. (3) A traumatic event or cerebral hypertension that displaces arachnoid islets. (4) An origin from undifferentiated or multipotential mesenchymal cells, such as fibroblasts, Schwann cells, or a combination of these, perhaps explaining the diverse pathologic spectrum found in meningiomas. Consequently, by one mechanism or another, arachnoid cells are identified outside the neuraxis and give rise to meningiomas in extracranial locations, including ear, sinonasal tract, scalp and soft tissues [3,8,10,14,[19][20][21][25][26][27][28][29][30][31][32][33][34][35][36][37][38][39].
Up to 20% of intracranial meningiomas may have extraneuraxial extension [10,14,[19][20][21][25][26][27][28][29][30][31][32][33][34][35][36][37][38][39][40][41][42], including the skull, scalp (all cutaneous sites), orbit, upper airway involvement (nasal cavity, paranasal sinuses, nasopharynx), soft tissues, and ear and temporal bone. However, when the scalp, orbit, sinonasal tract, oral cavity, and soft tissues are excluded, the incidence decreases to less than 1% [19-21, 25, 27, 37, 38, 40, 43-45]. Based on our findings, we believe the majority of our cases are arising de novo from multipotential stem cell, although a number of the cases within the ear and temporal bone and sinonasal tract may develop from misplaced remnants of the pacchionian bodies. While there are a number of reported cases of ''ectopic'' or ''heterotopic'' meningiomas, most of these were reported prior to the modern improvement in radiographic techniques which serves to more definitively exclude intracranial disease. It is important to exclude an intracranial component radiographically or during surgery to yield the best possible management and follow-up. The possibility of an intracranial tumor must always be considered if long term management is to achieve its intended goal [8,10,19,20,34,46].
Both genders were equally affected in this clinical series in the aggregate (74 females and 72 males), although ear and temporal bone tumors were more common in females than males (2:1) and males were more commonly affected than females for skin scalp lesions (1.7:1). This is in contrast to the reported predilection of intracranial tumors for females [1].
Females tend to be significantly older (mean 45.6 years) than males (mean 39.6 years, P = 0.01), with the exception of soft tissue tumors where males tend to be older (mean 42 years) than females (mean 31 years). Further, patients with tumors of the skin scalp tend to present at a younger mean age (36.2 years) than tumors which develop in the ear and temporal bone (mean 50.1 years) and sinonasal tract (mean 47.1 years). This may be due to a greater facility for clinical identification of a skin scalp lesion rather than an inner ear or sinonasal tract tumor. The overall average age at presentation of our patients (42.4 years) is not dissimilar from the middle-aged figure used for intracranial meningiomas without any extracranial extension. No patients in this clinical series had any syndrome-associated meningiomas.
Meningothelial meningiomas are the most common tumor type identified in this clinical series. Interestingly, 95% of ear and temporal bone lesions were meningothelial, but only 43% of soft tissue sites showed this pattern. This suggests that a greater degree of vigilance is required to render this diagnosis when looking at tumors in such unusual anatomic sites as the parotid gland, Achilles tendon or the pelvis. Nearly one-third of meningiomas in the soft tissues are atypical, further emphasizing the need for careful consideration of this diagnosis in unusual locations. Atypical meningiomas [1,22,47] were diagnosed in 11 cases in this series. The soft tissue sites, including the neck and pelvis, were most commonly affected, but the sinonasal tract and scalp skin were also affected by a few tumors each. Interestingly, atypical histology did not portend a worse outcome for these extracranial sites as it may for intracranial lesions [19,20,24]. It may be they are discovered at an early stage of development, allowing for earlier surgery than their intracranial counterparts.
Axiomatic, the immunohistochemical profile of extracranial meningiomas is indistinguishable from intracranial lesions; therefore, no ancillary technique can separate direct extension from ectopic/extracranial meningioma. Surprisingly, staining for CK7 was detected in 21.8% of cases. Expression of CK7 in meningiomas has been previously reported [19,20], specifically in secretory meningioma [7]. In our series, expression was localized to prepsammoma bodies, similar to those observed in secretory meningiomas, but secretions were not seen in any of the cases in this series. Therefore, it may be that the specific CK7 staining pattern may help with yielding a diagnosis of meningioma in these extracranial sites.
When receiving these cases in consultation, 24% of the cases were incorrectly diagnosed, with a larger proportion submitted as ''undetermined.'' This results in inadequate or inappropriate management, especially if an intracranial primary has not been excluded. The differential diagnosis of extracranial meningiomas includes a variety of benign and malignant neoplasms dependent upon the anatomic site of involvement. Paraganglioma, schwannoma and metastatic carcinomas are the most frequent misdiagnoses for ear and temporal bone tumors [3,19,48]; carcinoma, melanoma, olfactory neuroblastoma, and aggressive psammomatoid ossifying fibroma for sinonasal tract lesions [6,20,[49][50][51]; dermatofibroma, melanoma, fibrosarcoma, leiomyosarcoma, and synovial sarcoma for soft tissue and skin lesions, especially for Grade II or III tumors, and for the non-meningothelial types of meningiomas.
The general histologic features and immunohistochemical findings can usually separate between these tumors [3,6,9]. Specifically, paraganglioma will show more of a nested, Zellballen pattern, with granular cytoplasm and chromogranin, synaptophysin, CD56, and S-100 protein immunoreactivity (highlighting their respective compartments). Schwannoma will demonstrate different degrees of cellularity, with areas of myxoid change, perivascular hyalinization, wavy nuclei, and strong, diffuse S-100 protein immunoreactivity. Metastatic carcinomas or carcinoma in general tends to show more pleomorphism, a much higher mitotic rate, and will be immunoreactive with a variety of keratins. Whereas psammoma bodies can be seen in papillary carcinomas (thyroid, lung, ovary), the growth pattern of meningioma tends not to be papillary in these extracranial sites. Melanoma (cutaneous or mucosal) can mimic any tumor type, replete with intranuclear cytoplasmic inclusions and whorled architecture. However, pigment, prominent, irregular nucleoli, and melanocytic markers (including S-100 protein, HMB-45, melan-A, MART1, tyrosinase, microphthalmia transcription factor) will help to make the distinction. Olfactory neuroblastoma occurs with direct proximity to the cribriform plate region, usually maintains a lobular growth pattern of small to intermediate cells with scant cytoplasm, has a fibrillary background, exhibits rosette and/or pseudorosette formations, and displays characteristic immunohistochemical features (chromogranin, synaptophysin, neuron specific enolase and CD56 tumor cell staining and S-100 protein sustentacular staining) that are easy to separate from meningiomas. An aggressive psammomatoid ossifying fibroma is an uncommon lesion which may be confused with meningioma because both lesions occur in young to middle-aged patients and may be associated with proptosis. Both lesions have abundant psammoma bodies, but meningiomas lack associated osteoclasts and osteoblasts. Furthermore, the background stromal pattern is storiform and more compact than a meningioma and does not have the same immunophenotypic characteristics as a meningioma. High grade tumors with pleomorphism, necrosis, and increased mitotic activity with a loss of architecture, may be very difficult to separate from skin or soft tissue sarcomas. Usually, small foci of recognizable meningioma will be present, showing a vague whorling pattern. The weak EMA may help, although EMA will be reactive in synovial sarcoma. Specific sarcomas will generally exhibit specific histologic, immunophenotypic or molecular results that allow for their categorization apart from meningioma.
In general, the prognosis of extracranial meningiomas appears to be excellent, with an overall median survival of 28 years. This is tempered by the specific anatomic site, histologic type, tumor grade, gender, and age of the patient. The recurrence rate for meningiomas after total excision varies from 7% to 84% depending upon the number of years of follow-up [13,19,22,25,47,52,53], with our rate of 23.6% falling within this range. This is similar to intracranial meningiomas which have a recurrence rate of up to 20% and a mean survival around 7 years [1,25,52].
In this study, there was little difference between the 5 and 10-year disease-free survival rates (91.2% versus 90.1%, respectively), indicating that once the patients survived disease free for 5 years, they were unlikely to die with tumor. Furthermore, this same trend was noted at 15, 20, 25 and 30 years (90.1%, 86.2%, 86.2%, and 75.4%, respectively). Nine of the thirteen patients who died of their disease, died within the first 5 years. The remaining four patients died at 5.9, 15.6, 17.3, and 28.4 years, respectively. One patient was still alive with local disease at 27.5 years. This finding supports the slow, indolent growth of extracranial meningiomas, but also suggests very long term clinical follow-up is necessary to achieve these outcomes. Our findings are different from meningiomas in general in which the recurrence rate increases with protracted follow-up (6% recurrence at 5 years and 20% recurrence at 15 years [52]). Therefore, meticulous surgical extirpation of extracranial meningiomas is important to minimize the recurrence rate, without the necessity of adjuvant therapy. While surgery is the treatment of choice, there are a number of challenges due to the invasiveness of the tumors and the complexity of the anatomy within the sinonasal tract and ear and temporal bone, although scalp and soft tissues lesions are no less difficult to remove if they are adjacent to vital structures. It may be necessary to utilize a multidisciplinary approach with a combination of intracranial, temporal bone, maxillofacial, and skull base techniques to achieve total resection, possibly including widely exenterative procedures to achieve this end [10,16,29,31,37,[54][55][56][57][58][59][60][61].
Five patients received radiation in this series. Radiation therapy has been suggested to yield a possible improvement in survival in meningiomas of the central nervous system [13,53,57,62]. However, this does not seem to be the case in this series, although the numbers are limited: One patient with recurrence, died of disease (0.5 years); three patients are alive without evidence of disease (mean 14.2 years), and one patient died without evidence of disease (15.6 years). Further evaluation may be necessary in a larger group of patients.
In our cases, and in those of the literature, when recurrences developed, they usually arise in the same anatomic site as the primary lesion and depending on time interval, may represent residual disease rather than recurrent tumor. Metastatic disease did not occur in any of our patients nor did we find any convincing cases in the literature of extracranial meningiomas resulting in metastasis.
Meningiomas that arise outside the skull are for the most part benign tumors. Symptoms are associated with the location of the lesion. Imaging studies may provide some suspicion of diagnosis, but confirmation relies on pathologic examination. The mainstay of treatment is complete excision of the tumor by appropriate approaches, often requiring multidisciplinary approach. The histologic features, while generally ''characteristic'', especially for meningothelial tumors, may be atypical, requiring separation from other tumor types, by both histologic and immunohistochemistry studies. Separation from other tumors is essential as extracranial meningioma seems to have an excellent long term prognosis with only limited recurrence.
Extension into eustachian tube 1
Size (in cm) Range 0.5-8.0 Mean 2.3
* P value from log rank test
* CI confidence interval ** Variables in the reduced model were selected from among the variables in the full model using the backwards stepwise variable selection algorithm implemented in SPSS version 14. Variables with P \ 0.10 were retained in the final model
The opinions or assertions contained herein are the private views of the authors and are not to be construed as official or as reflecting the views of the
Historically, cultural accounts and descriptions of blood banking in Britain have been associated with notions of altruism, national solidarity and imagined community. While these ideals have continued to be influential, the business of procuring and supplying blood has become increasingly complex. Drawing on interview data with donors in one blood centre in England, this article reports that these donors tend not to acknowledge the complex dynamics of production and exchange in modern blood systems. This, it is argued, is congruent with nostalgic narratives in both popular and official accounts of blood services, which tend to bracket these important changes. A shift to a more open institutional narrative about modern blood services is advocated, as blood services face current and future challenges.
Blood donation has long been viewed as a dramatic symbol of interdependence. In Europe, the emergence of formal blood banking systems in the post-war era was inextricably associated with notions of altruism and solidarity. The association of blood banks with altruism, solidarity and imagined national communities continued well beyond the post-war years (Rabinow, 1999). When Richard Titmuss (1997: 311) wrote his famous analysis of blood donation as a 'gift relationship', based in part on empirical research undertaken in Britain in the late 1960s, he could still write in universalist terms that blood services represented 'one practical and concrete demonstration of fellowship relationships institutionally based in Britain in the National Health Service and the National Blood Transfusion Service'. He wrote at a time when donated blood was predominantly used whole; the National Blood Service (NHS) was held to be a practical symbol of mutual interests between citizens in Britain, and trust in the capacity of NHS authorities to supply safe blood was not publicly questioned.
Blood banking was set to change. By the mid-1970s, blood could be stored and transported more easily, facilitating the beginning of a more flexible system for supplying blood. The development of fractionation techniques allowed for blood to be broken down into components and reconfigured, leading to the great majority of donated blood being used in the manufacture of blood products (Farrugia, 2006). Blood was to become the raw material for an emerging industry that was to expand into markets throughout the world (Starr, 1998). The flexibility allowed for by these developments had been a goal of pioneer blood bankers, to allow for provisioning in mass emergencies. As the blood products industry developed further, blood plasma in particular became a commodity that could be imported and exported more easily, and the supply chain became more complex to manage (Starr, 1998). We can see in retrospect that, from the 1970s point onwards, the work of supplying blood products would be in some respects at odds with the traditional image of blood services, especially with the notion of national blood banking.
Notwithstanding these substantive changes, blood services received relatively little public and political scrutiny until the 1980s, when the HIV/AIDS crisis brought the problems, risks and hazards inherent in blood banking to wider attention. Some officials and organizations were prosecuted for their failure to act to protect the recipients of blood after the risks of HIV infection were known (Casteret, 1992;Rabinow, 1999;Starr, 1998). Although others emerged relatively unscathed from the crisis, with it came the dawning recognition of the possibility of vested interests in the supply of blood products.
More recent sociological writing about blood banks is anchored with an awareness of the unfolding consequences of the transmission of the HIV virus through blood, and of the risks associated with receiving blood products. This in turn opens up scrutiny of the impact of ongoing processes of modernization, globalization and commercialization in blood systems. Even where blood services are organized by national state authorities, the impact that global commerce, disease and travel are such that they can no longer be regarded as having national boundaries (O 'Neill, 2003). These changes may be theorized in terms of the threat they pose to the association of blood banking with ideals of solidarity and altruism (Waldby and Mitchell, 2006). The discussion of the apparent erosion of mutuality in this context can be linked to a broader consideration of the extent to which social life is becoming 'de-mutualized', and the ways in which there is resistance to that process (Williams, 2002). However, the question of blood safety has long been on the agenda for blood bankers, who from the outset had to countenance and manage the risks of recipients contracting infections from blood (Berridge, 1997).
While the blood service in the UK is in the process of modernizing to 'create a service fit for the 21st century', this process is focused on organizational and technical innovation (NHS Blood and Transplant, 2007: 8). Blood donors are widely described as 'altruistic', yet little research has been undertaken to explore the broader social meanings of donation today. This article focuses on blood donors' narrative accounts, and explores their potential significance for contemporary blood services.
Recent sociological work about blood donation offers some theoretical and methodological approaches around which the discussion of the data presented in this article can be organized. I begin with Healy's work on cultural accounts of 'procurement' of blood and organs. Healy approaches the question of blood donation by exploring how the meaning of blood is shaped by organizations and the regimes surrounding them. Drawing on large scale quantitative data, he began by exploring how donation patterns in state run, red-cross and blood banking regimes in Europe differ significantly from one another, developing an 'institutional perspective on altruism', by showing how these 'collection regimes produce their donor populations by providing different opportunities for donations' (Healy, 2000(Healy, : 1633)). More recently, looking at organizations responsible for 'procuring' organs, Healy (2006) shows how the production of a 'cultural account' of donation is an important facet of their work. I take from this approach a reminder that cultural accounts of blood donation mediate the meanings of such projects; these meanings have practical consequences for the procurement and supply of such tissues.
While civic virtue is often stressed in the public and political accounts of blood donation, we can consider this practice as simultaneously a 'private' and a 'public' act (Valentine, 2005). Bearing in mind the recent history of gay men's exclusion from donating blood, Valentine's (2005: 115) analysis centred on a critique of the construction of blood donation as a 'participatory space of belonging'. One element of her critique is the consideration of the impact of practices of exclusion/inclusion for eligibility as donors. Such practices are central to the operation of blood services, and have become progressively so since 1983. Valentine makes the point that private experiences and accounts of donating blood -or not donating -may differ from public and more expressly ideological narratives.
Eliciting 'private' accounts about blood donation has its methodological difficulties, however. This is partly by virtue of the status of venepuncture as a ubiquitous technology in hospitals. In addition, many of the uses of blood for diagnosis, research and treatment, are long established and may be seen as unproblematic by patients and staff alike. Pfeffer and Laws (2006) show how people in a London teaching hospital tended to 'turn away from' the mundane technology of venepuncture, and to take it for granted. The respondents in Pfeffer and Laws' study -who included patients and staff -often saw submission to venepuncture within hospital as appropriate, or expected; this acceptance of the taking and circulating of blood can be seen in terms of a 'law of the place' ( de Certeau, 1988: 118, cited in Pfeffer and Laws). Their analysis also suggests that patients in their London hospital site had expectations surrounding the use of blood: they expected that the blood would remain within the hospital, that blood test results would be fed back to patients, and that the blood would not be used by commercial bodies. These expectations were sufficiently strong and consistent that Pfeffer and Laws (2006: 3022) describe them as an 'implicit contract'. This and other recent work underlines the significance of attending to the nuances and complexities of donors' accounts.
This article focuses on an analysis of interviews with blood donors to one of the UK's four blood services, the National Blood Service in England and North Wales (NBS), whose primary remit is to supply blood components and other tissue products for use in medical care in the NHS. 1 The aim of these interviews was to revisit and explore some of the frameworks and assumptions of blood donors, given the very limited sociological literature on these. One driver for doing this was that 'Titmussian' ideas about blood donation were being invoked in policy discussions about the development of a new national biobank in the UK (Tutton, 2004). Although the NBS has a remit to conduct research and development on current and novel applications of blood, I did not intend to focus specifically on this here. 2 Rather, these interviews sought to explore the understandings and expectations of contemporary blood donors, in as open ended a way as possible. A full account of the development and formulation of these interests can be accessed online (Busby, 2004b).
The fieldwork for this project included participant observation and discussions with donors and staff of the NBS at the same blood centre, and at several other locations in same UK city. Over the course of four months, many discussions, and conversations took place, and interviews were conducted with NBS donors at one city centre blood centre. In addition, observation, discussion with centre staff and analysis of NBS publications was undertaken. This article draws on data from 26 interviews with donors, who numbered approximately equal numbers of men and women, ranging in age from their late teens to late 60s, with most being in their 30s, 40s and 50s. (The identity numbers given refer to the sequence of all discussions with donors held in this centre. The analysis refers to 26 donors with whom longer interviews were held.) Before the fieldwork began, the proposed arrangements for interviews were reviewed by an NHS Research Ethics Committee, which agreed that it may go ahead.
The aim of this approach being to develop a narrative account of the interviewees' experience as donors, and their view on this involvement, these interviews were 'semistructured'. Interviews often began with a question about when a donor had first donated blood -'Can you tell me how you first became a blood donor?' Instead of the fixed questions which were often expected, a topic guide was used to guide the interviews, covering the following themes: what is done with the blood; views on payment for blood donation; information; concerns or worries about giving blood especially at the outset. Donors were also asked about their views on research, beginning with a question about whether they would see giving blood for research differently from giving blood to help people directly. These issues were not pursued with detailed questioning, if interviewees did not seem comfortable with discussing them. There were two grounds for this reticence on my part, one methodological, the other ethical. The primary aim was to gain an overall narrative account of the donor's perspectives. Too many questions, or questions that were too specific being repeated, seemed to impede the narrative flow of the account. I was constrained too by a concern not to intervene in people's understandings of blood systems. I knew little about the research uses of blood. I was not able to answer questions that people might (and did) put to me in response to my own questions. Given that both I and the NBS considered their blood donations to be important, I did not wish to dwell on research uses, if the donor wished to avoid talking about this. At the end of the interview, people who had been donating blood over some years were invited to talk about how the experience had changed over time.
The approach to analysis can be described in three stages.
Grounded theory techniques were used to explore donors' accounts in the discussions and interviews conducted in the course of the early weeks of the fieldwork (Strauss and Corbin, 1990). Several key phrases emerged from these analyses of preliminary discussions with donors: for example, 'blood bank', 'limits of expertise'. These sensitizing themes influenced how the work was taken forwards in a number of ways. The first of these themes refers to an image of traditional blood banks that anchored accounts of donating blood for the NBS, and is explored further in this article. This image of blood banks is predicated on a model in which such facilities were organized on a local, regional or national basis. It is the idea of 'self-sufficiency' in blood supplies within a community, rather than the exact boundaries of the blood bank, that is important here. 3 The second theme refers to donors' responses to my many questions about blood, and how it is used: interviewees often drew my attention to what they felt were the limits of their expertise. A common response to my questions was to indicate that donors could not be expected to know about or understand research: in asking them I was 'going beyond' their expertise. In the early interviews my questions about what kind of information people looked at, and what they thought about the uses which were made of blood frequently elicited fairly brisk, even dismissive replies: 'I give my blood and now it's gone and I have no views' said one man (NBS 17). Another captured the spirit of these discussions when she said that she had 'no idea' what the blood would be used for (NBS 25). Many commented that they trusted the blood service and the relevant authorities to make the best use of their blood, and did not feel the detail to be their concern. These recurring comments, often couched in terms of 'trust' challenged my assumption that I could analyse people's decisions in relation to the information they had access to about blood services seemed to have been misplaced. The importance of information as a keystone for the consent process for tissue donors has been prominent in UK government policy in recent years. Yet it emerged that in this particular case, of blood donation, donors did not see information as the basis of their consent for their blood to be used. Most donors stressed that the details of what happens to blood after it is donated are not of pressing concern. One long-term blood donor explained his difficulty with my questions to me in this way: 'The problem is once I've given it what they the NBS do is it's up to them. I mean once I walk through the door I forget about the NBS' (NBS 100). This constitutes a methodological problem for the researcher who would like to understand the donors' point of view about the uses of blood through an interview. As Hoeyer (2003) has observed of the process of researching with donors, there is a sense in which an interviewer constructs a request to the donor to discuss something which is a physical practice often performed by turning away.
The second stage of analysis involved the use of a data matrix, to summarize the donors' circumstances and the context of their interview account -for example, the length of time they had been a donor, and the particular issues which they had raised. 4 Added to the biographical and contextual summaries for individuals interviewed, notes on specific points in the transcribed data were coded and briefly noted. A summary note of their views on the uses of blood was made in each case. This matrix was also used to include other discussions held in the course of observation at the blood centre. Having been coded and written manually, these data were stored on a spread sheet, which enabled them to be indexed for subsequent recall and analysis. Each interviewee, along with other donors with whom I had brief discussions, was given a unique number, which is the number referred to in this text. Where the interviews are cited in this article, 'I' stands for the interviewer (myself), and R stands for the respondent.
Once these donors' accounts were thus summarized, the final stage of the analysis involved relating the particularity of these findings back to wider themes in the literature. These wider themes from the literature, elaborated in the section on the sociological literature above, include the significance of mutuality in tension with the demutualization thesis in this context. Healy's (2006) argument that organizations procuring human tissues necessarily organize 'cultural accounts' of donation that support their work, and Pfeffer and Laws' (2006) emphasis on the implicit nature of the agreements that exist between patients undergoing venepuncture and hospitals, are used to draw out the significance of these donors' accounts.
Most donors attributed their initial involvement in blood donation to having the opportunity to give blood at a local venue, rather than to a more active moral rationale in the first instance. Next, they talked about the influence of other individual donors who they knew, and sometimes expressed a sense of obligation in this context. For example, some of the interviewees talked in terms of 'replacing' someone who had been a longstanding blood donor, but had become unable to continue due to ill-health, or someone who had died. When asked more about their continuing to donate blood, donors' responses became in a sense more their own stories, with a range of reasons for doing so being expressed. For some this sense of responsibility began with awareness of someone in their own family who had been ill, so that the blood donation was imagined as being 'to help somebody like my mum'. But it had then extended to 'future patients'. Reasons for donating blood are intertwined and not easily separated.
The satisfaction which could be gained from giving blood was often mentioned: one woman, who first gave blood because her teacher at college suggested that she and other students do so, pointed out that 'there's not much else you can do that gives you that kind of buzz'. Another talked about blood donation as a sociable experience which she looked forward to (NBS 41). She said that she enjoyed 'being part of something' as she lay down quietly on her own but with others similarly engaged nearby. Although it might be sociable, enjoyable, an important element of the experience for some was that it offered a refuge from commercialization: part of what makes it special, she continued, is 'the idea that nobody can make any money out of … it just makes you feel good' (NBS 41).
Being able to help was a source of satisfaction for many, when contrasted with other situations where they might feel less able to help. Others couched their involvement in blood donation more in terms of duty or obligation. Blood donation (of course) is something which is done with others in mind. However it was not seen as a sacrifice: many pointed out that giving blood did not cost them much. For some, the relationship to others was invoked through a dramatic, usually traumatic reason for giving blood: a family member being ill in hospital, a friend being diagnosed with cancer. For many though the awareness of blood being needed was expressed in terms of a growing awareness of the vulnerability of others when they encounter accidents, serious illness, unexpected operations -and an awareness that these unexpected disasters could happen to themselves and to those close to them. Thus the interdependence symbolized by the possibility of needing donated blood at a time of catastrophic illness or accident featured in virtually all of these accounts.
Among this group, the most common occasion for coming in to make a first blood donation was related to people's work, in that they had either attended a session with a work group or colleague or at NBS mobile sessions at their place of work. The days of mobile visits to large industrial workplaces were fondly remembered by some of the older donors, mainly men now in their 60s. Among those interviewed were men who had worked in the docks, in large factories, in big Royal Mail sorting offices in the city centre or for other companies since disappeared or privatized, like BT. One of these explained his experience of being of a blood donor over the years as follows:
A sense of wanting to, a sense of being part of this system which keeps the blood available because it might be me one day who needs it or someone in my family and if it isn't there then that's going to be a problem. So it's just something I think we should all do really. (NBS 10) For these donors, there was a sense that the blood centre had become a participatory space and one that was no longer available through their work associations.
Many donors (approximately half) specifically used the phrase 'blood bank' in the course of these interviews, and others referred indirectly to this idea: It's really a case of why not … I know people who have done it and I have a friend who has leukaemia, or who had leukaemia shall we say and you know you hear about other people who are blood donors because their lives have been saved as a result of being recipients of blood ... All right I may be doing it the other way round, I hope that I never have to be the recipient but I'm willing to give because if I don't need it others will. (NBS 83) Here the arrangement of donating blood is envisaged as being equally of use to oneself and to others. It is envisaged as an arrangement in which the risk that we may need blood is shared with others who may likewise become vulnerable. The mutuality of such an arrangement is enhanced by the thought that the health screening for blood donation can function as a health check. Often donors said it was recognizing the possibility that they might need blood themselves which enabled them to accept repeated, detailed and to my mind intrusive screening questions before they gave blood. Donors often pointed out that the fact of their being asked these detailed questions would be a source of reassurance that the blood was as safe as possible if they themselves needed blood in the future.
Notwithstanding their knowledge of the part played by the blood service in passing on infected blood to some recipients before HIV tests were implemented for blood donors, these donors entrusted their blood to the blood service. The possibility of anyone needing a blood transfusion unexpectedly was often present in these accounts, and therefore donors also saw themselves or those close to them as candidate recipients. Among these donors then, we can say that the notion of blood banking as a mutual arrangement, as part of a strategy of pooling resources for responding to the risk of catastrophic illness or accident, were prominent. A narrative about mutual interests was still possible for these donors. This is so despite the risks that have come to be associated with receiving blood, and which risks are emphasized to donors in the course of lengthy donor screening procedures.
The donors' accounts reflected the idealistic ethos that blood should be universally available. However, while performing their role as donors in this context, some contradictions do occur: first, while no-one seriously suggested that certain categories of people should not receive blood, ambivalence about those who do not give was sometimes expressed. Comments about those who should not receive blood were always expressed laughingly, jokingly: 'I hope that people who won't give blood don't get it (laughs). It sounds awful that but I think God you know I give blood, and if I want it it's there' (NBS 8). Within this system, it was felt that you could not reasonably specify exactly who should have the blood, despite the fact that there might be some people who you wished could not have it. Likewise, most felt that you could not specify what the donated blood could be used for. This was not a case of having no views on priorities for medical treatment and research. It was rather that the nature of the transaction was one of entrusting the blood to the bank to make the best use of.
Blood services were understood to be part of the NHS -as indeed they are at a statutory level. Donors felt that their voluntary donation was an intrinsic part of that system. Sometimes exceptions to the ideal of entitlement were mentioned in this context. There was for example the issue of patients in private hospitals: for some this posed a challenge to the ethos of the universal system, and it was felt that perhaps they should pay for the blood. Similarly, it was often said that blood should be used in this country. However, the case for these candidates (private patients, foreigners) to be excluded from receiving blood tended to evaporate if it was thought that surplus blood might go unused.
The extent to which the blood service is embedded in the NHS was emphasized in many of the donors' accounts. Often this point was made in talking about instances in which blood might be required for accidents or serious conditions. It was expected that emergencies were dealt with by the NHS. 'Operations, transfusions, babies ' were the examples usually given of points of crisis when blood might be needed. None of the examples given by any of the interviewed donors made reference to people being treated in the private sector. As with Pfeffer and Laws' account of patients' understandings of the circulation of blood in an NHS hospital, this was an implicit understanding, that only became clear when exceptions or 'scandals' were mentioned.
Comparisons were made with other countries in which emergency treatments were paid for by the patient, whereas here:
at least one thing at least if you're seriously ill you get to, you don't have to worry about the bill at the end because you know it's there. I mean if you're really seriously ill, a road accident, you're seen to straight away and when you work it out there's blood there. (NBS 100)
In the more discursive accounts the importance of the relationship with the NHS was made explicit, as was the ethos underlying voluntary blood donation. So here in my discussion with a female donor in her 30s: I: what's special or particular about blood donation that money shouldn't come into it? R: I think it comes down to that word donation in my head, it's like something you feel like, you feel good about it because it's a thing you do sort of in a voluntary sort of process and you enjoy doing because it's something you give the National Health isn't it? (NBS 41) It was, one respondent replied, a problem of where to draw the line -the consequences of paying for blood at point of donation would cascade through the system altering it substantially: People would say oh we've got to pay for this blood, why not pay for organs. And basically you're coming into a private health system or paid for private health system rather than a national health system and unfortunately, I'm being, by social inclination I'd prefer to have a national health service and a national blood transfusion service that's funded by the people without paying. (NBS 57)
Occasionally a scandal was referred to in which blood or organs had been traded in this country or abroad. Moral disapproval about these breaches of expectation was clear cut in these cases. But the extent to which blood is moved around different sites, fractionated, reconstituted as blood products, traded in an internal NHS market or imported from outside the national blood service, was barely present in these accounts. The importing of plasma products sourced from other countries, notably from the USA, was not discussed. These donors primarily characterized the use of blood as 'for emergencies'. This is in contrast to the reported use of blood by the blood service in England and Wales, according to whom approximately 8 per cent of blood is used in emergencies. 5 As we have seen, Healy proposes that organizations responsible for procuring human tissues for medical applications necessarily have to mobilize and present a narrative that promotes these activities as desirable. In this case, the accounts of blood donation produced by the NHS, with politicians and publics, tend to avoid acknowledging the complexities of producing and exchanging blood products. This information is not secret; it is included in corporate communications -publicly accessible to someone searching for it online for instance, but nor is it given any space in communication targeted at publics, patients or donors. This facet of the NBS's work rarely appears in the now many and varied corporate communications to donors. Another example of this absence of acknowledgement is to be found in the graphs depicting national blood stocks, of fresh donated blood in England, on a weekly basis on its website. 6 The website does not offer analogous graphics on the import and export of plasma and blood products that are likewise central to the business of a modern blood service.
When these donors entrusted the NBS with their blood, they did not expect to be involved in making decisions or drawing boundaries about the use of the blood: it was for the organization to make appropriate decisions about its uses. However, they sometimes made it clear that they did not expect it to be used in other contexts.
R: When you read about it and you think do they sell it abroad and you don't, you're not giving it to do that are you? I: No. R: Do you know what I mean? Not for them to make a profit about it in some way. I know I know it probably gets ploughed back into the NHS and they need it but you don't give it for that do you? (NBS 8)
Donors' understanding of the place of the blood service in the NHS was sustaining their commitment in several ways: through it they could imagine others' need for blood, they could trust the organization to which it was given, and they expected, or at least hoped, that the blood would be used within this health system. In addition, several people mentioned the role that expert scientific and experts committees would play here. It was not that donors stated categorically that non-nationals should not receive their blood. Nor did they often feel sufficiently informed or expert to be certain of the moral boundaries around new developments that might involve the use of blood for research. However, where discussions arose about the use of blood by commercial companies -for whatever purpose -this kind of use was not seen to be within the terms of a national health service. Importantly then, the association of the blood service with the NHS played an important part in defining and delimiting the imagined uses of blood.
When asked about the uses to which the blood is put, donors were sometimes puzzled or alternatively embarrassed at their lack of detailed knowledge, as in the following example.
I: What can you tell me from what you know about the kinds of ways the blood is used once it's collected? R: I think the main way is probably, just for acute care I would think, operations for people who would need regular blood transfusions. I would think that my blood would probably last about five minutes you know when it's been cleared because just for certain operations that's really my basic understanding of it. Just for surgeons to carry out operations. I wouldn't know what else it could be used for. I'm unaware of any other uses for it. (NBS 57)
It seemed that there was a gap between donors' formal consent, as indicated on their signed declaration, and the uses for the blood which were prominent in their explanations to me -such as use in operations and other such emergencies. One woman explained the uses of blood as follows:
No I'm very lazy, I mean I know it can be used obviously for donation, for transfusions This sense of 'having faith in them', or placing implicit trust in the organization, was one of the most consistent findings from these interviews with NBS donors. I observed that the detailed and specific information that was provided in the blood centre was not accorded much attention by donors during the time they spent waiting, and it seemed to play a small part in their interview accounts. Before I began interviewing NBS donors I was aware that each time blood was given, the donor signed a written consent form. The consent form formally indicated their written agreement for their blood to be tested, and then used for suitable purposes by the NBS, with research approved by an ethics committee being one of the uses set out in the donor information. Yet, when donors spoke about how the blood was used, few spontaneously mentioned the possibility that it might be used for research.
There are some resonances with Pfeffer and Laws' study of venepuncture among hospital patients, in that for these donors, donated blood stored within the NBS blood centres, blood banks and NHS hospitals is also seen as 'matter in place' (Pfeffer and Laws, 2006: 3022). The implicit nature of this arrangement, underlined by the extent to which donors often disavowed knowledge of or responsibility for the uses of their blood makes it difficult to establish categorically what donors expect. A common response to questions about the use of blood for research was to indicate that donors could not be expected to know about or understand research, and to reaffirm principles that had already been stated:
R: So therefore if blood and research and different things then like … I said I've come to the session, I've give my blood right, I've had my cup of tea and my biscuit and like I say when I go through the door what the Blood Service do with that it's up to them. I: You leave it behind mm. R: I mean if they say to me 'Your pint of blood is going to research.' I: They won't, I mean that's not … R: But I'm saying, if they said to me 'Well you've give a pint of blood today, it's going for research', it would probably save some lives, not save one life it might save three lives all well and good providing it's used for people in this country and nobody else. I don't like this idea that my blood is going to some, any Tom, Dick and Harry outside the UK that's what I'm saying. I: So it's very important to you that it's a national blood system? R: It is a national thing and I think, and I feel very strongly. If they said to me 'Oh well we're going to start sending the blood all over the place' I wouldn't have it. It's the same as paying, if you want to do that no. (NBS 100)
According to this donor, the blood service -or the NHS -was to put the blood to the best use, which might include research.
In this analysis, donors' responses to interview questions are taken to illuminate some of the dynamics of involvement in blood donation. An important finding is that donors' trust in the NBS is informed by its history and by the reputation of the NHS.
As Misztal (1996: 156) has written:
Habit, reputation and memory … all [are] means of preserving the past experience in order to construct a more predictable, reliable and legible present. They are all different but complementary strategies designed to help us to acquire a general sense of trust in the social world.
Trust is a valuable resource for blood services. However, the gap between the nostalgic accounts of blood banking, and the contemporary organization of blood services and industries is troubling, both in moral and in practical terms. As new challenges face those supplying blood sufficient in quantity and in quality, this disjuncture is likely to become more evident. The recent investigations of the Archer Inquiry into the impact of contaminated blood on people with haemophilia and others affected are testimony both to the complexity of blood services, and to the legacy of problems associated with a lack of openness about these systems (Boseley, 2009). Blood services have had major changes to confront, in the interests of risk management and patient safety. It is fair to say that questions of narratives or cultural accounts have, understandably, not been foremost in the minds of NBS leaders and of policy makers. However, a decrease in the number of blood donors is of concern. The policy response to this has been primarily orientated towards improving 'customer satisfaction' with the convenience of donor centres and appointments, and towards advertising for new donors (NHS Blood and Transplant, 2007). Some sociological reflections on the 'implicit contract' between donors and the service may flesh out these dynamics further (Pfeffer and Laws, 2006: 3022).
The business of procuring and supplying blood has become increasingly complex and is an international one. For example, the Department of Health has for some years had a policy of sourcing plasma from the USA, following recognition of the possible risk of variant Creutzfeldt-Jakob disease (vCJD) infectivity posed by plasma from UK-based donors. A commercial division of NBS, the Bio Products Laboratory markets plasma products both to the NHS and to over 40 countries worldwide (Bio Products, 2009). Thus, donated blood can circulate widely, globally. Blood is managed and regulated as a commodity -a process that arguably contributes to safety for recipients. Yet the extent of these transactions and circulations is often elided in NBS literature for donors. At a time when there is new emphasis on openness and transparency between the wider NHS and its publics, the mismatch of nostalgic narratives and modern practices seems problematic.
Foundational stories about blood banking in Britain emphasize ideals of altruism, mutuality, national solidarity and the provision of whole blood for use in medical emergencies. Modern blood services in the UK on the other hand are characterized by the use of blood products, including products from plasma procured in an international market. The risks of receiving blood products are explicitly recognized and managed under modern formal safety regimes. These developments challenge the more literal images of reciprocity and solidarity that have historically been associated with collective narratives about blood banking.
These donors' accounts are resonant of an older paradigm of 'blood banking'. They may also reflect the very limited public discussion about dilemmas in managing, balancing and communicating risk in 'blood systems'. While we have seen extensive research and public consultation in the area of novel biological products, there has been little engagement with patients and publics in the more 'established' field of blood products. A shift to a more open institutional narrative about modern blood services may be timely, as blood services face current and future challenges.
Funding from the
Helen Busby is a sociologist who is interested in healthcare systems. She is especially interested in social aspects of the donation, 'banking' and applications of human biological materials in medicine and in research. She is currently Senior Research Fellow in the Department of Sociology and Social Policy at the University of Nottingham, UK. Her current research is concerned with public policy and ethical issues in relation to the banking of stem cells from umbilical cord blood.
1 Previously, blood services in England and Wales were organized on a regional basis by NHS authorities. The transition to a national system began in 1993. In 2005, a new NHS umbrella body, NHS Blood and Transplant Authority, was given managerial responsibility for the NBS. 2 A second group of interviews was conducted with people who donated blood as a sample for genetic analysis, to a university research project. These interviews, which are reported elsewhere, addressed donors'/participants' thinking about 'genetic research' in some depth (Busby, 2004a). 3 The ideal of self-sufficiency within national communities has been an important theme in blood policies in Britain, as in the rest of Europe, albeit one that is in tension with the reality of trade in blood products (Farrell, 2006;Hagen, 1993). See Archer et al. (2009: 26-46) for a discussion of the tensions inherent in policies aiming at national 'self-sufficiency' of blood supply in the United Kingdom. 4 See Ritchie and Spencer (1994) on the use of a data matrix as a tool in qualitative data analysis. 5 See How Blood Is Used at http://www.blood.co.uk/pages/e18used.html. 6 http://www.blood.co.uk/StockGraph/stocklevelstandard.aspx.