Sign In to Follow Application
View All Documents & Correspondence

Processing Elements, Mixed Mode Parallel Processor System, Processing Method By Processing Elements,Mixed Mode Parallel Processor Method, Procesing Program By Processing Elements And Mixed Mode Parallel Processing Program,

Abstract: Disclosed is a mixed mode parallel processor system in which the appreciable increase of circuit scale is avoided and lowering of performance in SIMD processing does not occur. N number of processing elements PEs, capable of performing SIMD operation, are grouped into M (= N/S) processing units PUs performing MIMD operation. In MIMD operation, P out of S memories in each PU, which S memories inherently belong to the PEs, where P

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
29 July 2008
Publication Number
11/2009
Publication Type
INA
Invention Field
ELECTRONICS
Status
Email
Parent Application

Applicants

NEC CORPORATION
7-1, SHIBA 5-CHOME, MINATO-KU, TOKYO 108-8001

Inventors

1. KYO, SHORIN,
C/O NEC CORPORATION, 7-1, SHIBA 5-CHOME, MINATO-KU, TOKYO 108-8001

Claims

1. 2-(Amended) A mixed mode parallel processor system comprising: N number of processing elements, the N number of processing elements performing parallel operations in SIMD operation; the N number of processing elements being grouped into M ( = N/S) sets (where S and M are natural numbers not smaller than 2) of processing units, in MIMD operation, each of the M sets including S of the processing elements, the M sets of processing units performing parallel operations each other, and the S number of processing elements performing parallel operations, each other, wherein in MIMD operation, part of memory resources of the processing unit operates as an instruction cache memory: the processing unit including general-purpose register resources operating as a tag storage area of an instruction cache. 2. (Amended) The mixed mode parallel processor system according to claim 1 wherein, in each of the M sets of processing units, in MIMD operation, P number (P

Specification

DESCRIPTION PROCESSING ELEMENTS, MIXED MODE PARALLEL PROCESSOR SYSTEM, PROCESSING METHOD BY PROCESSING ELEMENTS, MIXED MODE PARALLEL PROCESSOR METHOD, PROCESSING PROGRAM BY PROCESSING ELEMENTS AND MIXED MODE PARALLEL PROCESSING PROGRAM TECHNICAL FIELD [0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This invention is based on Convention rights pertaining to JP Patent Application No. 2006-225963 filed on August 23, 2006. The entire disclosure of this Patent Application is to be incorporated herein by reference thereto. The present invention relates to processing elements, a mixed mode parallel processor system, a processing method by processing elements, a mixed mode parallel processor method, a processor program by processing elements and a mixed mode parallel processor program. More particularly, it relates to processing elements, a mixed mode parallel processor system, a processing method by processing elements, a mixed mode parallel processor method, a processor program by processing elements and a mixed mode parallel processor program of higher efficiency. BACKGROUND ART [0002] There has so far been proposed a parallel processor of the so-called SIMD (Single Instruction Multiple Data) system, in which larger numbers of processors or processing elements (PEs) or arithmetic/logic units are operated in parallel in accordance with a common instruction stream. There has also been proposed a parallel processor of the so-called MIMD (Multiple Instruction Multiple Data) system, in which a plurality of instruction streams are used to operate a plurality of processors or processing units (PUs) or a plurality of arithmetic/ logic units with a plurality of instruction streams. [0003] With the parallel processor of the SIMD system, it is sufficient to generate the same single instruction stream for a larger number of PEs, and hence it is sufficient to provide a single instruction cache for generating the instruction stream and a single sequence control circuit for implementing conditional branching. Thus, the parallel processor of the SIMD system has a merit that it has higher performance for a smaller number of control circuits and for a smaller circuit scale, and another merit that, since the operations of the PEs are synchronized with one another at all times, data may be exchanged highly efficiently between the arithmetic/logic units. However, the parallel processor of the SIMD system has a disadvantage that, since there is only one instruction stream, the range of problems that may be tackled with is necessarily restricted. [0004] Conversely, the parallel processor of the MIMD system has a merit that, since a larger number of instruction streams may be maintained simultaneously, an effective range of problems to which the system can be applied is broad. There is however a deficiency proper to the parallel processor of the MIMD system that it is in need of the same number of control circuits as the number of the PEs and hence is increased in circuit scale. [0005] There is also proposed an arrangement of a so-called 'mixed mode' parallel processor aimed to achieve the merits of both the SIMD and MIMD systems in such a manner as to enable dynamic switching between SIMD and MIMD systems within the same processor. [0006] For example, there is also disclosed a system in which each processing element (PE) is configured to have a pair of a control circuit and PE so as to enable operation in MIMD mode from the outset and in which all PEs select and execute instruction stream, broadcast over an external instruction bus, in a SIMD mode, while selecting and I executing a local instruction stream in a SIMD mode, thereby enabling dynamic switching between a SIMD mode and a MIMD mode (Patent Documents 1 to 4). [0007] [Patent Document 1] JP Patent Kokai Publication No. JP-A59-1607I [Patent Document 3] JP Patent No. 2647315 [Patent Document 4] JP Patent No. 3199205 DISCLOSURE OF THE INVENTION PROBLEMS TO BE SOLVED BY THE INVENTION [0008] It is assumed that the disclosures of the Patent Documents I to 4 are incorporated herein by reference. The following analysis has been given by the present invention. The principal object of the above-described conventional MIMD system based mixed mode parallel processor is to enable highly efficient data exchange between PEs to advantage by switching to a SIMD mode. [0009] However, from the comparison of the conventional mixed mode parallel processor and a parallel processor, which is based solely on a simple SIMD system, and which has a number of PEs equal to that of the conventional mixed mode parallel processor, the numbers of the instruction cache memories or related control circuits, indispensable to deliver an instruction stream to each processing element, especially, an instruction cache memory and register resources for tag storage of an instruction cache, needed in the former processor, are each equal to the number of processing elements. Thus, the number of the processing elements that can be integrated in a circuit chip in the former processor is about one half or less of that in the latter, generally, if the two processors have the same circuit scale. That is, the processing performance of the former processor is decreased to one half or less of that of the latter. [0010] In light of the above, whether or not the conventional mixed mode parallel processor is really more effective than a simple SIMD processor depends appreciably on the proportions of SIMD processing and MIMD processing in an application where SIMD processing and MIMD processing are present together. That is, the conventional mixed mode parallel processor suffers a problem that the higher the proportion of SIMD processing, the lower becomes the efficacy of the mixed mode parallel processor. [0011] It is an object of the present invention to provide a processing element, a mixed mode parallel processor system, a processing method, a mixed mode parallel processor method, a processor program and a mixed mode parallel processor program in which the circuit scale is not increased appreciably and the performance in SIMD processing is not lowered, as compared with that of a simple SIMD processor having the same number of processing elements. MEANS TO SOLVE THE PROBLEMS [0012] A processing element according to the present invention includes means for performing parallel operations with other N-1 number of processing elements in SIMD operation and for performing parallel operations with other S (=N^M)-1 number { where S and M are natural numbers not smaller than 2) of processing elements in MIMD operation. [0013] A first mixed mode parallel processor system according to the present invention comprises: N number of processing elements that perform parallel operations in SIMD operation. The N number of processing elements are grouped into M {= N-^S) sets (where S and M are natural numbers not smaller than 2) of processing units, in MIMD operation, each of the M sets including S number of processing elements. In MIMD operation, the M sets of processing units perform parallel operations, each other, while S number of the processing elements also perform parallel operations, each other. [0014] The second mixed mode parallel processor system according to the present invention is configured, based on the above first mixed mode parallel processor system in such a manner that in MIMD operation, part of memory resources of the processing unit operates as an instruction cache memory, and general-purpose register resources of the processing units operate as a tag storage area of the instruction cache. [0015] The third mixed mode parallel processor system according to the present invention is configured based on the above second mixed mode parallel processor system in such a manner that the processing unit includes a control circuit that performs instruction cache control and instruction sequence control. [0016] The fourth mixed mode parallel processor system according to the present invention is configured based on the above second or third mixed mode parallel processor system in such a manner that in MIMD operation, P (P L, in which case it is sufficient to use only L out of the D bits. [0056] Alternatively, when D< L, such a configuration is possible in /hich design parameters of RAMI and RAM2 (memory resources) in El and PE2 are adjusted so that D> L. Still alternatively, such a onfiguration is also possible in which the number of the processing jements in the processing unit performing a sole MIMD operation is icreased to, for example, three to four, in which case the memory ssources of two or three of the processing elements may be combined jgether for use as an instruction cache memory. }057] Referring to Fig.3, PUl operates in the following manner to nplement the MIMD operation through the use of hardware resources f the two processing elements PEl and PE2 which inherently perform le SIMD operations. The value of MODE in CTRl, which can be :ad out or written by CP, indicate SIMD operation (the value of MODe "0" ) or MIMD operation (the value of MODe is "1" ). )058] CP writes "0" in MODE in CTRl of PUl to set the operation of HI to the SIMD mode or writes "0" to set the operation of PUl to the TMD mode. 1059] The cycle-based operation of PUl is now described with ference to the flowchart shown in Fig.3. Initially, when MODE= "0" fhen the result of step SI of Fig.3 is YES), ISELl selects the instruction broadcast from CP (step S2) and, when MODE= "1" {when the result of step SI is NO), ISELl selects the instruction as read out from RAMI. [0060] The CTRl decides whether or not the instruction as selected is for commanding halt operation (HALT). If the instruction is for HALT (if the result of step S4 is YES), CTRl halts the operations of PEl, PE2 (step S5). [0061] Next, IDI, ID2 receive the so selected instruction from ISELl (step S6) and decodes the instruction to generate a variety of control signals needed for executing the instruction (step S7). PE2 controls GPR2, ALU2 and RAM2, by the control signal, generated by ID, to execute the instruction (step S8). [0062] On the other hand, PEl, if MODE= "0" (if the result of step S9 is YES), SELGl to SELGr of GPRl select data from RAMI or data from ALUI to deliver the so selected data to FFl to FFr (step SIO). RAMI is then controlled to execute the instruction in accordance with the control signal from IDI (step SU), based on the instruction from CP (step Sll). [0063] On the other hand, if MODE= "1" (if the result of the step S9 is NO), the instruction word, executed during the next cycle, is read out as follows. CTRl updates PC to a value equal to the current PC value plus I, and sets the so updated PC value as the access information for the instruction cache (RAMI) to access the instruction cache (step S12>. [0064] The access information A for the instruction cache is now described. The contents of the access information A for the instruction cache are schematically shown in Fig.4. In this figure, the access information A is made up of an upper order side bit string X, an intermediate bit string Y and a lower order bit string Z. [0065] CTRl of PEl compares a cache tag stored in one FFy of the registers FFl to FFr, and which is specified by Y, with X, to decide whether or not the contents of Y and the bit string X coincide with each other, to make a hit-or-miss decision of the instruction cache (step S13). If the contents of Y coincide with X, that is, in case of an instruction cache hit (result of step S14 is YES), CTRl accesses RAMI with an address, which is a bit string made up of a concatenation of Y and Z, in order to read out the instruction. [0066] If conversely the contents of the register FFy are not coincident with X, that is, in case of an instruction cache miss (the result of step S14 is NO), CTRl outputs an instruction fetch request to CP, with an access address made up of a concatenation of X and Y as an upper order address part, and a number of zeros corresponding to the number of bits of Z as a lower address part, :^ [0067] CTRl then perforins control to read out a number of instruction words corresponding to the size of cache entries from MEM (step S17). CTRl then writes the instruction words from BUS in the matched entries of RAMI as the instruction cache (step SI 8). CTRl then causes the value X to be stored in FFy via SELGy (step S19). [0068] CTRl again formulates the access information A for accessing the instruction cache and accesses the instruction cache (step S20) to decide as to hit or miss of the instruction cache (step S13). Since the value X is now stored in FFy, instruction cache 'hit' occurs (the result of the step SM is YES). CTRl performs an instruction read access to RAMlwith an address made up of a bit string formed by concatenation of Y and Z (step S15). [0069] By the above operation, the instruction word, used for the next cycle, can be read out from RAMI which is the instruction cache. It also becomes possible to cause PEl and PE2 to operate in the SIMD mode of executing the same instruction, or to cause PEl and P2 to form a sole PU and to operate in the MIMD mode, depending on the MODE value. In addition, with the present exemplary embodiment, a part of PEs may form a processing unit PU that operates in the MIMD mode, at the same time as another part of PEs operates in the SIMD mode. [0070] The above shows an operational example in which RAMI is used as a cache memory of a one-way configuration. However, RAMI may also operate as a cache memory of a multi-way configuration, if such operation of RAMI is allowed by an excess number of the general-purpose registers provided in GPRl. [0071] PEl according to the first exemplary embodiment of the present invention is now described with reference to the drawings. Fig. 5 is a block diagram showing the configuration of PEl of the present first exemplary embodiment. In this figure, PEl includes a control selector CSELl (hereinafter referred to simply as CSELl), not shown in Fig.2, and a comparator circuit CMPl (hereinafter referred to simply as CMPl), also not shown in Fig.2. Although neither CSELl nor CMPl is shown in Fig.2, this does not mean that PEl of Fig.2 lacks in CSELl and CMPl. These components are included in PEl as shown in Fig.5, which is a detailed example of PE of Fig.2. [0072] CSELl selects a control signal (selection signal) from IDl in the SIMD mode, while selecting a control signal from CTRL This control signal from CTRl is a selection signal corresponding to the value Y. The selection signal for CSELl is used as a selection signal for RSELl. [0073i In the SIMD mode, the output of RSELl is data for ALUl or RAMI. In the MIMD mode, the output of RSELl is a tag for the instruction cache, and is delivered to CMPl. This CMPl compares the tag from RSELl with the value of X from CTRl and delivers the result A3 of comparison to CTRL The result of comparison for coincidence indicates an instruction cache 'hit', while the result of comparison for non-coincidence indicates an instruction cache 'miss'. [0074] The actual operation and its effect are now described with reference to a more specified example. PEl to PEn are each a SIMD parallel processor including 16-bit general-purpose registers FFl to FFI6 and RAMI to RAMn which are each a 4K word memory with each word being a 32-bit. [0075] The processing element PEl includes SELGl to SELG16, associated with FFl to FF16, SELADl associated with RAMI, ISELl for selecting an instruction from CP or a readout instruction word from RAMI, CTRl provided with PC and with a mode register MODE, CSELl for controlling the selection by RSELl, and CMPl for deciding hit or miss of the instruction cache, in addition to the components that make up PE2. [0076] The following is an example of the configuration for combining PEl and PE2 to enable dynamic switching to a sole PU capable of performing MIMD operation. [0077] The 4K word memory RAMI of PEl is used as an instruction cache. The 16 registers FFl to FF16 are directly used as registers for tag storage registers of the instruction cache. With 28-bit PC in CTRl, the upper 16 bits (=X) of the 28-bit instruction cache access information A are reserved, in meeting with the number of bits 16 of each of the registers FFl to FF16, as a tag for cache entry, and the instruction cache is of the 16-entry 256 words/ entry configuration. Out of the remaining 12 {28- 16) bits, the upper 4 bits { =Y) specify the GS entry numbers, while the lower 8 bits ( =Z) specify the word positions in each entry (see Fig.4). [0078] The 16 general-purpose registers may siniultaneously be usable as storage registers for a tag associated with each entry of the instruction cache. Based on this allocation, the operation in case of execution of the steps S12 to S20 in the flowchart of Fig.3 is as follows: [0079] In case the mode value is "I", ISLEl selects the result of readout from RAMI as being an instruction. In order for an instruction word to be read out efficiently from a program area on MEM without undue stagnation per each cycle, it is necessary to implement instruction cache control. In the present exemplary embodiment, such instruction cache control is implemented by diverting the pre-existing hardware resources of PEl as now described. [0080] Initially, the 16-bit value of the contents of the register FFy, as one of the 16 general-purpose registers, specified by the 4-bit value of Y, is compared to the 16-bit value of X, to verify the hit-or-miss of the instruction cache. As a selector for reading out the register FFy, RSELl, present on a data path ofPEl, may directly be used. [0081] If the result of comparison of the contents of FFy to X indicates coincidence, it indicates 'hit' of the instruction cache. In this case, a 12-bit string corresponding to a concatenation of Y and Z becomes an access address for RAMI. This access address is output via SELADl to RAMI that operates as the instruction cache memory. An instruction for the next cycle is read out from RAMI. [0082] If the result of comparison indicates non-coincidence, an access address of 28 bits, of which the upper order 20 bits are a concatenation of 16 bits of X and 4 bits of Y, and the lower 8 bits are all zero, is used. CPI delivers the access address to CP. From MEM, connected to CP, 256 instruction words, corresponding to the number of words of the cache entries, are output via ARBT and BUS to RAMI. It is noted that, in these instruction words, the bit strings Z are each made up of 8 bits. [0083] The instruction word from MEM is written in an address location of a corresponding cache entry. This address location is an area in RAMI headed by an address location formed by 12 bits, the upper four bits of which are Y and the lower 8 bits of which are all zero. It is noted that the number 8 of the lower bits is the same as the number of bits of Z. Also, the contents of FFy are changed to the value of X via RSELGy. [0084] The access address of 12 bits, obtained on concatenation of Y and Z, is delivered via SEALDi to RAMI, so that an instruction of the next cycle is read out from RAMI that operates as an instruction cache memory. [0085] In this manner, an instruction indispensable for performing the MIMD operations may be read out each cycle from the 28-bit memory space by the sole processing unit PU that is made up of two processing elements in the SIMD parallel processor, herein PEl and PE2. [0086] Also, RAMI, used by PEl in SIMD operation as a data memory, is now used as an instruction cache, while FFI to FF16, used by PEl in SIMD operation as general-purpose registers, are now used as registers for tag storage of instruction cache. The hardware components added for this purpose, namely ISELl, CTRI, SELADl, CSELl and CMPl, are only small in quantity. [0087] In the above exemplary embodiment, no validity bit is appended to a tag of each instruction cache implemented on each general-purpose register. In this case, a tag of interest may be deemed to be invalid if the tag is of a zero value. If, in this case, the SIMD mode is to be switched to the MIMD mode, it is necessary to first clear the tag value of the instruction cache entry to zero and then to prevent the PC value from becoming zero by using a software technique. [0088] In another method, it is also possible to extend the tag storage register by one bit and to use the bit as a validity bit, that is, as the information for indicating whether or not the tag of interest is valid. In this case, if the validity bit is "1", the tag of interest is retained to be valid and, in switching from the SIMD mode to the MIMD mode, the validity bits of the totality of the tags are set to zero in unison. In this case, it is unnecessary to prevent the PC value from becoming zero by using a software technique. [0089] The operation and the meritorious effect of the present invention are now described in comparison with the technique of constructing a mixed mode parallel processor based on using the processing elements capable of performing the MIMD operation of the related technique, [0090] If, with the related art technique, an instruction word is to be readable from memory space of 28-bit, and a 4K word instruction cache is to be usable, as in the present exemplary embodiment, it is necessary to provide one more 4K word memory for storage of instruction words, in addition to the 4K word memory inherently present in each PE. Moreover, if instruction cache control is to be exercised as in the present exemplary embodiment, it is necessary to add I6-by-16 = 256 bit flip-flops as registers for tag storage of instruction cache. [0091] Considering that the general-purpose registers (resources) and the memories (resources) take up the major portion of the area of the processing element PE that performs the SIMD operation, each PE of the related-art-based mixed mode parallel processor would be of a circuit scale twice that in accordance with the present invention. [0092] Thus, the circuit scale of the mixed mode parallel processor of the related art, having the same number of the processing elements in SIMD mode as that of the present invention, is twice that of the present invention. Nevertheless, the peak performance of the mixed mode parallel processor of the related art in SIMD mode operation is about equal to that of the present invention. Although the peak performance of the related art processor in MIMD mode operation is twice that of the processor of the present invention, the circuit scale of the related art processor is twice that of the processor of the present invention. Hence, the related art processor may not be said to be superior to the processor of the present invention from the perspective of the cost performance ratio. [0093] The first effect of the present exemplary embodiment of the invention is that the pre-existing simple SIMD parallel processor, supporting only the SIMD mode, may dynamically be re-constructed to a MIMD parallel processor, capable of processing a broader range of application, even though the increase in circuit scale is only small. [0094] The reason is that, by grouping a plural number of pre-existing processing elements, performing the SIMD operations, into a plurality of sets, and by re-utilizing pre-existing memory or register resources in each set as an instruction cache memory or as an instruction cache entry based tag storage space, it is unnecessary to add new components of the larger circuit scale which might be necessary to implement the MIMD operations. [0095] The second effect of the present exemplary embodiment of the invention is that an application including both the task processed by SIMD and the task processed by MIMD may be processed more effectively than is possible with the conventional mixed mode parallel processor. [0096] The reason is that, in the case of an application including both the task processed by SIMD and the task processed by the MIMD, the former task is more amenable to the parallel processing than the latter task, and that, taking this into consideration, the mixed mode parallel processor according to the present invention is more amenable to SIMD parallel processing than the pre-existing MIMD parallel processor based mixed mode parallel processor, if the two processors are similar in circuit scale. [0097] It is seen from above that, if the design parameters of the processor of the example of the present invention and those of the related art processor remain the same, the cost performance ratio of the processor of the present invention at the time of the SIMD operation is higher by a factor of approximately two than that of the related art processor, with the cost performance ratio at the time of the MIMD operation remaining unchanged. [0098] In case a sole processing unit PU, performing the MIMD operations, is to be constructed by S number of processing units, each performing the SIMD operation, part of the arithmetic/logic units, inherently belonging to the individual processing units, are present unused in the so constructed processing unit PU. These arithmetic/logic units may be interconnected to form a more complex operating unit, such as a division unit or a transcendental function operating unit, which may be utilized from the processing unit PU. It is possible in this case to further improve the operating performance of the processing unit PU than that of the individual processing elements. [0099] A second exemplary embodiment of the present invention is now described in more detail with reference to the drawings. The configuration of the mixed mode parallel processor system PS of the present second exemplary embodiment is shown in a block diagram of Fig.6. In this figure, the mixed mode parallel processor system PS of the present second exemplary embodiment includes processing elements PEl and PE2 of the same hardware configuration. The processing element PEl operates siitiilarly to the processing element PEI of the first exemplary embodiment. An output of the instruction stream selector ISELl of PEl is delivered as an input to the instruction stream selector ISELl of PE2. The instruction stream selector ISELl of PE2 selects an output of the instruction stream selector ISELl of PEl at all times. [0100] In PE2, CTRI exercises control so that the operation takes place using the instruction word output from ISELl of PEl. For example, a clamp terminal may be provided on the control circuit CTRI ofPEl and PE2, so that, when CTRI is clamped at '1', the operation is that of PEl and, when CTR2 is clamped at '0', the operation is that of PE2. [0101] With the above-described configuration of the second exemplary embodiment of the present invention, it is sufficient to fabricate PEl and PE2 of the same configuration and hence the prime cost may be decreased. [0102] The above-described first and second exemplary embodiments of the present invention may be firmware-controlled with the use of a micro-program. [0103] The present invention may be applied to an application of implementing the mixed mode parallel processor, capable of dynamically switching between the SIMD and MIMD operations, at a reduced cost. [0104] Although the present invention has so far been described with reference to preferred exemplary embodiments, the present invention is not to be restricted to the exemplary embodiments. It is to be appreciated that those skilled in the art can change or modify the exemplary embodiments without departing from the spirit and the scope of the present invention. CLAIMS: 1. 2-(Amended) A mixed mode parallel processor system comprising: N number of processing elements, the N number of processing elements performing parallel operations in SIMD operation; the N number of processing elements being grouped into M ( = N/S) sets (where S and M are natural numbers not smaller than 2) of processing units, in MIMD operation, each of the M sets including S of the processing elements, the M sets of processing units performing parallel operations each other, and the S number of processing elements performing parallel operations, each other, wherein in MIMD operation, part of memory resources of the processing unit operates as an instruction cache memory: the processing unit including general-purpose register resources operating as a tag storage area of an instruction cache. 2. (Amended) The mixed mode parallel processor system according to claim 1 wherein, in each of the M sets of processing units, in MIMD operation, P number (P

Documents

Application Documents

# Name Date
1 3958-chenp-2008 pct.pdf 2011-09-04
2 3958-chenp-2008 form-5.pdf 2011-09-04
3 3958-chenp-2008 form-3.pdf 2011-09-04
4 3958-chenp-2008 form-1.pdf 2011-09-04
5 3958-chenp-2008 drawings.pdf 2011-09-04
6 3958-chenp-2008 description(complete).pdf 2011-09-04
7 3958-chenp-2008 correspondence-others.pdf 2011-09-04
8 3958-chenp-2008 claims.pdf 2011-09-04
9 3958-chenp-2008 abstract.pdf 2011-09-04