Machine Methods and Data Analysis in Studies on Comparative Psycholinguistics
Sign in to track this film in your collection or want list.
Genre: Educational
Creator: to be added
Format: 16mm
Sound: sound
Description: A report on Machine Methods and Data Analysis in Studies on Comparative Psycholinguistics. Shows lots of footage of data entry, punch cards, old computer tape drives
Transcription
the film you were about to see was prepared in order that the social scientists of the country's participating in this research project might have an opportunity to view some of our data analysis procedures in the course of this film we shall attempt to illustrate step by step each of the analysis phases as they are accomplished here at the University of Illinois although we began collecting data from a broad range of countries in 1960 research completed much earlier had shown that a stable structure of affective or connotative meanings existed for American English-speaking peoples the discovery of such a structure made it possible to describe with Precision a very wide range of seemingly quite different cognitive behaviors as instances of combinations of a rather small number of basic Dimensions these results encouraged the group at the University of Illinois to attempt to determine if such a framework might be found in all human beings if such a framework could be found great advances in the investigation of International Communication and attitudes would be possible an attempt to determine such a common framework may be likened to the establishment of an International System of weights and measures providing the same advantages to communication and understanding that such a common metric would provide it was with this motivation that Charles eosg good director of The Institute of communications research and principal investigator of this project Murray s meon in charge of statistical analyses and design and William K Archer ethnol linguist in charge of field contacts and linguistic problems began the current research project with support provided by the human ecology fund although we began the project in 1960 with only six participating countries before the first year was concluded we found that we could increase this sample to 16 because of the development of automatic data processing techniques of which we shall see more later on in the film we see Professor Archer indicating some of these 16 sites in a conference with the entire staff of the project in this next scene we see Mr yasumasa Tanaka explaining to the staff some of the intricacies of the Japanese frame construction used to elicit qualifiers in Japanese for each language such a frame conforming to the rules of the language is constructed these frames are designed to elicit adjective word associ iation responses each of the 100 nouns some of which you see at the right of the Blackboard is used as a stimulus to which High School subjects are asked to respond each of the persons tested is instructed to supply a single qualifier Association to each of these 100 nouns presented to them in their own language for example the demonstration subject in this scene has responded in Japanese on the first two lines of the form with the words good to house and happy to girl a total of 100 high school students are used in these testings in each of the countries Mrs viia Kumari shagam the resident field worker in myor India is shown here examining the forms for the conid language the responses of these students are collated by the field staff in a form as shown here you will note that the form indicates the noun pain at the top of the page for which each of the items in the body of the table were elicit as Associates we will follow one of these associates the conid word hatu which appears as the first item of the form through each of the analysis steps the original words obtained in the language of the subjects appear in the left column the next column contains a transliteration of the item into the Roman alphabet the next column shows a rough translation of these words hu being translated as severe the last column indicates the number of subjects who gave this particular response in the case of hu 21 subjects Associated the quality severe with the noun pain since there are 100 subjects the total number of entries at the bottom of this column including blank responses and non- adjectives should be 100 there are 100 such sheets in all corresponding to the 100 nouns used in testings after the initial corelation these forms are sent off to the United States for further analysis we are fortunate in having some of the most advanced data processing equipment here at the University of Illinois we have several installations containing almost every conceivable type of automatic equipment here we see Mr John limber in the foreground working on the data at one of these installations our first step after receiving the summary sheets prepared by the field worker is to transfer the information from those sheets to IBM Punch Cards the machines you are now seeing are called key punches and work very much like typewriters except that the record is made in terms of punches in a card rather than printed our machine operators headed by Mr William May are quite proficient at these kinds of operation they can very rapidly transfer the data to the cards you see passing through the machine here is the finished card for the word ha you the first number which appears at the left of the card indicates the code number of the noun to which the subjects responded in this case pain was quoted as noun number 06 then follows the romanized qualifier with the number of subjects giving this response indicated last just as it was on the form sent To Us by the field worker following the completion of the transferral of the data to these Punch Cards the next step is to redo the entire task a second time we always key punch the data twice in order to check for errors in the transfer process the machine used for the second operation is called a verifier and although similar to the key punch machine it does not actually punch holes in the cards but indicates when a hole already punched does not correspond to the key which was hit the second time when this happens a red light comes on and the operator cannot proceed until the error has been noted if all the information in the card is correct the machine places a small Notch at the edge of the card cards containing errors will not be notched by the machine so they later can be removed and corrected after this step is completed still a second verification step is made on the cards all cards are printed on the IBM 407 a basic accounting machine which senses the punches and converts them into type this machine obtains its instructions from a wiring board as to what and how the information contained in the cards is to be printed this wiring board which we have just seen Mr a shagon putting into the machine functions something like a telephone switchboard interconnecting the appropriate machine instructions and information paths all the information in a card is printed simultaneously the machine processes 150 cards or 18,000 symbols every minute this as we shall see is slow by mod standards although a typist would have to type at better than a rate of 3,000 words per minute to keep up with the machine there is our friend hu again as printed from the card after we are satisfied that the data has been transferred accurately we are finally prepared to start the analysis our first step is to order each of the qualifier responses according to the number of times it occurred and the number of nouns to which each was given as an associate we have called these two characteristics the frequency and diversity of qualifiers we seek to order the qualifiers from those having the greatest frequency and diversity to those with the least of these two characteristics in order to do this we calculate the information Theory measure H for each qualifier this measure gives higher scores to qualifiers which have a maximum of both the frequency and diversity characteristics in order to calculate the formula you saw indicated we must turn to a higher speed calculating machine than the machine you saw printing the information from cards this faster machine is an IBM 1401 as a demonstration of the speed of this machine we have instructed the machine to add by ones from the time we start until the time we stop the computer we will let the machine run for about 5 seconds and then we'll determine how high it has been able to count in that period of time we will start timing from now the 5 seconds are up and the Machine indicates it reached 2156 in those 5 Seconds the instru instructions for the 1401 are written in a special programming language which the machine reads from cards the H formula in this machine language looks like this the program and data cards are ordered for input and then fed into the computer the data are stored on magnetic tape during calculations the data are transferred to the calculating unit here the calculations are carried out the printer attached to this computer prints at the rate of 600 lines per minute which would correspond to a typist rate of approximately 14,000 words per minute the next step is to select from these ordered lists of qualifiers a relatively small number which can be considered most representative in order to do this we correlate each of the qualifiers with every other qualifier in terms of its distribution of occurrences we select only those qualifiers which are high and frequency and diversity and independent in distribution of occurrence we typically obtain somewhere in the neighborhood of 1,000 qualifiers from each country in order to compare every qualifier this requires the calculation of approximately 1 half million correlation coefficients accordingly we have to turn to a much faster computer than any we have shown so far the fastest computer currently available at the university is the IBM 7090 the 7090 computer utilizes several of the 1401 systems we saw earlier simply to read and write data onto magnetic tapes these tapes which you are now seeing are used to store the information used for calculation there are 10 such tape units attached to this computer each tape is capable of holding approximately 72 million symbols which means that a total of some 700 million symbols could be recorded in the entire system the tapes are read much as they would be on a tape recorder EX except that the speed of play is much faster the tape drops down from each reel into a set of columns maintained under vacuum to buffer the rapid changes of Direction the calculations on these stored data are performed at incredible speed in one minute of operation this computer can perform 2 million multiplications of numbers eight digits long after a few minutes calculation the 7090 records on a magnetic tape a list of qualifiers selected from the total list the tape is printed on one of the auxiliary 1401 systems the printed list indicates why qualifiers were selected and discarded it should be emphasized that all of these procedures have been entirely automatic the words in all cases have been in the native languages of the subjects this final list of qualifiers is next shipped back to the country from which they originated in the case of the Indian data hu was selected and is one of the items returned to India it is at this stage that the fuel worker obtains opposites for each qualifier in order to construct qualifier scales this generalized FL flowchart summarizes the analysis steps we have seen so far when the data is received from the field it is key punched and verified a printout for final verification is made a program for computation is written and stored on Magnetic Tape along with the qualifier data through the 1401 calculations are made on the 7090 the results are transferred back to tape and printed interpretation and recording of the results are made then they are sent back to the country from which they originated along with instructions for the next phase in the next phase the same 100 nouns which were used to obtain the qualifier responses are now used as concept terms to be rated on the qualifiers and their opposites as selected in the previous stage each qualifier and its opposite form a scale with seven step divisions to indicate the degree to which the concept relates to the Quality expressed by that scale these pictures show examples of the form of this task in English and Lebanese Arabic the sample Lebanese Arabic form as shown here was used to record the average ratings of the Arabic subjects for the concept at the top of the page these forms were then shipped back to the United States for further analysis upon receipt the booklets are taken to the computer facilities the data is transferred to cards the cards are verified the printing check is made the data is ordered and then input through a 1401 to the 7090 computer where a program computes the correlations between each of the scales and then performs a factor analysis the factor Dimensions being calculated by the computer in these scenes represent the basic qualification Dimensions which were implicit in the subject's judgments of the 100 Concepts as they were actually rated on the scales these implicit Dimensions may be discovered by determining the similarities and dissimilarities between the ways in which the subjects actually Ed the rating scales the computer is determining the minimum set of Dimensions which are necessary to specify how these subjects were responding much as we know that we can exactly specify all cardboard boxes with knowledge of only the three dimensions of height width and depth so with the results of such calculations we hope to to specify the unique dimensions of attitudes which in turn specify the differences in the ways in which people view concept meanings our findings to date have indicated that the attitudes held by diverse peoples from all parts of the world speaking many different languages can be represented by a single and uniform set of Dimensions the basic Dimensions we have found are three in number the first of these common Dimensions we have called evaluation this dimension compris his judgments as to the worth of Concepts it is characterized by qualities or terms such as nice awful sweet sour Heavenly hellish and good bad the second dimension is one of the potency of Concepts it is represented by such terms as strong weak powerful powerless and big little the third dimension has been called activity and represents such terms as fast slow and noisy quiet if we place the dimensions found in each country alongside those found in all other countries as in this result of a pancultural factor analysis it is clear that these basic and characteristic dimensions are present in all countries the particular qualifications representing each change from language to language but the meaningful sense of these terms Remains the Same in all languages when viewed as instances of the three basic dimensions for example we find that our word ha you falls on the same Dimension characterized in English by such items as bad and awful differences in meanings of concepts of course also exist these differences are important and are what make the world a good strong and active place in which to live what this research we hope has added is that we may study such exciting differences with a common metric a meter which can be used internationally
Online Copy: https://www.youtube.com/watch?v=nK_ytMUM92w
Metadata Source:YouTube
1 user has this film:
AV Geeks Archive
Related films:
- Turn The Other Cheek (1958) · Family Films.
- What The Frost Does - Background For Reading And Expression (1953) · Coronet Instructional Films
- Use and Care of Books (1979) · Centron Corporation
- We Explore Ocean Life. (1962) · Coronet Instructional Films
- A New leash on life (1988) · Humane Society of the United States. Walter J. Klein Company.
- Little Smokey (1965) · to be added
- Final Factor (1982) · to be added
- World of the Weed (1968) (1968) · National Educational Television
Original permalink · Record added: 2023-02-27 16:31:15