The following results from linear discriminant analysis (run in R) can be found at the bottom of this file and are summarized as follows: [1] m m m m m m m m m m m m i.e. all 12 disputed papers are classified into the "madison" class Note that the four features tested are the top 2 function words ("upon" and "there") and the top 2 word length features (2-letter words and 1-letter words). These rankings were computed using contigency tables. ***************************************************************** > test <- read.table("madham.txt", sep=""); > train <- read.table("disputed.txt", sep=""); > cl <- factor(c(rep("m",14), rep("h",18))) > test V1 V2 V3 V4 1 0.000 2.003 27.036 206.943 2 0.000 0.000 18.250 212.915 3 0.369 0.738 21.033 202.583 4 1.207 0.905 29.261 211.161 5 0.000 0.000 20.353 236.943 6 0.000 1.328 18.254 216.064 7 0.000 0.282 26.494 204.622 8 0.720 1.081 18.732 211.095 9 0.000 0.583 30.877 218.177 10 0.000 1.379 18.276 203.793 11 0.000 1.418 10.875 201.891 12 0.000 0.000 17.625 200.383 13 0.000 1.097 14.991 209.506 14 0.000 1.077 23.694 210.016 15 3.788 1.263 24.621 224.116 16 1.876 3.752 26.266 189.962 17 4.885 3.996 20.870 215.364 18 1.511 1.007 23.162 212.991 19 2.026 1.520 26.849 219.352 20 2.406 3.208 28.869 212.911 21 3.265 4.198 22.388 215.019 22 3.853 3.302 27.518 238.305 23 2.523 3.364 23.129 241.800 24 2.377 6.452 20.034 232.937 25 1.580 3.686 27.383 232.227 26 4.965 2.483 22.344 231.380 27 4.955 2.252 24.324 232.432 28 2.127 4.558 24.309 216.651 29 1.764 1.764 27.043 230.453 30 2.458 4.425 28.024 226.647 31 3.015 4.020 30.151 235.176 32 5.068 1.521 36.999 231.627 > train V1 V2 V3 V4 1 0.000 1.217 21.290 226.277 2 0.908 0.000 17.257 205.268 3 0.000 2.093 25.118 225.536 4 0.000 0.000 21.727 217.273 5 0.000 0.926 23.611 208.796 6 1.003 0.501 19.549 232.581 7 0.000 2.458 29.007 210.914 8 0.000 2.454 29.448 192.638 9 0.000 1.816 18.611 226.509 10 0.000 0.960 28.325 207.393 11 0.000 0.000 34.439 220.076 12 0.000 2.639 22.757 213.391 > cl [1] m m m m m m m m m m m m m m h h h h h h h h h h h h h h h h h h Levels: h m > z <- lda(test, cl) Error: couldn't find function "lda" > library(MASS) > z <- lda(test, cl) > z Call: lda.data.frame(test, cl) Prior probabilities of groups: h m 0.5625 0.4375 Group means: V1 V2 V3 V4 h 3.024556 3.1539444 25.79350 224.4083 m 0.164000 0.8493571 21.12507 210.4351 Coefficients of linear discriminants: LD1 V1 -0.78164820 V2 -0.58883622 V3 -0.06020593 V4 -0.01141801 > predict(z, train) $class [1] m m m m m m m m m m m m Levels: h m $posterior h m 1 1.160428e-03 0.9988396 2 1.612292e-04 0.9998388 3 2.227611e-02 0.9777239 4 4.739776e-05 0.9999526 5 4.569786e-04 0.9995430 6 4.370276e-03 0.9956297 7 6.637281e-02 0.9336272 8 3.267559e-02 0.9673244 9 2.535042e-03 0.9974650 10 1.457602e-03 0.9985424 11 1.180378e-03 0.9988196 12 2.614513e-02 0.9738549 $x LD1 1 1.9897887 2 2.4793572 3 1.2519606 4 2.7829001 5 2.2210002 6 1.6602417 7 0.9698486 8 1.1543287 9 1.7957185 10 1.9331885 11 1.9855576 12 1.2112739