Test Set 1: 

61 docs 500+ words each drawn from Compass archive. Authors include:

Aaron Kopitz			4
Bart Stupak			2
Doug Anger			15
Erika Dykstra			5
John Petkus			8
Justin Schnurer			4
Kenneth Casperson		4
Kevin Kainula			2
Liz Sippl			3
LSSU International Education 	2
Mackenzie Barrett		2
Mary Gilray			2
Nancy Marsh			2
PR				2
SI				2
Stephanie Rice			2

Feature set generated with:
../featex13.pl Compass.testing Compass.training


Test Set 2: 

Same as set 1, with an additional 69 docs 300-500 words added to the test set. Model is unchanged. Distribution of docs in added set is:

Aaron Kopitz			9
Doug Anger			8
Erika Dykstra			3
John Petkus			7
Justin Schnurer			9
Kenneth Casperson		6
Liz Sippl			7
LSSU International Education	1
Mackenzie Barrett		1
Mary Gilray			1
PR				7
SI				1
Stephanie Rice			9

Feature set generated with:
../featex13.pl Compass.testing Compass.training Compass.300+


Test Set 3:

All docs with word count >= 300. Both long (500+) and short (300-500) word docs used in training set.

Feature set generated with
../featex13.pl Compass.testing Compass.training


Test Set 4:

New Training/test split on the docs w/ word count >=500


Test Set 5:

New Training/test split on the docs w/ word count >=500


Test Set 6:

New Training/test split on the docs w/ word count >=500


Test Set 7:

Compass.500+ as training set, Compass.300+ as test set


Test Set 8:

Compass.300+ as training set, Compass.500+ as test set


Test Set 9:

Compass.300+


Test Set 10:

Compass.300+ for training/testing, Compass.500+ added to testing only


Test Set 11:

Compass.500+


Test Set 12 - 15:

Compass.300+


Test Set 16-19

Compass.300+ & Compass.500+


Test Set 20-22:

Compass.500+ Training
Compass.300+ Testing


Test Set 23-26

Compass.300+ Training
Compass.500+ Testing


Test Set 27

Dup of TS16 using BMRclassify to get result vectors


Test Set 28

Data set from Shenzhi Li. 24 authors. ~9000+ docs


Test Set 29

Data set from Shenzhi Li broken down by author. 75/25 training/test taken on all authors. 12 authors selected for testing. One model constructed using training data from all authors, another using only 12 authors not selected for testing