TheChetan icon

ML Basics

TheChetan | PRO | 12/31/20 05:03:34 PM UTC (Edited) | 0 ⭐ | 1067 👁️ | Never ⏰ | []
text |

4.84 KB

|

None

|

0 👍

/

0 👎

Features:
---------
	Basically all ml algorithms need to run on a specific set of data.
	The data has to be in the most convinent form for the machine to read.
    Its our job to supply this kind of data.
	These are commonly called as features. Supply values that you think the algo might find useful in classification.
	For example: 
		[1] If you were dealing in audio, then amplitudes of selected frequencies can act as features.
		[2] For a movie dataset, the genere, length, bugget, actors, language etc are all features.
 	Lets do for some sentiment analysis of audio:
	When Sad, you (probably) speak in a lower tone
	When Angry, you speak in a higher tone. Loudness is higher etc
	When Happy, you speak in a higher tone, loudness is lower
 	For such a dataset you need these "features": [1] frequency, [2] rate of change of frequency, [3] Loudness etc
	This above is some dummy data, but bear with it
 	-------------------------------------------------------------------------------------
	| Slno. | freq1 | freq2 | fre3 | ... | freq n |  d_f1 | d_f2 | ...| df_n | loudness |
	-------------------------------------------------------------------------------------
	| tone1 | 3     |   2   |    ......  |   0    |   4    |  ...............|  9       |
	-------------------------------------------------------------------------------------
	| tone2 | 4     |   7   |    ......  |   1    |   8    |  ...............|  4       |
	-------------------------------------------------------------------------------------
	....
	-------------------------------------------------------------------------------------
	| tonem | 4     |   5   |    ......  |   1    |   8    |  ...............|  9       |
	-------------------------------------------------------------------------------------
 	I hope you get the idea.
  Label:
------
	Well if you need to classify a set of things, each possible output is called a label.
	Generally these are represented as vectors. If you had 3 outputs, most commonly you would choose your outputs as:
 	Output 1 (Sad)	: [1, 0, 0]
	Output 2 (angry): [0, 1, 0]
	Output 3 (happy): [0, 0, 1]
 	So features are the inputs and labels are you output.
	The end result of whatever algorithm you use will produce an output like:
	[0.8, 0.1, 0.1] -> So this gets classified as sad
	[0.4, 0.6, 0.0] -> This is angry 
 	This values are analogous to output probability values, so 0.4 -> 40% chance for sad
 	See: https://stackoverflow.com/questions/40898019/what-is-the-difference-between-a-feature-and-a-label
  Classifier:
-----------
	This is the actual algorithm that converts you input vector space to output vector space.
	Nice big balckbox that does the conversion.
	Actually the math behind each of these classifiers is pretty intense and fun.
	Forest is one such classifier, I dont remember much else. The wikipedia link is extensive.
 	raw inputs                         Feed to
	(audio)     ->    Features   ->    classifier    ->   Output vectors   -> Human readable output
      |      ...    .
    |    ....    .
    |   ...     o
    |  .. .     oo
    |...        oo
    | .          oo
    |.           oo
    ------------------
     Assume your features (unrelated to the sentiment analysis problem) plotted on a graph look like the above
    the "dots" are for one type of data and the "o" are for the other type.
    Depending on the algorithm that's used in the classifier, you a rule (polynomial equation) will be created
    that will determine what each future value is classified as.
     |      ...|   .
    |    .... | .
    |   ...   | o
    |  .. .   | oo
    |...      | oo
    | .       |  oo
    |.        |  oo
    ------------------
     Assume that your classifier comes with this line as the rule to decide if
    any future value is type '.' or type 'o'. Everything to the left of the line
    is '.' and to the right is 'o'.
     The end result of the training is this rule (or more commonly called model).
    You then use your test dataset to see how accurate your predictions were.
   Confusion matrix:
-----------------
    So the first step is to train your classifier and the next step is to measure how accurate the predictions are.
    A confusion matrix is a convinent way to plot the results. It represents the predictied values vs the actual values.
    For example:
     ------------------
    |    |  .  |  o  |
    ------------------
    |  . |  90 |  10 |
    ------------------
    |  o |  05 |  45 |
    ------------------
     Your Y axis represents the actual values and your X axis represents the predicted values.
    There are 100 (.) used in the test, of which your classifier predicted 90 of them correctly and 10 incorrectly.
    There are 50 (o) used in the test, of which your classifier predicted 45 of them correctly and 5 incorrectly.
 PS: Forest I dont remember

Comments