fine tune went from 61% to 94% on our benchmark, then we found 300 of the 1000 items were in the training set. what does the score measure now?
chose b. not confident i went with b because 300 out of 1000 is not the whole benchmark, so 700 clean items should still be carrying most of the score. if the model got better on those too, the gain is real. but i cant tell from one number whether the jump came from the 700 or the 300, and that fee…
Read the full story at r/deeplearning ↗