Description of the dataset and using linear models

For this project, we will be using the credit card fraud detection dataset from Kaggle. The dataset can be downloaded from https://www.kaggle.com/dalpozz/creditcardfraud. Since I am using the dataset, it would be a good idea to be transparent by citing the following publication:

  • Andrea Dal Pozzolo, Olivier Caelen, Reid A. Johnson, and Gianluca Bontempi, Calibrating Probability with Undersampling for Unbalanced Classification. In Symposium on Computational Intelligence and Data Mining (CIDM), IEEE, 2015.

The datasets contain transactions made by credit cards by European cardholders in September 2013 over the span of only two days. There is a total of 285,299 transactions, with only 492 frauds ...

Get Scala Machine Learning Projects now with the O’Reilly learning platform.

O’Reilly members experience books, live events, courses curated by job role, and more from O’Reilly and nearly 200 top publishers.