What is big data?
Extremely large datasets that traditional tools can’t process efficiently, often handled with tools like Hadoop and Spark.
Extremely large datasets that traditional tools can’t process efficiently, often handled with tools like Hadoop and Spark.
The process of detecting and correcting (or removing) inaccurate, corrupted, or irrelevant data from a dataset.
A plot that shows the diagnostic ability of a binary classifier as its discrimination threshold changes.
A performance table for classification models showing true positives, false positives, true negatives, and false negatives.
A dimensionality reduction technique that transforms features into a smaller set of uncorrelated variables.
Creating new input features from raw data to improve model performance
A technique to assess model performance by dividing the data into training and validation sets multiple times.
A standard process model for data mining with 6 phases: Business Understanding, Data Understanding, Data Preparation, Modeling, Evaluation, and Deployment
Overfitting: Model is too complex and memorizes training data. Underfitting: Model is too simple and fails to learn the data pattern
Matplotlib, Seaborn, Plotly (Python), and Tableau, Power BI (business intelligence tools).