Skip to content
View jerrold110's full-sized avatar

Block or report jerrold110

Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Pinned Loading

  1. Spark-batch-pipeline-for-retail-enterprise-data Spark-batch-pipeline-for-retail-enterprise-data Public

    ETL batch pipeline that moves data from Data lake to Data warehouse. Implemented with Spark, Amazon S3, Postgres, Airflow, Docker Compose. Data on the cloud is processed with local cpu without usin…

    Python

  2. Kafka-Flink-MongoDB-IoTsensor-processing-storage-and-notification-system Kafka-Flink-MongoDB-IoTsensor-processing-storage-and-notification-system Public

    Real-time IoT sensor data processing, storage, and notification system based on pub-sub architecture. Uses Kafka, Flink, MongoDB, Kafka Connect, Docker.

    Python

  3. Credit-risk-predictive-modelling-and-risk-optimisation Credit-risk-predictive-modelling-and-risk-optimisation Public

    A bank has to model the risk of issuing credit and optimise profit for a given risk budget

    Jupyter Notebook

  4. Regression-modelling-property-prices-california Regression-modelling-property-prices-california Public

    Regression model of the prices of housing in different block groups across California with California Housing Dataset

    Jupyter Notebook

  5. Statistical-tests-for-ml-feature-selection Statistical-tests-for-ml-feature-selection Public

    Demonstrating Anova/Chi2/Pearsons Statistical tests for feature selection in classification and regression (linear and polynomial) models

    Jupyter Notebook