Tuesday, March 3, 2020

Building a XGBoost model

1. pre-process :

           1.1 Boruta to clean data column-wise
           1.2 TomekLinks to clean data row-wise

2. training (use a random sample to do this step if the data size is too big):

          2.1 create a basic with fixed learning rate and n_estimators:
                                       XGBRFClassifier(objective='binary:logistic',
                                                                     n_estimator=X,
                                                                     learning_rate=0.1,
                                                                     n_jobs=-1)
         
          2.2 grid search for optimal 'max_depth' and 'min_child_weight'.
                          -- use scoring='roc_auc'
                          -- use cv = RepeatedStratifiedKFold(n_splits=3, n_repeats=3)

          2.3 get 'current_best' by using gridsearch.best_estimator, then fit current_best again.

         2.4 then grid search for 'gamma' with current_best, when it's done, get 'current_best' by using gridsearch.best_estimator, then fit current_best again.

         2.5 then grid search for 'subsample' and 'colsample_bytree', when it's done, get 'current_best' by using gridsearch.best_estimator, then fit current_best again.

         2.6 then grid search for 'learning_rate' , when it's done, get 'current_best' by using gridsearch.best_estimator, then fit current_best again.


3. final training: use all data to fit the mdl with all the optimized params.

4. Evaluatiion.
                 
               

Friday, February 28, 2020

Assign bin labels to new values during model inference

In model development:

import pandas as pd
import numpy as np
np.random.seed(42)

bins = [0, 10, 15, 20, 25, 30, np.inf]
labels = bins[1:]
ages = list(range(5, 90, 5))
df = pd.DataFrame({"user_age": ages})
df["user_age_bin"] = pd.cut(df["user_age"], bins=bins, labels=False)

# sort by age 
print(df.sort_values('user_age'))


In production, I will need to put individual age values to its corresponding bins. Here's how to do it:

# a new age value
new_age=30

# use this right=True and '-1' trick to make the bins match
print(np.digitize(new_age, bins=bins, right=True) -1)

Friday, June 14, 2019

speed up loading local csv file into AWS RDS MySQL database

tricks I learned today:
1. use 'LOAD LOCAL INFILE'
2. 'SET AUTOCOMMIT=0'  - and manually commit at the end.

Friday, March 15, 2019

Thursday, March 14, 2019

Python function: format dollars

 def format_dollar(s):
      """takes in a str or a number and format it as dollar format
      i.e. u'24567.0' --> u'$24,567'
     """


     s = str(s) # in case input is not string
     

     try:
         i = int(s.split('.')[0])
         output = "$" + "{:,}".format(i)
     except:
         output = s

     return output

Tuesday, February 5, 2019

AWK: single quote eche line and add comma in the end

File.csv looks like:

line1
line2
line3

Use:

cat file.csv | awk -v a="'" '{print a$0a ","}'

to make it look like:

'line1',
'line2',
'line3',

Thursday, January 10, 2019

Python: Notes on Fluent Python

1.

2. List comprehension

a = [['-'] * 3 for i in range(3)]

b = [['-']*3] *3

What is the difference between a and b?

3. Inplace method

Inplace method returns None and does not create a new object. For example:

lst = [5,4,3,2,1]
lst.sort() # return None


4. Sort a list of strings by length

fruits = ['apple', 'grape', 'orange', 'banaba', 'dragon fruit']
sorted(fruits, key=len)

5. recursion

def factorial(n):
    return 1 if n<2 else="" factorial="" n-1="" n="" p="">print(factorial(5))


6. from operator import itemgetter, attrgetter, methodcaller

my-alpine and docker-compose.yml

 ``` version: '1' services:     man:       build: .       image: my-alpine:latest   ```  Dockerfile: ``` FROM alpine:latest ENV PYTH...