In model development:
import pandas as pd
import numpy as np
np.random.seed(42)
bins = [0, 10, 15, 20, 25, 30, np.inf]
labels = bins[1:]
ages = list(range(5, 90, 5))
df = pd.DataFrame({"user_age": ages})
df["user_age_bin"] = pd.cut(df["user_age"], bins=bins, labels=False)
# sort by age
print(df.sort_values('user_age'))
In production, I will need to put individual age values to its corresponding bins. Here's how to do it:
# a new age value
new_age=30
# use this right=True and '-1' trick to make the bins match
print(np.digitize(new_age, bins=bins, right=True) -1)
MATLAB applications, tutorials, examples, tricks, resources,...and a little bit of everything I learned ...
Friday, February 28, 2020
Friday, June 14, 2019
speed up loading local csv file into AWS RDS MySQL database
tricks I learned today:
1. use 'LOAD LOCAL INFILE'
2. 'SET AUTOCOMMIT=0' - and manually commit at the end.
1. use 'LOAD LOCAL INFILE'
2. 'SET AUTOCOMMIT=0' - and manually commit at the end.
Friday, March 15, 2019
Python function: check if an object is an float
def isfloat(value):
try:
float(value)
return True
except ValueError:
return False
Thursday, March 14, 2019
Python function: format dollars
def format_dollar(s):
"""takes in a str or a number and format it as dollar format
i.e. u'24567.0' --> u'$24,567'
"""
s = str(s) # in case input is not string
try:
i = int(s.split('.')[0])
output = "$" + "{:,}".format(i)
except:
output = s
return output
"""takes in a str or a number and format it as dollar format
i.e. u'24567.0' --> u'$24,567'
"""
s = str(s) # in case input is not string
try:
i = int(s.split('.')[0])
output = "$" + "{:,}".format(i)
except:
output = s
return output
Tuesday, February 5, 2019
AWK: single quote eche line and add comma in the end
File.csv looks like:
line1
line2
line3
Use:
cat file.csv | awk -v a="'" '{print a$0a ","}'
to make it look like:
'line1',
'line2',
'line3',
line1
line2
line3
Use:
cat file.csv | awk -v a="'" '{print a$0a ","}'
to make it look like:
'line1',
'line2',
'line3',
Thursday, January 10, 2019
Python: Notes on Fluent Python
1.
2. List comprehension
a = [['-'] * 3 for i in range(3)]
b = [['-']*3] *3
What is the difference between a and b?
3. Inplace method
Inplace method returns None and does not create a new object. For example:
lst = [5,4,3,2,1]
lst.sort() # return None
4. Sort a list of strings by length
fruits = ['apple', 'grape', 'orange', 'banaba', 'dragon fruit']
sorted(fruits, key=len)
5. recursion
def factorial(n):
return 1 if n<2 else="" factorial="" n-1="" n="" p="">print(factorial(5))
6. from operator import itemgetter, attrgetter, methodcaller
2>
2. List comprehension
a = [['-'] * 3 for i in range(3)]
b = [['-']*3] *3
What is the difference between a and b?
3. Inplace method
Inplace method returns None and does not create a new object. For example:
lst = [5,4,3,2,1]
lst.sort() # return None
4. Sort a list of strings by length
fruits = ['apple', 'grape', 'orange', 'banaba', 'dragon fruit']
sorted(fruits, key=len)
5. recursion
def factorial(n):
return 1 if n<2 else="" factorial="" n-1="" n="" p="">print(factorial(5))
6. from operator import itemgetter, attrgetter, methodcaller
2>
Monday, December 31, 2018
Pandas: groupby and find the most frequent item
Say I have this dataframe:
order_id | class
1 | furniture
2 | book
2 | furniture
2 | book
3 | auto
3 | auto
3 | electronics
3 | pet
and to get the most frequent class of each order:
df.groupby('order_id').agg({'order_id': lambda x: x.value_counts().index[0]})
order_id | class
1 | furniture
2 | book
2 | furniture
2 | book
3 | auto
3 | auto
3 | electronics
3 | pet
and to get the most frequent class of each order:
df.groupby('order_id').agg({'order_id': lambda x: x.value_counts().index[0]})
Subscribe to:
Posts (Atom)
my-alpine and docker-compose.yml
``` version: '1' services: man: build: . image: my-alpine:latest ``` Dockerfile: ``` FROM alpine:latest ENV PYTH...
-
It took me a while to figure out how to insert a space in Mathtype equations. This is especially useful when you write an equation with mult...
-
In this post, I am trying to solve the problem given in the comments of one of the old post. Here's the problem, if I understand it co...
-
Recently I got a very long column of data and it contains lots of NaN. I found the finite function very useful to help me remove all the NaN...