Zhizhong John DIng
This is a cheatsheet and tutorial for the Numpy library.
NumPy is more compact, convenient than Python Lists. The vector and matrix operations in NumPy is well and efficiently implemented. NumPy is much faster and it also have more functionality such as FFTs, convolutions, fast searching, basic statistics, linear algebra, histograms, etc.
- In Numpy, dimension are called axes
- Numpy's class type is called ndarray
import numpy as py
a = np.array([[1,2,3],[4,5,6],[7,8,9],[10,11,12]])| Operator | Discription | Example | Return |
|---|---|---|---|
| ndarray.ndim | Return number of axis (dimension) | a.ndim | 2 |
| ndarray.shape | Return shape (m,n) | a.shape | (4,3) |
| ndarray.size | Return total number of elements | a.size | 12 |
| ndarray.dtype | Return object describing the type | a.dtype | int 32 |
| ndarray.itemsize | Return the size in terms of each element | a.itemsize | 4 |
| ndarray.data | Return the buffer (memory address) | a.data | <memory at 0x000000016D29B40> |
| ndarray.astype(type) | Convert to another data type | a.astype(str) | [['1' '2' '3'],['4' '5' '6'],['7' '8' '9'],['10' '11' '12']] |
| len(ndarray) | Return number of rows | len(a) | 4 |
ndarray.astype(type) is a great tool to convert data to int, float or str.
a = np.array([])
Create from python list, python tuple or python list of tuples, or python list of list. Results are same no matter from list of list or from tuple of list. Recommend using List to avoid confusion
| Operator | ndarray.shape | Return |
|---|---|---|
| a = np.array([1,2,3,4]) | 4 | [1 2 3 4] |
| a = np.array([[1,2],[3,4],[5,6]]) | (3,2) | [[1 2],[3 4],[5 6]] |
Note: You can also use tuple to create numpy Array
a = np.array((1,2,3,4))
a = np.array([(1,2),(3,4),(5,6)])
a = np.array(((1,2),(3,4),(5,6)))But personally I don't recommend using tuple. Because using list is consistent with python list for either array and matrix, easy to remember.
| Operator | Description | Return |
|---|---|---|
| a = np.arange(10,30,5) | same as python range | [10 15 20 25] |
| a = np.linspace(0,2,4) | start, stop, number | [0. 0.667 1.333 2] |
| a = np.fromfunction(f,(5,4)) | Construct an array by executing over each coord |
def f(x,y):
return 10*x+y
a = np.fromfunction(f,(5,4))
print(a)[[ 0. 1. 2. 3.]
[10. 11. 12. 13.]
[20. 21. 22. 23.]
[30. 31. 32. 33.]
[40. 41. 42. 43.]]
| Operator | Description | Return |
|---|---|---|
| a = np.empty((3,2)) | Create empty matrix | |
| a = np.empty(3) | Create empty matrix | |
| a = np.ones(4) | Create list (vector) | [1. 1. 1. 1.] |
| a = np.ones((3,2)) | Create matrix | [[1. 1.] [1. 1.] [1. 1.]] |
| a = np.zeros((3,2)) | Create matrix | [[0. 0.] [0. 0.] [0. 0.]] |
| a = np.eye(3) | Create Identity matrix (n) | [1., 0., 0.], [0., 1., 0.], [0., 0., 1.] |
| a = np.random.rand(3,2) | Create random matrix | uniform distribution between [0,1) |
| a = np.random.random((3,2)) | same as last one | uniform distribution between [0,1) |
| a = np.random.randn(3,2) | Create random matrix | normal distribution, mean 0 |
Original data: test_import.csv
Latitude,Longitude,Elevation
48.89016000,2.689270000,71.0
48.89000000,2.689730000,72.0
48.88987000,2.689810000,72.0
48.88924000,2.689570000,67.0
48.88934000,2.690050000,67.0
48.88949000,2.691400000,65.0
import csv
import numpy as np
data_path = 'test_import.csv'
with open(data_path, 'r') as f:
reader = csv.reader(f, delimiter=',')
# get header from first row
headers = next(reader)
# get all the rows as a list
data = list(reader)
# transform data into numpy array
data = np.array(data).astype(float)output looks like:
print(headers)
print(data.shape)
print(data[:3])['Latitude', 'Longitude', 'Elevation']
(6, 3)
[[48.89016 2.68927 71. ]
[48.89 2.68973 72. ]
[48.88987 2.68981 72. ]]
Comment:
- This method could be slow, if the data file is huge.
- next() returns the next item from the iterator
Two functions:
numpy.loadtxtnumpy.genfromtxt
numpy.genfromtxt is recommended over the other because np.genfromtxt can read CSV files with missing data and gives you options like the parameters missing_values and filling_values that help with missing values in the CSV.
Original data: test_import.csv
X,Y,Name,small,large,circular,mini_hoop,total_rack
982903.56993819773,205129.99858243763,1 7 AV S,5,0,0,0,5
987330.41607135534,191302.73030526936,1 BOERUM PL,1,0,0,0,1
983210.95318169892,199016.51343409717,1 CENTRE ST,10,0,0,0,10
985897.83954019845,207157.88527469337,1 E 13 ST,1,0,0,0,1
1010993.9694659412,252137.33960694075,1 E 183 ST,0,0,2,0,2
987774.37089210749,210586.44665901363,1 E 28 ST,1,0,0,0,1
import numpy as np
data_path = "test_import2.csv"
types = ['f8', 'f8', 'U50', 'i4', 'i4', 'i4', 'i4', 'i4']
data = np.genfromtxt(data_path, dtype=types, delimiter=',',names=True)
a = data['X']output looks like:
print(data)
print(data['X'])[( 982903.5699382 , 205129.99858244, '1 7 AV S', 5, 0, 0, 0, 5)
( 987330.41607136, 191302.73030527, '1 BOERUM PL', 1, 0, 0, 0, 1)
( 983210.9531817 , 199016.5134341 , '1 CENTRE ST', 10, 0, 0, 0, 10)
( 985897.8395402 , 207157.88527469, '1 E 13 ST', 1, 0, 0, 0, 1)
(1010993.96946594, 252137.33960694, '1 E 183 ST', 0, 0, 2, 0, 2)
( 987774.37089211, 210586.44665901, '1 E 28 ST', 1, 0, 0, 0, 1)]
[ 982903.5699382 987330.41607136 983210.9531817 985897.8395402 1010993.96946594 987774.37089211]
Comment:
- This method returns a tuple list
names = Trueto access the header and use it as column name to return a specific column
Note: if the data is generated in excel file, should export the data as csv file.
- get array from data
import pandas as pd
path = "test_import2.csv"
df=pd.read_csv(path, delimiter = ',', header = 0)
a = df['X'].valuesoutput looks like:
print(a)
print(type(a))
[ 982903.5699382 987330.41607136 983210.9531817 985897.8395402 1010993.96946594 987774.37089211]
<class 'numpy.ndarray'>
- get matrix from data
import pandas as pd
path = "test_import2.csv"
df=pd.read_csv(path, delimiter = ',', header = 0)
a = df.iloc[:,:2].valuesoutput looks like:
print(a)
print(type(a))
[[ 982903.5699382 205129.99858244]
[ 987330.41607136 191302.73030527]
[ 983210.9531817 199016.5134341 ]
[ 985897.8395402 207157.88527469]
[1010993.96946594 252137.33960694]
[ 987774.37089211 210586.44665901]]
<class 'numpy.ndarray'>
Rule: all operations for array (vector) are element wise.
a = np.array([20,30,40,50])
b = np.arange(4) # b = np.array([0,1,2,3])| Operator | Description | Return |
|---|---|---|
| a+b / np.add(a,b) | Addition | [20 31 42 53] |
| a-b / np.substract(a,b) | Subtraction | [20 29 38 47] |
| a/b / np.divide(a,b) | Division | [inf 30. 20. 16.67] |
| a*b / np.multiply(a,b) | Multiplication | [ 0 30 80 150] |
| b**2 | Square | [0 1 4 9] |
| np.sqrt(b) | Square Root | [0. 1. 1.4141.732] |
| np.sin(a) | Treat element as Radian | [ 0.912 -0.988 0.745 -0.262] |
| np.cos(a) | Treat element as Radian | |
| np.log(a) | Base 2 | |
| np.dot(a,b) / np.sum(a*b) | Dot product | 260 |
| a<35 | comparison | [True True False False] |
Rule: except matrix product, all operations are element wise.
a = np.array([[1,1],[0,1]])
[[1 1]
[0 1]]
b = np.array([[2,0],[3,4]])
[[2 0]
[3 4]]| Operator | Description | Return |
|---|---|---|
| a+b / np.add(a,b) | Addition | [[3 1] [3 5]] |
| b-a / np.substract(b,a) | Subtraction | [[ 1 -1] [ 3 3]] |
| b/a / np.divide(b,a) | Division | [[ 2. 0.] [inf 4.]] |
| a*b / np.multiply(a,b) | Multiplication | [[2 0] [0 4]] |
| b**2 | Square | [[ 4 0] [ 9 16]] |
| np.sqrt(b) | Square Root | [[1.41421356 0. ] [1.73205081 2. ]] |
| np.sin(a) | Treat element as Radian | |
| np.cos(a) | Treat element as Radian | |
| np.log(a) | Base 2 | |
| a@b /a.dot(b) | Matrix dot product | [[5 4] [3 4]] |
| a<35 | comparison | [[ True True] [ True True]] |
cumsum min exp etc needs update
under construction
| Operator | Description |
|---|---|
| a.reshape(2,6) | Won't change array itself |
| a.resize(2,6) | Will change array itself |
| b = np.ravel(a) | Return 1D flattened array, itself doesn't change |
| a.flatten() | Return 1D flattened array, itself doesn't change |
Comment:
- A key difference between
flatten()andravel()is thatflatten()is a method of anndarrayobject and hence can only be called for true numpy arrays. In contrastravel()is a library-level function and hence can be called on any object that can successfully be parsed. - Turn 1D array into 2D array:
import numpy as np
A = np.arange(8)
A = A.reshape(1,8)
#or
A = A.reshape(8,1)def f(x,y):
return 10*x+y
a = np.fromfunction(f,(3,4))print(a)[[ 0. 1. 2. 3.]
[10. 11. 12. 13.]
[20. 21. 22. 23.]]
b = a.reshape(2,6)
print(b)
print(a)[[ 0. 1. 2. 3. 10. 11.]
[12. 13. 20. 21. 22. 23.]]
[[ 0. 1. 2. 3.]
[10. 11. 12. 13.]
[20. 21. 22. 23.]]
b = a.resize(2,6)
print(b)
print(a)None
Comment: resize works on the array itself, it returns None
[[ 0. 1. 2. 3. 10. 11.]
[12. 13. 20. 21. 22. 23.]]
b = np.ravel(a)
print(b)
print(a)[ 0. 1. 2. 3. 10. 11. 12. 13. 20. 21. 22. 23.]
[[ 0. 1. 2. 3.]
[10. 11. 12. 13.]
[20. 21. 22. 23.]]
b = a.flatten()
print(b)
print(a)[ 0. 1. 2. 3. 10. 11. 12. 13. 20. 21. 22. 23.]
[[ 0. 1. 2. 3.]
[10. 11. 12. 13.]
[20. 21. 22. 23.]]
| Operator | Description | Graph |
|---|---|---|
| np.hstack((A,B)) | Stack horizontally | AB |
| np.vstack((A,B)) | Stack vertically | A B |
| np.column_stack((A,B)) | Stack 1-D arrays as columns into a 2-D array | A.T B.T |
| np.append(A,B) | form 1D array of AB | AB |
| np.append(A,B, axis=0) | stack vertically | A B |
| np.append(A,B, axis=1) | stack horizontally | AB |
Comment:
- I don't recommend using concatenate. It's basically same as operators mentioned above by changing axis. And it could be confusing.
- For
np.append, when axis is specified, values must have the correct dimension(2D and 2Dor1D and 1D)
import numpy as np
A = np.arange(2,6)
B = np.arange(1,5)*2
print(A,'\n',B)[2 3 4 5]
[2 4 6 8]
print(np.hstack((A,B)))[2 3 4 5 2 4 6 8]
print(np.column_stack((A,B)))[[2 2]
[3 4]
[4 6]
[5 8]]
print(np.vstack((A,B)))[[2 3 4 5]
[2 4 6 8]]
np.append([[1, 2, 3], [4, 5, 6]], [[7, 8, 9]], axis=0)[[1, 2, 3]
[4, 5, 6]
[7, 8, 9]]
np.append([1, 2, 3], [[4, 5, 6], [7, 8, 9]])[1, 2, 3, ..., 7, 8, 9]
- add a 1D vector to 2D array (stack horizontally)**
A_2d = np.array([[2, 3],
[4, 5],
[6, 7],
[8, 9]])
B_1d = np.array([1, 2, 3, 4])
B_1d_reshape = B_1d.reshape(len(B_1d),1)
new_matrix = np.hstack((A_2d,B_1d_reshape))
print(new_matrix)[[2 3 1]
[4 5 2]
[6 7 3]
[8 9 4]]
- Put together multiple arrays into a matrix
# matrix: t, x, y, v
t = np.array([0,1,2,3,4])
x = np.array([-2,-1,0,1,2])
y = np.array([-3,-1.5,0,1.5,3])
v = np.array([10,20,30,40,50])
new_matrix = np.hstack((t.reshape(len(t),1),
x.reshape(len(x),1),
y.reshape(len(y),1),
v.reshape(len(v),1)))
print(new_matrix)[[ 0. -2. -3. 10. ]
[ 1. -1. -1.5 20. ]
[ 2. 0. 0. 30. ]
[ 3. 1. 1.5 40. ]
[ 4. 2. 3. 50. ]]
Same as python list.
import numpy as np
a = np.arange(10)**2
print(a)[ 0 1 4 9 16 25 36 49 64 81]
print(a[2])4
print(a[2:5])[ 4 9 16]
print(a[::-1])[81 64 49 36 25 16 9 4 1 0]
Same as python 2D list
def f(x,y):
return 10*x+y
a = np.fromfunction(f,(5,4))print(a[2,3])23
print(a[1:4,2])[12. 22. 32.]
print(a[1:3,:])[[10. 11. 12. 13.]
[20. 21. 22. 23.]]
print(a[-1])[40. 41. 42. 43.]
a.flat is a 1D iterator over the array
def f(x,y):
return 10*x+y
a = np.fromfunction(f,(3,4))
print(a)
print(a.flat[6])[[ 0. 1. 2. 3.]
[10. 11. 12. 13.]
[20. 21. 22. 23.]]
12.0