Showing posts with label python. Show all posts
Showing posts with label python. Show all posts

Tuesday, December 4, 2018

Python regex to match decimal numbers

Here is a Python regex to match regular decimal numbers (without the e or exponent part):
r'([+-]?([0-9]+(\.[0-9]*)?|\.[0-9]+))'

Wednesday, October 3, 2018

Timing code in Python

To time a particular code block in Python, do this:

import time

t = time.time()
#<code_block>
print(f'code-block-name:{time.time()-t})

[Courtesy: this]




Friday, September 21, 2018

Specify data types when creating Pandas dataframe from csv

To specify data types for various columns when creating a Pandas dataframe, do this:

df = pd.read_csv(filename, dtype={'col1':np.int64, 'col2':str, ...})

If converters are given, they will be used instead of data type conversion.

Convert Python / numpy datetime to Unix timestamp

To convert Python datetime to Unix / POSIX timestamp (float), do this:

mytime.timestamp()

[Courtesy: this]

To convert numpy.timestamp to Unix / POSIX timestamp (float), do this:

mytime.astype('uint64')

[Courtesy: this]

Changing dimensions of plot in matplotlib

To specify dimensions of a plot in matplotlib (e.g. to get longer x or y axis), do this:

fig = plt.figure(figsize=(20,3)) #x-axis = 20", y-axis=3"
ax = fig.add_subplot(111)
ax.plot(x, y)

[Courtesy: this]

Show summary statistics for a Pandas dataframe

To see the summary statistics for a Pandas dataframe, do this:

df.describe()

This will show things like: count, mean, standard deviation, etc.

To see column details (data types), do this:

df.info()


Cast a Pandas dataframe column to timestamp

To cast / convert a column in a Pandas dataframe to the timestamp datatype, do this:

df['timestamp'] = pd.to_datetime(df['timestamp'])


Jupyter-notebook: plot matplotlib graphs inline

To plot matplotlib graphs inline in Jupyter-notebook, do this:

import matplotlib.pyplot as plt
%matplotlib inline

Now, plt.plot(...) while show the graph inline.

Jupyter-notebook module reload

To reload modules being used in a jupyter-notebook (when they have been changed outside), do this:

%load_ext autoreload
%autoreload 2

This will automatically reload the module every time it is changed.

[Courtesy: this]

Monday, November 17, 2014

Passing lists in numba

If you plan to @jit() a Python function using numba, and one of the arguments is a list, it will be treated as an object, and the jit-ted function will probably be slower than the original. Instead, explicitly specify the argument data types (e.g. int32, double, unit64, etc.) within the @jit([data types]) declaration and, importantly, when calling the function, remember to convert the list into a numpy array of the data type specified in the jit declaration. For all your efforts, you should be rewarded with a good speed-up (if you have long-running loops).

Wednesday, October 15, 2014

Debugging in iPython

To debug a function func in module mod.py, do the following at the iPython prompt.
from IPython.core.debugger import Pdb
ipdb = Pdb()
import mod
ipdb.runcall(mod.func, [args, for, func])

This will start the iPython debugger at the first line of func.

Saturday, September 27, 2014

Code formatting in Blogger

I found this answer on stackoverflow about how to format source code in blogs on blogger, which pointed this tutorial with detailed instructions. It worked for me with Python! Many Thanks to David Craft and Alex Gorbatchev.

Also note the useful tip here about loading only the format styles that you need for your blog, to speed up page load.

How-to (for Python)

Add the following to Blogger template above the </head> tag.







In the blog editor, switch to HTML mode and insert code within <pre> as follows.

You can find more brushes here.

A matplotlib example

The following matplotlib script shows some cases where the matplotlib defaults were not sufficient for the job, and how to customize the relevant properties. The main tweaks were:
  • changing the font type from Type 3 to True Type
  • scaling the y values
  • using latex syntax in the labels
  • changing the fontsize for axis labels, axis ticks, and legends
  • two legends
  • using proxy objects (lines or patches) for the legends
  • using axis coordinates for locating the legend
  • semi-transparent legends and grid lines
Here is the python code and the resulting figure:

import matplotlib.pyplot as plt
import matplotlib as mpl
import matplotlib.patches as mpatches
import numpy as np
import json
import sys
import os

# change font type to True Type to avoid Type 3 fonts
# (which are not allowed by some conferences)

mpl.rcParams['pdf.fonttype'] = 42


# x = 1-d array with x values
# ys = six 1-d arrays with y values (one for each line we want to plot).
#      Of these 3 are of one kind, and the other 3 are of another kind.
# xlabel, ylabel = labels for the x and y axis

x, ys, xlabel, ylabel = get_plot_info(datafile)


# I want to scale the y values so that the y axis is easier to read.

ys = ys / 10   # ys is a numpy array


# The plot will be shrunk when it is included in the paper, so that the 
# default fontsize becomes too small. Select a larger fontsize globally.

fontsize = 30


# Plot the first three lines in blue with different line styles;
# then plot the other three in green with similar line styles.

plt.plot(x, ys[0], 'b-', lw=3)
plt.plot(x, ys[1], 'b--', lw=3)
plt.plot(x, ys[2], 'b:', lw=3)
plt.plot(x, ys[3], 'g-', lw=3)
plt.plot(x, ys[4], 'g--', lw=3)
plt.plot(x, ys[5], 'g:', lw=3)


# Set axis labels with larger fontsize.
# Mention the scaling done to the y values.

plt.xlabel(xlabel, fontsize=fontsize)
plt.ylabel(ylabel + r' $\div 10$', fontsize=fontsize)  # latex syntax works!


# Increase the fontsize of the axis ticks

ax = plt.gca()
for labx in ax.get_xticklabels():
    labx.set_fontsize(fontsize)
for laby in ax.get_yticklabels():
    laby.set_fontsize(fontsize)


# If you want the lines to reach the left and right extremities of the graph,
# reset the x-limits.
ax.set_xlim( (min(x), max(x)) )

# Instead of showing 6 entries in the legend (one for each of the 6 lines),
# we show one legend for the color code, and another for the line style.
# ('MCTM-...' are methods, and 'en' etc. are languages for which we ran
# the method.)

# first legend (for color code)

blue_patch = mpatches.Patch(color='blue')
green_patch = mpatches.Patch(color='green')

# loc=(x,y) are the coordinates of the lower left corner of the legend,
# where (0,0) if the lower left corner of the axes, and (1,1) is the
# upper right corner.
# framealpha is the transparency level (0=transparent, 1=opaque).

leg1 = plt.legend((blue_patch, green_patch), ('MCTM-DSGNP','MCTM-D'), 
                  loc=(.3,.6), fontsize=fontsize, framealpha=.5)
plt.gca().add_artist(leg1)

# second legend (for line style)

# plot empty arrays in black (to avoid the colors used above).

l_en, = plt.plot([], [], 'k-', lw=3, label='en')
l_hi, = plt.plot([], [], 'k--', lw=3, label='hi')
l_hir, = plt.plot([], [], 'k:', lw=3, label='hir')
plt.legend((l_en, l_hi, l_hir), ('en','hi', 'hir'), 
                  loc='lower right', fontsize=fontsize, framealpha=.5)

# Add grid lines, but make them semi-transparent (otherwise they appear too
# prominent when the figure is shrunk).

plt.grid(alpha=.5)


# The large fontsize pushes axis labels out of the figure; but matplotlib
# offers a function to auto-correct this.

plt.tight_layout()


# The saving format is chosen automatically based on the file extension.
plt.savefig(figname+'.pdf')

Monday, November 25, 2013

Very Quick File Server

Just cd into the desired directory, and run python -m SimpleHTTPServer. You get the message "Serving HTTP on 0.0.0.0 port 8000 ...". Now go the browser and type localhost:8000, and start browsing files/folders in that directory. Share the link http://<your IP addr>:8000 to let your friends browse as well.

[Tested on Ubuntu 10.04 (64-bit) with Python 2.6.5.]

Monday, June 24, 2013

Python tips

To compile a python script without executing it, do
python -m py_compile my_script.py
[Courtesy: stackoverflow]

Wednesday, June 5, 2013

Installing Python regex on Ubuntu 12.04

To install the alternative regular expression module regex (which has richer Unicode support than re) for Python:
  • Install python-dev from the Software Center.
  • Download the archive, extract, and run sudo python setup.py install.

Thursday, May 2, 2013

Unicode tips for Python

  • To use non-Latin characters in regular expressions, use u'...' instead of r'...', even if you have to escape every backslash; e.g. the regex u'(?u)[०-९]\\s' matches a Devanagari digit followed by whitespace.
  • Remove zero-width joiners/non-joiners from Unicode text to get a normalized representation; otherwise words that are rendered the same in a browser/editor will be stored differently, and will not be equal on comparison; e.g. use the regex u'[\\u200D\\u200C]' and replace all matches with u'' (the empty string).

Thursday, January 17, 2013

File compression in python

Normally, you can do
 f = open(fname, 'r')  
 data = f.read()  
 f.close()  
to read a file, and
 f = open(fname, 'w')  
 f.write(text)  
 f.close()  
to write a file.

To use gzip compression, do
 import gzip  
 f = gzip.open(fname, 'rb')  
 data = f.read()  
 f.close()  
to read a gzip-compressed file, and
 f = gzip.open(fname, 'wb')  
 f.write(data)  
 f.close()  
to write a gzip-compressed file.

To compress pickled objects, just use the file handle returned by gzip.open(); everything else remains the same.

(Courtesy Doug Hellmann.)

Friday, March 23, 2012

Debugging and Profiling in Python

To debug hello.py, do
$ python -m pdb hello.py

You will immediately get the debug prompt, with the execution stopped before the first line of the program. You can now use:

l - to display source code around the line to be executed next
n - to execute the next line of code
s - to step into the next line of code
b <line no.> - to set breakpoints
b - to list breakpoints
clear <breakpt no.> - to clear breakpoints
c - to continue execution
tbreak <line no.> followed by c - to execute till a particular line
p <var> - to print contents of variables,
vars(obj) - to print contents of objects (same as obj.__dict__)
dir(obj) - to list attributes/methods of objects
help - to get more help.


 To profile hello.py, do
 $ python -m cProfile hello.py

The program will be run and you will get a profiler report on the console.

To debug a particular part of the code, do
import cProfile
profile = cProfile.Profile()
profile.enable()
# do something
profile.disable()
profile.dump_stats('something.prof')

To view the dump file, use cprofilev.

Monday, February 20, 2012

Literal strings in Python

Some of the many ways one can construct strings in Python:

s = "qwe\nasd\nzxc"  #gives 'qwe\nasd\nzxc'
s = 'qwe\nasd\nzxc'  #gives 'qwe\nasd\nzxc'
s = 'qwe\
\nasd\
\nzxc'  #gives 'qwe\nasd\nzxc'
s = '''qwe
asd
zxc'''  #gives 'qwe\nasd\nzxc'
s = '''qwe\\n
asd
zxc'''  #gives 'qwe\\n\nasd\nzxc'
s = r'qwe\n\
asd\
zxc'  #gives 'qwe\\n\\\nasd\\\nzxc'; => \n in the string is always accompanied by a '\'.
s = u'qwe\u0020asd'  #gives a Unicode string 'qwe asd' (the Unicode code for space is 20)
s = ur'qwe\u0020asd'  #gives a Unicode string 'qwe asd'
s = ur'qwe\\u0020asd'  #gives a Unicode string 'qwe\\\\u0020asd' (the desired format for use in regular expressions.
s = u'qwe\u0020asd'.encode('utf-8')  #gives an encoding of the Unicode string in the specified encoding ('utf-8')
s1 = unicode(s, 'utf-8')  #gives a Unicode string after decoding the content of s according to specified encoding