converting octal strings to unicode

flamingivanova · Dec 24, 2004

I have several ascii files that contain '\ooo' strings which represent
the octal value for a character. I want to convert these files to
unicode, and I came up with the following script. But it seems to me
that there must be a much simpler way to do it. Could someone more
experienced suggest some improvements?

I want to convert a file eg. containing:

hello \326du

with the unicode file containing:

hello Ödu

----------8<---------------------------------------
#!/usr/bin/python

import re, string, sys

if len(sys.argv) > 1:
file = open(sys.argv[1],'r')
lines = file.readlines()
file.close()
else:
print "give a filename"
sys.exit()

def to_unichr(str):
oct = string.atoi(str.group(1),8)
return unichr(oct)

for line in lines:
line = string.rstrip(unicode(line,'Latin-1'))
if re.compile(r'\\\d\d\d').search(line):
line = re.sub(r'\\(\d\d\d)', to_unichr, line)
line = line.encode('utf-8')
print line

----------8<---------------------------------------

Christos TZOTZIOY Georgiou · Dec 24, 2004

I have several ascii files that contain '\ooo' strings which represent
the octal value for a character. I want to convert these files to
unicode, and I came up with the following script. But it seems to me
that there must be a much simpler way to do it. Could someone more
experienced suggest some improvements?

decoded_string = "\326du".decode("string_escape")
unicode_text = unicode(decoded_string, "latin-1")

Christos TZOTZIOY Georgiou · Dec 24, 2004

I have several ascii files that contain '\ooo' strings which represent
the octal value for a character. I want to convert these files to
unicode, and I came up with the following script. But it seems to me
that there must be a much simpler way to do it. Could someone more
experienced suggest some improvements?

(hope I cancelled the previous off-by-one-backslash post...)

your_string = "\\326du"
decoded_string = your_string.decode("string_escape")
unicode_text = unicode(decoded_string, "latin-1")

Unicode strings as arguments to exceptions	3	Jan 16, 2014
split lines from stdin into a list of unicode strings	0	Aug 28, 2013
converting to and from octal escaped UTF--8	9	Dec 3, 2007
Convert unicode escape sequences to unicode in a file	1	Jan 11, 2011
Converting EBCDIC to Unicode	3	Sep 28, 2010
Ascii to Unicode.	4	Jul 28, 2010
FAQ 4.3 Why isn't my octal data interpreted correctly?	0	Jan 20, 2011
groveling over a file for Q:: and A:: stmts	3	Jul 24, 2012

converting octal strings to unicode

flamingivanova

Christos TZOTZIOY Georgiou

Christos TZOTZIOY Georgiou

Ask a Question

Similar Threads

Members online

Forum statistics

Latest Threads