encoding misunderstanding

Tim Arnold · Jul 27, 2007

Hi, I'm beginning to understand the encode/decode string methods, but
I'd like confirmation that I'm still thinking in the right direction:

I have a file of latin1 encoded text. Let's say I put one line of that
file
into a string variable 'tocline', as follows:
tocline = 'Ficha Datos de p\xe9rdida AND acci\xf3n'

import codecs
tocFile =
codecs.open('mytoc.htm','wb',encoding='utf8',errors='replace')
tocline = tocline.decode('latin1','replace')
tocFile.write(tocline)
tocFile.close()

What I think is that tocFile is wrapped to insure that anything
written to it is in utf8
I decode the latin1 string into python's internal unicode encoding and
that gets written out as utf8.

Questions:
what exactly is the tocline when it's read in with that \xe9 and \xed
in the string? A latin1 encoded string?
Is my method the right way to write such a line out to a file with
utf8
encoding?

If I read in the latin1 file using
codecs.open(filename,encoding='latin1') and write out the utf8 file
by
opening with
codecs.open(othername,encoding='utf8'), would I no longer have a
problem -- I could just read in latin1 and write out utf8 with no
more worries about
encoding?

thanks,
--Tim
p.s. sorry if you see this twice--my newsreader is flaky right now.

encode/decode misunderstanding	3	Jul 26, 2007
encoding error	1	Feb 20, 2013
Problem with a login script, SESSION user rights and put this together so it works with the other pages and MySQL. Code examples.	2	May 5, 2023
encoding problem with BeautifulSoup - problem when writing parsedtext to file	9	Oct 6, 2011
Cyrillic text from file - set utf8 in cmd, unknown characters output anyway	0	Nov 11, 2022
A few questiosn about encoding	103	Jun 9, 2013
encoding problem	11	Dec 19, 2008
Python 2.1 / 2.3: xreadlines not working with codecs.open	3	Jun 23, 2005

encoding misunderstanding

Tim Arnold

Ask a Question

Similar Threads

Members online

Forum statistics

Latest Threads