Looking for help with Regular Expression

ProvoWallis · May 24, 2006

Hi,

I'm looking for a little advice about regular expressions. I want to
capture a string of text that falls between an opening squre bracket
and a closing square bracket (e.g., "[" and "]") but I've run into a
small problem.

I've been using this: '''\[(.*?)\]''' as my pattern. I was expecting
this to be greedy but the funny thing is that it's not greedy enough in
some situations.

Here's my problem: The end of my string sometimes contains a cross
reference to a section in a book and the subsections are cited using
square brackets exactly like the one I'm using as the ending point in
my original regular expression.

E.g., the text string in my data looks like this: <core:emph
typestyle="it">see</core:emph> discussion in
§ 512.16[3]]

But my regular expression is stopping after the first "]" so after I
add the new markup the output looks like this:

<core:emph typestyle="it">see</core:emph> discussion in
§ 512.16[3]</fn:note>]

So the last subsection is outside of the note tag. I want something
like this:

<core:emph typestyle="it">see</core:emph> discussion in
§ 512.16[3]]</fn:note>

I'm not sure how to make my capture more greedy so I've resorted to
cleaning up the data after I make the first round of replacements:

data = re.sub(r'''\[(\d*?)\]</fn:note>\[(\w)\]\]''',
'''[\1][\2]]</fn:note>''', data)

There's got to be a better way but I'm not sure what it is.

Thanks,

Greg

James Stroud · May 24, 2006

ProvoWallis said:
Hi,

I'm looking for a little advice about regular expressions. I want to
capture a string of text that falls between an opening squre bracket
and a closing square bracket (e.g., "[" and "]") but I've run into a
small problem.

I've been using this: '''\[(.*?)\]''' as my pattern. I was expecting
this to be greedy but the funny thing is that it's not greedy enough in
some situations.

Here's my problem: The end of my string sometimes contains a cross
reference to a section in a book and the subsections are cited using
square brackets exactly like the one I'm using as the ending point in
my original regular expression.

E.g., the text string in my data looks like this: <core:emph
typestyle="it">see</core:emph> discussion in
§ 512.16[3]]

But my regular expression is stopping after the first "]" so after I
add the new markup the output looks like this:

<core:emph typestyle="it">see</core:emph> discussion in
§ 512.16[3]</fn:note>]

So the last subsection is outside of the note tag. I want something
like this:

<core:emph typestyle="it">see</core:emph> discussion in
§ 512.16[3]]</fn:note>

I'm not sure how to make my capture more greedy so I've resorted to
cleaning up the data after I make the first round of replacements:

data = re.sub(r'''\[(\d*?)\]</fn:note>\[(\w)\]\]''',
'''[\1][\2]]</fn:note>''', data)

There's got to be a better way but I'm not sure what it is.

I do: Pyparsing.

from pyparsing import *
crossref = Suppress("[") + Word(alphanums, exact=1) + Suppress("]")
footnote = (
Suppress("[") + SkipTo(crossref) +
ZeroOrMore(crossref) + Suppress("]")
)

footnote.parseString("[§ 512.16[3]]").asList()

py> footnote.parseString("[§ 512.16[3]]").asList()
['§ 512.16', '3', 'b']

James

--
James Stroud
UCLA-DOE Institute for Genomics and Proteomics
Box 951570
Los Angeles, CA 90095

http://www.jamesstroud.com/

lao_mage · May 24, 2006

'''\[(.*?)\]'''
?-> when this char after(*, +, ?, {n}, {n,}, {n,m}), the match pattern
is not greedy

e.g.1
String: 512.16[3]]
Pattern:'''\[(.*)\]'''
This will match "[3]]"

e.g.2
String: 512.16[3]]
Pattern:'''\[(.*)?\]'''
This will match "[3]" and ""

Roger Miller · May 24, 2006

Seem to be a lot of regular expression questions lately. There is a
neat little RE demonstrator buried down in
Python24/Tools/Scripts/redemo.py, which makes it easy to experiment
with regular expressions and immediately see the effect of changes. It
would be helpful if it were mentioned in the RE documentation, although
I can understand why one might not want a language reference to deal
with informally-supported tools.

Looking for programmers!	3	Feb 9, 2024
Looking For Help	2	Feb 8, 2024
Regular Expression for the special character "\|" pipe	7	May 27, 2014
Looking for a partner to team up with	0	Sep 23, 2023
Help with datascraping script	1	Aug 26, 2024
New coder looking for critique on fun project.	6	Jul 20, 2023
Regular Expression : Bad Character Range	0	Dec 20, 2013
Using a function for regular expression substitution	5	Aug 29, 2010

Looking for help with Regular Expression

ProvoWallis

James Stroud

lao_mage

Roger Miller

Ask a Question

Similar Threads

Members online

Forum statistics

Latest Threads