Search code examples
regexhtml-content-extractiontext-extraction

How to extract values from HTML using RegEx?


Given the following HTML:

<p><span class="xn-location">OAK RIDGE, N.J.</span>, <span class="xn-chron">March 16, 2011</span> /PRNewswire/ -- Lakeland Bancorp, Inc. (Nasdaq:   <a href='http://studio-5.financialcontent.com/prnews?Page=Quote&Ticker=LBAI' target='_blank' title='LBAI'> LBAI</a>), the holding company for Lakeland Bank, today announced that it redeemed <span class="xn-money">$20 million</span> of the Company's outstanding <span class="xn-money">$39 million</span> in Fixed Rate Cumulative Perpetual Preferred Stock, Series A that was issued to the U.S. Department of the Treasury under the Capital Purchase Program on <span class="xn-chron">February 6, 2009</span>, thereby reducing Treasury's investment in the Preferred Stock to <span class="xn-money">$19 million</span>. The Company paid approximately <span class="xn-money">$20.1 million</span> to the Treasury to repurchase the Preferred Stock, which included payment for accrued and unpaid dividends for the shares. &#160;This second repayment, or redemption, of Preferred Stock will result in annualized savings of <span class="xn-money">$1.2 million</span> due to the elimination of the associated preferred dividends and related discount accretion. &#160;A one-time, non-cash charge of <span class="xn-money">$745 thousand</span> will be incurred in the first quarter of 2011 due to the acceleration of the Preferred Stock discount accretion. &#160;The warrant previously issued to the Treasury to purchase 997,049 shares of common stock at an exercise price of <span class="xn-money">$8.88</span>, adjusted for stock dividends and subject to further anti-dilution adjustments, will remain outstanding.</p>

I'd like to get the values inside the <span> elements. I'd also like to get the value of the class attribute on the <span> elements.

Ideally I could just run some HTML through a function and get back a dictionary of extracted entities (based on the <span> parsing defined above).

The above code is a snippet from a larger source HTML file, which fails to pare with an XML parser. So I'm looking for a possible regular expression to help extract the information of interest.


Solution

  • Use this tool (free): http://www.radsoftware.com.au/regexdesigner/

    Use this Regex:

    "<span[^>]*>(.*?)</span>"
    

    The values in Group 1 (for each match) will be the text that you need.

    In C# it will look like:

                Regex regex = new Regex("<span[^>]*>(.*?)</span>");
                string toMatch = "<span class=\"ajjsjs\">Some text</span>";
                if (regex.IsMatch(toMatch))
                {
                    MatchCollection collection = regex.Matches(toMatch);
                    foreach (Match m in collection)
                    {
                        string val = m.Groups[1].Value;
                        //Do something with the value
                    }
                }
    

    Ammended to answer comment:

                Regex regex = new Regex("<span class=\"(.*?)\">(.*?)</span>");
                string toMatch = "<span class=\"ajjsjs\">Some text</span>";
                if (regex.IsMatch(toMatch))
                {
                    MatchCollection collection = regex.Matches(toMatch);
                    foreach (Match m in collection)
                    {
                        string class = m.Groups[1].Value;
                        string val = m.Groups[2].Value;
                        //Do something with the class and value
                    }
                }