yinyueyouge yinyueyouge - 28 days ago 14
HTML Question

Java: How to unescape HTML character entities in Java?

Basically I would like to decode a given Html document, and replace all special chars, such as

"&nbsp"
->
" "
,
">"
->
">"
.

In .NET we can make use of
HttpUtility.HtmlDecode
.

What's the equivalent function in Java?

Answer

I have used the Apache Commons StringEscapeUtils.unescapeHtml4() for this:

Unescapes a string containing entity escapes to a string containing the actual Unicode characters corresponding to the escapes. Supports HTML 4.0 entities.