html 显示纯文本,标签也显示出来 15

类似<div></div>显示成文本或者〈imgsrc="">显示的是原文而不是图片。... 类似<div></div>显示成文本或者〈img src="">显示的是原文而不是图片。展开

 我来答

2个回答

#热议# 空调使用不当可能引发哪些疾病？

郭某人来此
推荐于2017-09-07 · TA获得超过1646个赞

知道答主

回答量：952

采纳率：100%

帮助的人：91万

我也去答题访问个人页

关注

展开全部

不知道这个用的着不！

在网页刚流行起来的时候，提取html中的文本有一个简单的方法，就是将html文本（包含标记）中的所有以“<”符号开头到以“>”符号之间的内容去掉即可。
但对于现在复杂的网页而言，用这种方法提取出来的文本会有大量的空格、空行、script段落、还有一些html转义字符，效果很差。
下面用正则表达式来提取html中的文本，
代码的实现的思路是：
a、先将html文本中的所有空格、换行符去掉（因为html中的空格和换行是被忽略的）
b、将<head>标记中的所有内容去掉
c、将<script>标记中的所有内容去掉
d、将<style>标记中的所有内容去掉
e、将td换成空格，tr,li,br,p 等标记换成换行符
f、去掉所有以“<>”符号为头尾的标记去掉。
g、转换&，&nbps;等转义字符换成相应的符号
h、去掉多余的空格和空行
代码如下：

using System;
using System.Text.RegularExpressions;
namespace Kwanhong.Utilities
{
/// <summary>
/// HtmlToText 的摘要说明。
/// </summary>
public class HtmlToText
{
public string Convert(string source)
{
string result;
//remove line breaks,tabs
result = source.Replace("\r", " ");
result = result.Replace("\n", " ");
result = result.Replace("\t", " ");
//remove the header
result = Regex.Replace(result, "(<head>).*(</head>)", string.Empty, RegexOptions.IgnoreCase);
result = Regex.Replace(result, @"<( )*script([^>])*>", "<script>", RegexOptions.IgnoreCase);
result = Regex.Replace(result, @"(<script>).*(</script>)", string.Empty, RegexOptions.IgnoreCase);
//remove all styles
result = Regex.Replace(result, @"<( )*style([^>])*>", "<style>", RegexOptions.IgnoreCase); //clearing attributes
result = Regex.Replace(result, "(<style>).*(</style>)", string.Empty, RegexOptions.IgnoreCase);
//insert tabs in spaces of <td> tags
result = Regex.Replace(result, @"<( )*td([^>])*>", " ", RegexOptions.IgnoreCase);
//insert line breaks in places of <br> and <li> tags
result = Regex.Replace(result, @"<( )*br( )*>", "\r", RegexOptions.IgnoreCase);
result = Regex.Replace(result, @"<( )*li( )*>", "\r", RegexOptions.IgnoreCase);
//insert line paragraphs in places of <tr> and <p> tags
result = Regex.Replace(result, @"<( )*tr([^>])*>", "\r\r", RegexOptions.IgnoreCase);
result = Regex.Replace(result, @"<( )*p([^>])*>", "\r\r", RegexOptions.IgnoreCase);
//remove anything thats enclosed inside < >
result = Regex.Replace(result, @"<[^>]*>", string.Empty, RegexOptions.IgnoreCase);
//replace special characters:
result = Regex.Replace(result, @"&", "&", RegexOptions.IgnoreCase);
result = Regex.Replace(result, @" ", " ", RegexOptions.IgnoreCase);
result = Regex.Replace(result, @"<", "<", RegexOptions.IgnoreCase);
result = Regex.Replace(result, @">", ">", RegexOptions.IgnoreCase);
result = Regex.Replace(result, @"&(.{2,6});", string.Empty, RegexOptions.IgnoreCase);
//remove extra line breaks and tabs
result = Regex.Replace(result, @" ( )+", " ");
result = Regex.Replace(result, "(\r)( )+(\r)", "\r\r");
result = Regex.Replace(result, @"(\r\r)+", "\r\n");
return result;
}
}//end class
}//end namespace

已赞过 已踩过<

评论收起

lx6692
2014-09-18 · 超过41用户采纳过TA的回答

知道小有建树答主

回答量：86

采纳率：0%

帮助的人：47.7万

我也去答题访问个人页

关注

展开全部

因为html解析是< 和 >这两个尖括号，所以不管你用什么方法带尖括号的都是显示不出来的,但是只要html页面加载时找不到<，>符号就可以用，但是实现不了你的需求。
举例：<div>不可识别,div是可以识别的,呵呵：）
希望帮到你：）

本回答被网友采纳

已赞过已踩过<

你对这个回答的评价是？
评论收起

推荐律师服务：若未解决您的问题，请您详细描述您的问题，通过百度律临进行免费专业咨询

html 显示纯文本,标签也显示出来 15

其他类似问题

为你推荐：