jericho解析html
2021-06-30 18:04
标签:jericho解析html 本文出自 “素颜” 博客,请务必保留此出处http://suyanzhu.blog.51cto.com/8050189/1945451 jericho解析html 标签:jericho解析html 原文地址:http://suyanzhu.blog.51cto.com/8050189/19454511.导入jar包
2.实现源代码
package com.zhishang.lucene;
import net.htmlparser.jericho.Element;
import net.htmlparser.jericho.HTMLElementName;
import net.htmlparser.jericho.Source;
import org.junit.Test;
import java.io.File;
import java.io.IOException;
/**
* Created by Administrator on 2017/7/8.
*/
public class HtmlBeanUtil {
@Test
public void parseHtml(){
String path = "G:\\data\\index.html";
try {
Source sc = new Source(new File(path));
Element element = sc.getFirstElement(HTMLElementName.TITLE);
System.out.println(element.getTextExtractor().toString());
System.out.println(sc.getTextExtractor().toString());
} catch (IOException e) {
e.printStackTrace();
}
}
}
下一篇:lucene创建索引