Python实现简单HTML表格解析的方法

   本文实例讲述了Python实现简单HTML表格解析的方法。分享给大家供大家参考。具体分析如下:

  这里依赖libxml2dom,确保首先安装!导入到你的脚步并调用parse_tables() 函数。

  1. source = a string containing the source code you can pass in just the table or the entire page code

  2. headers = a list of ints OR a list of strings

  If the headers are ints this is for tables with no header, just list the 0 based index of the rows in which you want to extract data.

  If the headers are strings this is for tables with header columns (with the tags) it will pull the information from the specified columns

  3. The 0 based index of the table in the source code. If there are multiple tables and the table you want to parse is the third table in the code then pass in the number 2 here

  It will return a list of lists. each inner list will contain the parsed information.

  具体代码如下:

  ?

  1

  2

  3

  4

  5

  6

  7

  8

  9

  10

  11

  12

  13

  14

  15

  16

  17

  18

  19

  20

  21

  22

  23

  24

  25

  26

  27

  28

  29

  30

  31

  32

  33

  34

  35

  36

  37

  38

  39

  40

  41

  42

  43

  44

  45

  46

  47

  48

  49

  50

  51

  52

  53

  54

  55

  56

  57

  58

  59

  60

  61

  62

  63

  64

  65

  66

  67

  68

  69

  70

  71

  72

  73

  74

  75

  76

  77

  78

  79

  80

  81

  82

  83

  84

  85

  86

  87

  88

  89

  90

  91

  92

  93

  94

  95

  96

  97

  98

  99

  100

  101

  102

  103

  104

  105

  106

  107

  108

  109

  110

  111

  112

  113

  114

  115

  116

  117

  118#The goal of table parser is to get specific information from specific

  #columns in a table.

  #Input: source code from a typical website

  #Arguments: a list of headers the user wants to return

  #Output: A list of lists of the data in each row

  import libxml2dom

  def parse_tables(source, headers, table_index):

  """parse_tables(string source, list headers, table_index)

  headers may be a list of strings if the table has headers defined or

  headers may be a list of ints if no headers defined this will get data

  from the rows index.

  This method returns a list of lists

  """

  #Determine if the headers list is strings or ints and make sure they

  #are all the same type

  j = 0

  print 'Printing headers: ',headers

  #route to the correct function

  #if the header type is int

  if type(headers[0]) == type(1):

  #run no_header function

  return no_header(source, headers, table_index)

  #if the header type is string

  elif type(headers[0]) == type('a'):

  #run the header_given function

  return header_given(source, headers, table_index)

  else:

  #return none if the headers aren't correct

  return None

  #This function takes in the source code of the whole page a string list of

  #headers and the index number of the table on the page. It returns a list of

  #lists with the scraped information

  def header_given(source, headers, table_index):

  #initiate a list to hole the return list

  return_list = []

  #initiate a list to hold the index numbers of the data in the rows

  header_index = []

  #get a document object out of the source code

  doc = libxml2dom.parseString(source,html=1)

  #get the tables from the document

  tables = doc.getElementsByTagName('table')

  try:

  #try to get focue on the desired table

  main_table = tables[table_index]

  except:

  #if the table doesn't exits then return an error

  return ['The table index was not found']

  #get a list of headers in the table

  table_headers = main_table.getElementsByTagName('th')

  #need a sentry value for the header loop

  loop_sentry = 0

  #loop through each header looking for matches

  for header in table_headers:

  #if the header is in the desired headers list

  if header.textContent in headers:

  #add it to the header_index

  header_index.append(loop_sentry)

  #add one to the loop_sentry

  loop_sentry+=1

  #get the rows from the table

  rows = main_table.getElementsByTagName('tr')

  #sentry value detecting if the first row is being viewed

  row_sentry = 0

  #loop through the rows in the table, skipping the first row

  for row in rows:

  #if row_sentry is 0 this is our first row

  if row_sentry == 0:

  #make the row_sentry not 0

  row_sentry = 1337

  continue

  #get all cells from the current row

  cells = row.getElementsByTagName('td')

  #initiate a list to append into the return_list

  cell_list = []

  #iterate through all of the header index's

  for i in header_index:

  #append the cells text content to the cell_list

  cell_list.append(cells[i].textContent)

  #append the cell_list to the return_list

  return_list.append(cell_list)

  #return the return_list

  return return_list

  #This function takes in the source code of the whole page an int list of

  #headers indicating the index number of the needed item and the index number

  #of the table on the page. It returns a list of lists with the scraped info

  def no_header(source, headers, table_index):

  #initiate a list to hold the return list

  return_list = []

  #get a document object out of the source code

  doc = libxml2dom.parseString(source, html=1)

  #get the tables from document

  tables = doc.getElementsByTagName('table')

  try:

  #Try to get focus on the desired table

  main_table = tables[table_index]

  except:

  #if the table doesn't exits then return an error

  return ['The table index was not found']

  #get all of the rows out of the main_table

  rows = main_table.getElementsByTagName('tr')

  #loop through each row

  for row in rows:

  #get all cells from the current row

  cells = row.getElementsByTagName('td')

  #initiate a list to append into the return_list

  cell_list = []

  #loop through the list of desired headers

  for i in headers:

  try:

  #try to add text from the cell into the cell_list

  cell_list.append(cells[i].textContent)

  except:

  #if there is an error usually an index error just continue

  continue

  #append the data scraped into the return_list

  return_list.append(cell_list)

  #return the return list

  return return_list

  希望本文所述对大家的Python程序设计有所帮助。

时间: 2024-08-29 11:23:01

Python实现简单HTML表格解析的方法的相关文章

Python实现简单截取中文字符串的方法

 本文实例讲述了Python实现简单截取中文字符串的方法.分享给大家供大家参考.具体如下: web应用难免会截取字符串的需求,Python中截取英文很容易: ? 1 2 3 >>> s = 'abce' >>> s[0:3] 'abc' 但是截取utf-8的中文机会截取一半导致一些不是乱码的乱码.其实utf8截取很简单,这里记下来作为备忘 ? 1 2 3 4 #-*- coding:utf8 -*- s = u'中文截取' s.decode('utf8')[0:3].e

python实现简单的TCP代理服务器_python

本文实例讲述了python实现简单的TCP代理服务器的方法,分享给大家供大家参考. 具体实现代码如下: # -*- coding: utf-8 -*- ''' filename:rtcp.py @desc: 利用python的socket端口转发,用于远程维护 如果连接不到远程,会sleep 36s,最多尝试200(即两小时) @usage: ./rtcp.py stream1 stream2 stream为:l:port或c:host:port l:port表示监听指定的本地端口 c:host

Python时区设置与获取本地时区方法

Python时区的处理 发现python没有简单的处理时区的方法,不明白为什么Python不提供一个时区模块来处理时区问题. 好在我们有个第三方pytz模块,能够帮我们解决一下时区问题. pytz简单教程 pytz查询某个的时区 可以根据国家代码查找这个国家的所有时区. >>> import pytz >>> pytz.country_timezones('cn') ['Asia/Shanghai', 'Asia/Harbin', 'Asia/Chongqing', '

python自定义解析简单xml格式文件的方法

  这篇文章主要介绍了python自定义解析简单xml格式文件的方法,涉及Python解析XML文件的相关技巧,非常具有实用价值,需要的朋友可以参考下: 因为公司内部的接口返回的字串支持2种形式:php数组,xml;结果php数组python不能直接用,而xml字符串的格式不是标准的,所以也不能用标准模块解析.[不标准的地方是某些节点会的名称是以数字开头的],所以写个简单的脚步来解析一下文件,用来做接口测试. ? 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17

Python实现简单状态框架的方法

 这篇文章主要介绍了Python实现简单状态框架的方法,涉及Python状态框架的实现技巧,具有一定参考借鉴价值,需要的朋友可以参考下     本文实例讲述了Python实现简单状态框架的方法.分享给大家供大家参考.具体分析如下: 这里使用Python实现一个简单的状态框架,代码需要在python3.2环境下运行   代码如下: from time import sleep from random import randint, shuffle class StateMachine(object

python实现将html表格转换成CSV文件的方法

  本文实例讲述了python实现将html表格转换成CSV文件的方法.分享给大家供大家参考.具体如下: 使用方法:python html2csv.py *.html 这段代码使用了 HTMLParser 模块 ? 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 5

python实现简单ftp客户端的方法

  本文实例讲述了python实现简单ftp客户端的方法.分享给大家供大家参考.具体实现方法如下: ? 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 #!/usr/bin/python # -*- coding: utf-8 -*- import ftplib import os

python创建一个最简单http webserver服务器的方法

  这篇文章主要介绍了python创建一个最简单http webserver服务器的方法,实例分析了Python操作http创建服务器端的相关技巧,需要的朋友可以参考下 ? 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 import sys import BaseHTTPServer from SimpleHTTPServer import SimpleHTTPRequestHandler Handler = SimpleHTTPRequestHandler Serve

用Python编写简单的定时器的方法

  这篇文章主要介绍了用Python编写简单的定时器的方法,主要用到了Python中的threading模块,需要的朋友可以参考下 下面介绍以threading模块来实现定时器的方法. 首先介绍一个最简单实现: ? 1 2 3 4 5 6 7 8 9 10 import threading   def say_sth(str): print str t = threading.Timer(2.0, say_sth,[str]) t.start()   if __name__ == '__main