ย้ายบล็อกไปที่ bact.cc แล้วนะครับ

พ.ร.บ.คอมพิวเตอร์
หยุด ร่างพ.ร.บ.คอมพิวเตอร์
พื้นที่เก็บข้อมูลออนไลน์ ฟรี 2GB จาก Dropbox (sync กับ Windows, Linux, Mac, iPhone, Android ฯลฯ ได้)
Showing posts with label XML. Show all posts
Showing posts with label XML. Show all posts

2007-10-03

YAiTRON/LEXiTRON "ancient" word

ถามวีร์และคนอื่น ๆ

ตอนนี้พยายามทำความสะอาด YAiTRON (LEXiTRON ฉบับ XML) เพิ่มเติมอยู่

ผมลองสั่งหาคำว่า “(คำโบราณ)” ในอีลีเมนต์ translation-similar ใน YAiTRON เจองี้ (ถ้าหาในอีลีเมนต์ translation เฉย ๆ จะเจอเยอะกว่านี้)

$ cat yaitron.xml | grep -n -e \<translation-similar.*\(คำโบ
832513:  <translation-similar lang="tha">ชาวเรือ, กะลาสี (คำโบราณ)</translation-similar>
947329:  <translation-similar lang="tha">(คำโบราณ) (ทางวรรณคดี)</translation-similar>
952697:  <translation-similar lang="tha">(คำโบราณ หรือทางวรรณคดี)</translation-similar>

อยากรู้ว่า เจ้าหมายเหตุว่า “(คำโบราณ)” หรืออะไรประมาณนี้เนี่ย มันควรจะไปเก็บอยู่ตรงไหนดีครับ ที่ตัวอีลีเมนต์ entry (ซึ่งเป็นอีลีเมนต์แม่ของ translation-similar) หรือว่าจะไปเก็บเป็นแอตทริบิวต์ของอีลีเมนต์ translation-similar หรือว่าอย่างอื่น ??
(ในสเปคปัจจุบันของ YAiTRON แนะนำให้เก็บหมายเหตุพวกนี้ลงอีลีเมนต์ชื่อ note แต่อยากให้มัน machine-readable อ่ะ)

เช่น ปัจจุบันเป็นงี้

<entry lang="eng">
 <headword>seafarer</headword>
 <translation lang="tha">คนเดินเรือ</translation>
 <translation-similar lang="tha">ชาวเรือ, กะลาสี (คำโบราณ)</translation-similar>
 <lexitron id="63739"/>
</entry>

ที่เสนอแบบที่ 1 (“คำโบราณ” บ่งชี้ “seafarer”):

<entry lang="eng" >
 <ancient>true</ancient>
 <headword>seafarer</headword>
 <translation lang="tha">คนเดินเรือ</translation>
 <translation-similar lang="tha">ชาวเรือ</translation-similar>
 <translation-similar lang="tha">กะลาสี</translation-similar>
 <lexitron id="63739"/>
</entry>

ที่เสนอแบบที่ 2.1 (“คำโบราณ” บ่งชี้ทุกคำใน translation-similar):

<entry lang="eng">
 <headword>seafarer</headword>
 <translation lang="tha">คนเดินเรือ</translation>
 <translation-similar lang="tha" ancient="true">ชาวเรือ</translation-similar>
 <translation-similar lang="tha" ancient="true">กะลาสี</translation-similar>
 <lexitron id="63739"/>
</entry>

แบบที่ 2.2 (“คำโบราณ” บ่งชี้เฉพาะ “กะลาสี”):

<entry lang="eng">
 <headword>seafarer</headword>
 <translation lang="tha">คนเดินเรือ</translation>
 <translation-similar lang="tha">ชาวเรือ</translation-similar>
 <translation-similar lang="tha" ancient="true">กะลาสี</translation-similar>
 <lexitron id="63739"/>
</entry>

คือไม่แน่ใจว่า “คำโบราณ” ในหมายเหตุ (ในวงเล็บ) เนี่ย มันเป็นตัวบ่งชี้อะไร
ตัวบ่งชี้ คำศัพท์ (seafarer) หรือว่าตัวบ่งชี้คำแปล (ชาวเรือ, กะลาสี)

ใครมีความเห็นไรมั่งครับ ?

อีกอันที่น่าสนใจ เอาไว้พิจารณาประกอบก็คือ มันมี entry แบบนี้ด้วยอันนึง:

<entry lang="eng">
 <pos>N</pos>
 <headword>weeds</headword>
 <translation lang="tha">เสื้อผ้าสีดำซึ่งเดิมเป็นชุดสวมใส่ของแม่ม่าย</translation>
 <translation-similar lang="tha">(คำโบราณ หรือทางวรรณคดี)</translation-similar>
 <lexitron id="80641"/>
</entry>

จะเห็นว่าใน translation-similar ไม่มีคำแปลอะไรอยู่เลย มีแต่หมายเหตุ แบบนี้ แปลว่าหมายเหตุใน translation-similar ไม่ได้บ่งชี้ตัวคำแปลใน translation-similar .. แต่บ่งชี้คำศัพท์ (weeds) น่ะสิ ?? คิดแบบนี้ได้ไหม ?

หรือ ... แต่เนื่องจากเจอแบบนี้แค่อันเดียว ก็อาจจะถือว่ามันเป็นข้อผิดพลาด ไม่ต้องสนใจ จะได้ไหม?

นอกเรื่อง: ข้อมูล LEXiTRON ที่ให้ดาวน์โหลดได้* ซึ่งเอามาทำ YAiTRON (โดยวีร์ ผ่านทางทางคุณพูนลาภอีกที) มันเป็นรุ่นเมื่อหลายปีก่อน คือ 2.0 (หรือก่อนหน้านั้น) ส่วนรุ่นล่าสุดที่ใช้บนเว็บ คือ 2.2 เค้าไม่มีให้ดาวน์โหลด

Dictionary Thai→English English→Thai
WordsSenses WordsSenses
LEXiTRON 2.135,000?53,000?
LEXiTRON 2.251,000?79,000?
YAiTRON32,35040,85453,53483233

* ต้องสมัครสมาชิกก่อนถึงจะดาวน์โหลดได้ แต่หน้าเว็บสำหรับสมัครก็ดันบอก PHP error อีเมลไปตามที่อยู่ที่แจ้งไว้ ก็ตีกลับ ... - -"

technorati tags: , , ,

2007-09-26

YAiTRON XSLT stylesheets

YAiTRON is a cleaned-up version of NECTEC's LEXiTRON in a well-formed XML format, created by Vee Satayamas. Its tag names are TEI-inspired.

technorati tags: , , , ,

2007-08-30

OOXML Advertorial -- NoOOXML

OOXML ทำเนียน

วันนี้เจอโฆษณา “Open XML” ใน ฐานเศรษฐกิจ ฉบับวันที่ 30 ส.ค. - 1 ก.ย. 2550 หน้า 34 (เซคชั่น "ตลาด-ตลาดภูมิภาค")

หน้าตาทำเหมือนเป็นบทความ ขึ้นหัวใหญ่ว่า

“ธุรกิจไทย คนไทย มีทางเลือกหรือไม่ในเวทีระดับโลก
ประเทศไทยควรโหวตรับมาตรฐานการจัดเก็บเอกสารใหม่หรือไม่...”

ในนั้นมียกคำพูดจากบุคคลในวงการไอทีต่าง ๆ เช่นจาก คุณฟูเกียรติ จุลนวล ผู้จัดการฝ่ายกลยุทธ์และแพลตฟอร์ม บริษัท ไมโครซอฟท์ (ประเทศไทย) จำกัด คุณสมเกียรติ อึ้งอารี ประธานกรรมการบริหาร บริษัท ซีเนียร์ คอม จำกัด นายกสมาคมอุตสาหกรรมซอฟต์แวร์ไทย (ATSI) คุณสุวิภา วรรณสาธพ ผู้อำนวยการเขตอุตสาหกรรมซอฟต์แวร์ประเทศไทย (ซอฟต์แวร์พาร์ค)

ตรงกลาง ๆ “บทความ” ตอนหนึ่งเขียนว่า

“ที่สำคัญ การโหวตครั้งนี้เป็นการทำให้ภาษาไทยได้เข้าไปเป็นหนึ่งในมาตรฐานโลก ซึ่งหากต่อไปจะมีการพัฒนาแอพพลิเคชั่นอะไรขึ้นมาแล้ว ภาษาไทยก็จะเป็นหนึ่งภาษาที่ถูกนำไปพิจารณาด้วย แม้ว่าจะมีผู้ใช้เฉพาะในประเทศไทยเท่านั้น”

... จริงหรือไม่ครับ ?
(แต่เป็นการใช้ภาษาที่ดูดีทีเดียว เขียนแบบให้ความหวังมากในที่แรก จะมีภาษาไทยแน่ ๆ ... แต่ในตอนสุดท้ายก็ทิ้งระยะความรับผิดชอบแบบนิ่ม ๆ .. จะถูกนำไปพิจารณาเท่านั้นแหละนะ ไม่ได้สัญญาอะไรมากกว่านี้)

อีกตอนหนึ่งเขียนว่า

“วันนี้เรากำลังมีทางเลือกในการที่จะมีอีกมาตรฐานที่ช่วยในการจัดเก็บเอกสารไว้ใช้งาน ประเทศไทยอาจจะไม่จำเป็นต้องรับรองให้ Open XML เป็นมาตรฐาน ISO ก็ได้ แต่ถามว่า วันนี้เรามีมาตรฐานที่เหมาะสมที่ช่วยในการจัดกับเอกสารที่เรามีใช้งานอยู่แล้วในองค์กร รวมถึงมาตรฐานที่ช่วยให้เอกสารของเราสามารถทำงานร่วมกับระบบงานต่าง ๆ ที่มีใช้งานอยู่แล้วในองค์กร”

อ่านโดยรวมทั้งหมดแล้ว จะเน้นกลุ่มองค์กรที่ใช้งานไมโครซอฟท์ออฟฟิศอยู่แล้ว และสร้างความไม่แน่ใจเกิดขึ้นว่า ถ้า Open XML ไม่ได้เป็นมาตรฐาน ISO แล้วเอกสารทั้งหมดของพวกเขา จะทำงานกับระบบอื่น ๆ ในโลกไม่ได้

ก็ติดตามตรวจสอบกันไปครับ ใครพูดจริงเท็จ พูดครึ่งเดียว พูดบิดเบือน ...

แล้ว สมอ. ของไทย จะเชื่อใคร โหวตให้ใคร เพื่อเห็นแก่ประโยชน์ของใคร ... ก็ดูกันไป

พวกเราจะไปมีส่วนร่วมอะไรได้ไหม ??


ลองอ่าน รวมความเห็น OOXML จากหลายฝ่าย ที่ Blognone

ถ้าใครพิจารณาแล้ว ไม่สนับสนุนการมีอีกมาตรฐาน (ตอนนี้ ISO มีมาตรฐานการจัดเก็บเอกสารอยู่แล้ว ชื่อว่า OpenDocument)
ก็ไปลงชื่อคัดค้านกัน ที่ No OOXML (ในนั้นมีเหตุผลให้อ่านเป็นข้อ ๆ เลย ลองอ่านดู)


(เลือกรูปอื่น ๆ ไปแปะเว็บ ได้จาก NoOOXML Banners)

technorati tags: , , ,

2006-09-06

Google n-gram are belong to YOU

กูเกิล แจกโมเดล n-gram ซึ่้งกูเกิลใช้ในงานวิจัยต่าง ๆ เช่น การแปลภาษาอัตโนมัติ การแก้ตัวสะกดอัตโนมัติ การสกัดสารสนเทศ ฯลฯ โดยโมเดลนี้สร้างจากคำมากกว่า 1 ล้านล้านคำ โดยจะแจกจ่ายผ่าน LDC ในรูปของ DVD 6 แผ่น

LDC นี่ เป็นหน่วยงานที่ทำงานด้านข้อมูลภาษาศาสตร์ พวกคลังข้อความ (corpus) ข้อมูลที่แจกจ่ายโดย LDC มีหลายประเภท บางประเภทต้องเป็นสมาชิก (เสียเงินค่าสมาชิกแพงอยู่) จึงจะเรียกดูได้ บางประเภทซื้อแยกต่างหากได้โดยไม่ต้องเป็นสมาชิก บางประเภทก็ฟรี — แต่กรณี DVD 6 แผ่นนี่ ยังไงคงต้องเสียค่าส่งแน่ ๆ

Google Research Blog announced:

... we decided to share this enormous dataset with everyone. We processed 1,011,582,453,213 words of running text and are publishing the counts for all 1,146,580,664 five-word sequences that appear at least 40 times. There are 13,653,070 unique words, after discarding words that appear less than 200 times.

Watch for an announcement at the LDC, who will be distributing it soon, and then order your set of 6 DVDs.

ใครอยากจะลอง เชิญได้เลย! :P

via information retrieval

tags: | | |

2006-07-13

DITA - Darwin Information Typing Architecture

The Darwin Information Typing Architecture () is an XML-based architecture for authoring, producing, and delivering technical information.
Wikipedia

tags:

2006-04-11

Open Source HTML Parsers in Java

Open Source HTML Parsers in Java, a list by Java-Source.net

NekoHTML, HTML Parser, Java HTML Parser, Jericho HTML Parser, JTidy, TagSoup, HotSax

แถม Nux เหมือนจะทำอะไรได้หลายอย่างสารพัดเกี่ยวกับ XML (เป็น wrapper ของตัวอื่น ๆ ด้วย)

2006-04-05

TIGER API 1.8 released

TIGER API is a library which allows Java programmers to easily access the structure of any corpus given as a TIGER-XML file.

oeze, one of the authors of TIGER API, has leave a message to us today:

BTW, Tiger API has moved. This is the new URL: TIGER API.

We have also included a section describing how to access corpora encoded in Penn Treebank format and other formats.

Thanks, oeze ! :)

link: http://tigerapi.org

2006-04-03

REXML Nodes and Elements

REXML, a Ruby-style XML toolkit

What's the difference between results from code (1) and (2) below ?
(element is an XML element)

Code (1), use Element#elements :


element.elements.each do |e|
 puts e.inspect
end

Code (2), use Element#to_a :


element.to_a.each do |e|
 puts e.inspect
end

Update: We can actually use just element.each .. no .to_a requied — thanks to P'Pok for this

Code (2) will give us texts, elements (as well as other nodes).
Where code (1) will give us only elements.

If our input is:


<p><b>bold</b> text</p>

Code (1) will give:


<b> ... </>

While code (2) will give:


<b> ... </>
" text"

This tiny difference already wasted me hours, shamed :(
I was thought that text is a kind of element, ... that's plain wrong, both text and element are kinds of node !

For several REXML tutorials/examples I've found, where I copied and pasted codes from for my quick-n-dirty-self-education, all of them show only the use of Element#elements but not Element#to_a.

This is probably because all of them only deal with a data-oriented XML, where a use of 'mix content' is rare (and indeed not recommended). But that's no longer true for document/text-oriented XML — for example, XHTML.

If you going to process a XML with mix content, beware of #elements.

Correct me if I do anything wrong here.
REXML veteran? Share! ;)


Tutorials: REXML Home | XML.com | developerWorks
API docs

2006-01-31

Eclipse XML plug-ins

ปลั๊กอินของ Eclipse กันอีกแล้ว เยอะจริง ๆ - - (T-T)

ผมลงตัวแรกกะตัว WTP ไม่ได้อ่ะ โห dependency ยุ่บยั่บ ตัวแรก (EclipseXSLT) นั่นจะเอา WTP ส่วน WTP ก็จะเอา JEM, JEM ก็บอกจะเอา RCP (คือไรมั่งฟะ - -) ... ไม่เอาล่ะ เลิก

2005-09-22

A survey of XML standards

World's full of standards. Some of them will go obsoleted when you finished reading the last one, seriously.

Thanks to mk for the link.

2005-08-30

Trang problem with Mustang

If you now trying Mustang (Java SE 6 development snapshots), be warned that it's not compatible with Trang (XML schema converter).

Bug ID: 6301903
REGRESSION: Cannot run Trang - CLASSPATH has no effect
State: In progress, bug

Java SE 5, 1.4 and 1.3 are doing fine.

2005-07-29

Regular expression in Relax NG Compact syntax

ถ้าอยากกำหนดรูปแบบข้อมูลที่จะอนุญาตให้ใส่ใน XML ของเรา ลงไปใน schema language เพื่อที่จะได้ตรวจได้โดย XML validator ใน Relax NG นี่ เราไปยืม pattern จาก W3C XML Schema datatypes มาใช้ได้

โดยตัว pattern นี่ มันก็คือ regular expression น่ะแหละ

ใช้งี้:

xsd:string { pattern = "regular_expression_here" }

เช่น ถ้าเราเขียน:

# ยืม datatypes ของ W3C XML Schema มาใช้
datatypes xsd = "http://www.w3.org/2001/XMLSchema-datatypes"

# กำหนดรูปแบบข้อมูลเอง
Percent.data = ( xsd:string{ pattern = "(100|[1-9]?[0-9])[%]" } )

Style = element font {
  attribute weight { Percent.data }, # เอามาใช้ตรงนี้
  attribute color { "green" | "greener" }, # ถ้าแค่นี้ ไม่ต้องใช้ pattern ก็ได้
  text
}

ตัวอิลิเมนต์ font ก็จะเอาไปใช้ได้ประมาณนี้:

<font color="greener" weight="80%">bact'</font>

นี่ไม่ได้เปิดหนังสือดู อาศัยมั่ว ๆ หาเอาจากเน็ต เลยยังไม่แน่ใจว่า pattern เนี่ย มันใช้กับอย่างอื่นนอกจาก xsd:string ได้รึเปล่า ?

2005-07-28

Co-Occurrence Constraints in RELAX NG

วิธีกำหนดเงื่อนไข &ldquoปรากฏร่วม” ใน RELAX NG: ถ้าค่าของโหนดนี้ เป็นแบบนี้ ตัวข้อมูลที่เหลือจะต้องเป็นแบบนี้ ... เช่น ในอิลิเมนต์ person ถ้าเกิดแอตทริบิวต์ type มีค่าเป็น “author” ตัวอิลิเมนต์ person จะต้องมีอิลิเมนต์ลูก ชื่อ dead ด้วยนะ; แต่ถ้าแอตทริบิวต์ type มีค่าเป็น “character” ตัวอิลิเมนต์ลูกที่ต้องมี ก็จะเปลี่ยนไป

เงื่อนไขลักษณะนี้ ทำด้วย DTD ไม่ได้เลย. ส่วน W3C XML Schema นี่ พอทำได้บ้าง แต่ยุ่งยาก

คำถามก็คือ แล้วทำไมไม่สร้างอิลิเมนต์แยกกันไปเลยล่ะ ? -_-" อย่างในตัวอย่างข้างบน ก็ทำเป็นสองอิลิเมนต์ไปเลย ก็น่าจะได้รึเปล่า ? หรือว่าทำแบบนี้แล้ว มันทำให้ดูทั่ว ๆ ไป (generic) มากขึ้น ?

อันนี้ก็ต้องแล้วแต่คนออกแบบ schema แล้วล่ะ ว่าจะเอาไง. ตัว schema language (อย่าง Relax NG) นั้น เปิดช่องให้ทำได้, แต่จะทำไม่ทำก็แล้วแต่.

2005-05-23

Piccolo SAX Parser

From benchmarks here and here, this Piccolo Java SAX parser performs really, incredibly, fast.

2005-04-15

Choosing an XML editor

Alastair Dunning posted this to the TEI list today :-

AHDS Literature, Languages and Linguistics has recently published a new Information Paper on XML editors.

With a large number of XML editors now available, this Information Paper serves as an introduction to the different features XML editors can have and the extent to which these features are implemented. It also presents the result of an evaluation exercise where different user groups tried a number of the editors.

The article is based on a study by Thijs van den Broek, Benchmarking XML editors, undertaken in 2004. The evaluation includes results of the survey van den Broek undertook via the TEI website.

AHDS Literature, Languages and Linguistics is hosted by the Oxford Text Archive.

PDF to XML

Papers

Tools

  • pdftohtml — PDF to HTML/XML conversion utility