Tuesday, September 13, 2011

Valuable Information

Today I stumbled upon the following twitter post:

Every human intervention in a business process introduces a 4% chance of error. - B. Beims  

Sounds interesting and relevant in the context I'm working in. Than I tried to verify the source and basis for this statement.

  • Using google to search for the statement
  • Using google to search for the author
  • Finding second / third source for this statement
To be honest I wasn't able to verify what I have to verify and therefore  use any bit of this information. Therefore I take this statement as a trigger for this blog post - better than nothing.

That is a example of todays most common topic today:
  • more and more "characters" are accessable and flowing around the world, like "Chinese whispers" posted, re-posted, extended, ....
  • less and less of the accessible" information" (in terms of percent based on the complete available total amount of "information") is relevant or valid
  • Shorten  / context less "information" does not lead to human usable information chain
That is just a fact and reality - everyone has to deal with. To improve your personal ability to make "characters" to "information" you still have to go the hard way:

Don't use and post a information which is not verified by at least
  • a second, independent source
    or
  • personal verification
    or
  • background information which provides you with considerable background to trace the information
If you do not have time for this kind of verification - just leave the "characters" as they are and mark them as irrelevant for you. This should make your personal information chain much cleaner and helps you to divide relevant from irrelevant information.

Don't forget: It is never cheap to gather valuable information. It was never and will never.












Sunday, September 04, 2011

Kevin Slavin: How algorithms shape our world



http://www.ted.com/talks/kevin_slavin_how_algorithms_shape_our_world.html

Writing code in most cases does not mean that you can ever control the usage and implication of the results.....

HTML5 and XML

HTML5 will be the main syntax for the Internet in the next few years and will replace the today most frequently used HTML 4.01. Main driver for this shift was Google and is now adapted by all major browser / OS vendors and organizations.

The main advantage of HTML5 are the new amount of build in features which reflect most of the todays common requirements for web based applications.


Why XHTML 1.0/1.1 failed so far? It mainly was much to strict for the web community - the web and also the world is a non perfect place and therefore HTML5 is much more suitable to fit into this world than the XML based approach of XHTML can provide.

Does this mean XML for the web has loose and does not make sense at all. No there is still space for the XML based standards beside HTML5:
Major advantage of using XML to express the content on the web is a much more easier way to integrate the resulting content into XML processing chains using regular XML transformation tool chains.

  • Easier Reuse content for different channels using XSLT / XQuery
  • Retrieve content as XHTML and extract only dedicated parts (views) required for different use-case
  • Store and request the content using XQuery based infrastructure
The drawback is that today only the newest browser support the mime-type "application/xhtml" therefore for a while Polyglot XHTML might be a good opportunity to deliver the mass and keep processing use-cases doable.

A good summary of Polyglot XHTML and related XML based alternatives for HTML5 can be found here: http://www.xmlplease.com/xhtml/xhtml5polyglot/

Thursday, September 01, 2011

analyse and process DTDs

Working with DTD is still a common task for XML (/SGML) driven use-cases. Knowing this it is very amazing that there is no well know DTD visualization tool available supporting this task.

The good old "Near&Far Designer" is gone many years ago and the source is probably lost in the space of Open Text (the company bought Microstar Software Ltd in 1999). This tool is still in use by many organization having to deal with SGML DTDs (e.g. in the military or aircraft industry).

DTD documentation


There are a few open source scripts out there which converting a DTD into HTML pages for documentation purpose which are available free of charge:
There is one tool out there supporting graphical visualization, documentation and a few function to report key function within the given DTD:

TreeVision (http://www.ovidius.com/meta/download/treevision.html) from German company Ovidius. The tool is available free of charge and provides a very sufficient way to analyze XML / SGML DTDs.


Convert to XML Schema alternatives


If you have to process the content of a DTD for specific use-cases like analyzing the model based on custom specific rules the easiest way is to convert the DTD to RELAX NG (XML syntax) or W3C Schema language. Both are based on XML and therefore can be processed using regular XML based tools.

The best tool to do support this is trang (http://www.thaiopensource.com/relaxng/trang.html source is hosted on http://code.google.com/p/jing-trang/) initially created by James Clark. Compared to commercial alternatives the result is very predictable and for many use cases as good as possible.

DTDs will still exists for many years just because of the many legacy applications created around them. The amount of support is limited but still exists....


Monday, August 01, 2011

lost in email threads....

One of the most time expensive daily tasks is to identify the email required for the the current task in mind.

You know that you already received a email for a particular topic and you want reference it, you require the technical details for a certain topic, ....

Using email tags and full text search of todays email clients is a quite sufficient help to get those kind of tasks done. But once you find a particular email, they is almost ever part of a thread back and forth and getting what you want requires to get the context of the found email. To get the complete context in emails thread isn't trivial, even with thread functions of the common email tools. Most of them are limited in what is shown in particular for long running threads:
  • you loosing the message context around the identified message because the threading function re-arrange the way your inbox is displayed
  • you don't have an easy to use visibility of what really happens, what are the timings for each mail in the task the corresponding sender
  • you do not have easy navigation without loosing the context
Few weeks ago I got aware of ThreadVis a add on for Thunderbird email client.

Pretty cool, it provides a visual graph of the email thread based on the currently selected email with different colors for different sender, length indication for time durations and direct popup help for the content of each item:


You see where you are, what was and after and who was the sender. Even threads you not receive are visible. A easy navigation between the emails, and popup previous for each of the thread items.
Viola, what else could you want? Of course there are things can be improved by the way the basic idea and implementation is worth to take a look at....

Monday, November 01, 2010

Open Source TMS: optentm2 / Open Source Localization framework: okapitools

The open standards in the world of localization becoming more and more mainstream. XLIFF as the main data model to carry localized information between the different automated and manual steps within the localization process, TMX as a format to exchange translation memories, ... the full list of standards can be find here http://www.opentag.com/okapi/wiki/index.php?title=Open_Standards.
This enables more and more open source or free available tools to support certain steps in the localization process without dealing with proprietary and complex formats and conventions.
Few weeks ago the new project "opentm2" released first stable release. Based on IBM TranslationManager this infrastructure claims:
OpenTM2 provides an open platform for managing translation related activities with enterprise level scalability and quality. It serves as an open yet comprehensive localization tool that provides that integration platform. Ultimately, the goal is to create a cost-efficient and high-quality localization deliverable.
Promising and  goes into the same direction as GlobalSight mentioned in a previous post: http://trent-intovalue.blogspot.com/2009/05/open-source-tms-globalsight.html.
In addition infrastructure helps to automate common tasks in localization and dealing with the mentioned standards are available as open source. The project "Okapi framework"  provides a pipeline and several useful steps in localization processes, like
 In addition useful tools to e.g. test text segmentation rules based on SRX and provide a Java based infrastructure which can be embedded in your own application makes this framework something you should look at if you work for / in localization process.

Hopefully the story continues.....

Link: Image Search Tools

Short post for useful list of image search tools posted here "7 Image Search Tools That Will Change Your Life". With the help of such tools you might be able to find images and not only search for it.....

DITA - Beyond OT

DITA-OT (Open Toolkit) is the reference implementation to transform DITA source into various output formats. The reference implementation is open source and maintained by increasing community:

The OT using Apache ANT as pipeline infrastructure. Thats OK in general, but there are several shortcomings once you have to integrate or extend the OT implementation for enterprise use cases (see e.g. discussion "Pipeline refactoring", "Diving Into Performance Improvements").
Because DITA-OT is "only" a reference implementation there might be other implementation out there already using a better approach because they started development once the limitation of the OT implementation already known?
Yes there are two available I'm aware of with different key aspects:
  •  XMLmind DITA Converter (see http://www.xmlmind.com/ditac/what_is_ditac.html)

    Key Aspects: Easier use, integration and improved output.
    Reality: Bad and monolithic design with hard coded java based pipeline and a small user community. No real advantage compared to existing DITA-OT.
  • DITA XProc Pipelines (see https://community.emc.com/docs/DOC-8740)

    Key Aspects: greater flexibility, extensibility, portability, performance.
    Reality:
    Implementation based on XProc (XML Pipeline Language). The design is one step in the right direction and shows the much higher scalability (functional and non functional) of this design, e.g. in case you want to use a extended semantics for validation you can add ISO-Schematron based rules for validation using existing "ISO Schematron schema for DITA" stylesheet and build-in "p:validate-with-schematron" step in XProc and add it into the existing pipeline.

    This implementation isn't perfect, it mainly use XProc markup for the implementation which makes the code long and hard to read / maintain. The usage of the right language for each step which is one advantage of XML pipelining isn't consequent implemented in this implementation. Based on currently available XProc Engines (http://tests.xproc.org/results/) this implementation is not yet ready for enterprise but this will change in the near future the more real world examples are available and used in production.
The XProc based approach is promising and maybe in the future the DITA-OT development also switch over to real XML pipelining. This makes any kind of integration and extension much easier than it is today.....

Tuesday, September 21, 2010

Usage of Wolfram|Alpha

Wolfram|Alpha started round about 1,5 years ago with big noise in the IT related press. Their goal to make "systematic knowledge immediately computable and accessible to everyone" (see http://www.wolframalpha.com/about.html) was promising from the beginning.

Up to now they make huge progress and their engine is a sufficient help for everyone using a computer on a daily basis.

Each day you have to answer question requires specific knowledge or algorithm. In many cases you can find specific Websites helping to answer those questions. But having one source, answer those questions, combine the results for as many domains as Wolfram|Alpha do is unique on the Net (AFAIK).

Samples

Once you want to use or share a certain query you can create a so called widget and bookmark it, embed it into your website, .....

creating a wizard to convert any given unix timestamp to a given named timezone (specific query shown above) i created the following widget (it takes me 10 min. of my limited time ;-)



Limitation

There are still lots of limitations, e.g. for many knowledge domains only US based sources are available. The combination of knowledge is still limited, a fancy query, like "weather the day Jimmy Hendrix died" does not work for many scenarios right now. But they getting better each day....

The end of power point....

My previous post (see http://trent-intovalue.blogspot.com/2010/09/end-of-documents.html) shows that a linear sequence of information is not a perfect match for the expression of complex technical content.
That is also true for presentations which are often supported by power point slides. If you already tried to introduce a complex topic and interact with the audience you might now that the right information is always a slide away.

Why all office application re-create power point instead of thinking of more flexible concepts? I'm not really sure. Few days ago i stumbled upon  Prezi which tries to goes a different route. The content is organized in a tree of topics and navigation is much more focused on context / audience and concrete situation.

Sample:

The end of documents....

Assumption

If you have to describe any technical subject you are aware that knowledge is hard to express in a straight linear sequence of information topics. In most cases you have to structure the information into information tropics and semantic connection between those => a network of information topics.

Nothing new and a trivial statement, you might say. All semantic concepts are using those principle buildup.

Yes, but why most technical subjects are still using documents or slide shows to express technical subjects?
Those media formats are linear by design. The reader has to follow the one and only linear flow defined by the author of the document. In the best case the author is able to find one of the sufficient linear paths through the network of information and the reader is therefore able to understand the described subject. But even in this case getting the hole picture, identify ways to extend the provided information, embed it to different subject etc. isn't possible or at least requires to re-construct the information tree in mind.

Ask yourself why you are using
  • A word processing software to define project information (requirements, design specification, test specification, ....)
  • Power Point to introduce a particular problem domain
  • A word processing software to trace a result of a workshop (also known as workshop protocol)
  • ....
The answer is simple. Because we simple get used to and there are no mainstream alternative media formats out there which can be used without at least one significant constrain (effort to implement and train, difficult to share, ....). The complete office suites still remains rooted in the old linear concepts. Even new players in this business adapting this paradigm (e.g. Google Docs).

Alternatives?

I'm pretty sure that in the future documents will be replaced with applications which providing a way to describe topics as short topics and makes it easy to connect those topics with semantic links (e.g. depends on, contains, .....). A document in this scenario is just one path through the network of information for one particular use case.  This kind of application can replace todays word processing software without loosing any important feature.

In the domain of technical writing topic based authoring (e.g. using DITA information architecture) becomes more and more popular. The main use case there is to (re-)use information as much as possible to reduce creation and information maintenance costs. In my point of view that is "only" a important side effect of having the content defined in a much more usable form. Not a linear sequence of information reflecting a group of authors view but as more or less complete set of information topics linked together. The todays results are still linear documents of some format (pdf, online help formats, ....) but that only corroborate the belief that linear documents are mainstream.

Open Issue


The usability of topic based authoring isn't sufficient today. It is more or less a hand crafted creation of enriched information. To dive into the mainstream usability is the most important factor. The creation and linking of information must be at least as easy as using e.g. Mind Mapping tools (e.g. FreeMind, MindManager) combined with easy to use structured topic content editor (e.g. tools like Xopus or XMAX goes into this direction).



There’s more to come? Lets see....

Tuesday, August 17, 2010

Open Source CCMS

stumbled upon "Calenco XML CMS" a open source (AGPL) CCMS (for further infos about the different kind of available CMS domains, see my post: http://trent-intovalue.blogspot.com/2010/03/stumbled-upon-microsoft-sharepoint-cms.html).
From the list of features published on their website it looks promising. It seems to me the first CCMS application available as open source. I not verified the application so far but I definitely will. Basic features of course but in case such a application get a strong user and developer community the business case for the small CCMS vendors might be tricky. Will see...
The now implemented support for DITA 1.1 and Docbook 4 and 5. The amount of covered feature is also a matter for the verification.

Wednesday, August 11, 2010

usage of open source

Working in  software development domain focused on XML processing you have many major and stable open source infrastructure to use.

But how do you decide to use a certain project for your specific needs? you have to consider the following basic questions:
  • what is the goal of the project and does it match to the goal of your usage?
    If you are not able to answer the question or if both goals doesn't match the roadmap is likely to mismatch => you cannot participate on improvements and in same cases are no more able to upgrade to newer version.
  • how big is the user community?
    the more people using the project the more use-cases are implemented and considered to be stable
  • Is your use case similar to a significant group of the user community?
    same rational as in point one and two
  • How many active developers working on the project?
    The activity multiplied with the amount of developers divided by the size of the project gives you an impression how mature the current code is and how sufficient development will be

    note: big is no value for itself. a small but focused developer community sometimes is a better choice as long as their interest for the project stays stable
  • How active is the development?
    Same rational as above
  • Does the license fits to my use-case or company policy?
  • And of course the most important one: does the project fits to my system requirements?
How to get those information? You can investigate the projects website, taking with the project community, the developers. And you might use Ohloh (http://www.ohloh.net/) which is a community for open source developers provides information around open source projects, the code and the community. Great source for open source....

XQuery Design Patterns

Nice summary of XQuery Design Patterns (see http://patterns.28msec.com/). This side also contains a link to a online XQuery engine based on Zorba (see http://try.zorba-xquery.com/).

The real power of XQuery comes with a database hosting the data for dynamic retrieval and processing. Today there are several databases out there supporting XQuery processing, even as open source.

Open Source
Commercial
  • The big 3 relational db vendors (MS, Oracle, IBM) provide XQuery support. This approach combines relational storage of xml fragments and is therefore only suitable in certain use-cases
  • MarkLogic (http://www.marklogic.com/)
    the most powerful XML DB currently available in my point of view.
    no online sandbox available, but a free developer edition can be downloaded from http://developer.marklogic.com/products
  • ....

Monday, May 17, 2010

Why, What and How

If you want to define a RFP (Request for Quotation) / RFP (Request for Proposal) or you have to answer those requests you are forced to define or understand requirements for a particular product or solution.

There are lot's of good sources showing how to write and handle requirements. But often they miss one important fact which makes either the creation, prioritization and understanding of requirements much easier. The Why instead of the What.

Rational
  1. First set the baseline for effectiveness. (Why)
  2. Than define what you need to be effective. (What)
  3. And at the end define the efficiency. (How)

What is the rational, the reason for a particular requirement. If you try to understand the Why or if you forced to define the Why the resulting requirements / usage of the requirement are much more valid than without doing this.

The "What" does not tell anyone the real intention of a solution. The why provides the motivation and business value and therefore the baseline for effectiveness which makes or makes not a requirement right to exists.

ToDo
  1. Before you start to define any requirement try to define the major "Whys"  you intent to solve. Keep in mind that there must no requirement which cannot be derived from those high level whys.
  2. derive each use-case / requirement from one of the major why areas and add a specific rational for the individual use-case / requirement.
  3. classify / prioritize the requirements based on the rational.
  4. Derive the "How". In this step the statement of efficiency is the key for selection of the right solution.
Result
  • The defined requirements are much easier to understand and alternatives can be much easier identified and qualified.
  • Prioritization can be much easier coordinated to management because the consequence of a decision for the "Why" and therefore the company / department goals are always clear even if someone does not have detailed knowledge of the subject matter.

    =>ensure effectiveness
  • Definition of the system requirements / How can focused on efficiency.
    =>ensure efficiency
To see the same idea from sales / marketing point of view watch: Start with Why: How Great Leaders Inspire Everyone to Take Action.

Sunday, May 16, 2010

facebook - openbook

if you have a facebook account or knowing people having account, try http://youropenbook.org/. I'm always amazed about the data people commit to companies like facebook.

try it out as long as it is available.

XSLT Version 2.1

XSLT Version 2.1 Draft published. As mentioned in the published draft the main goal is making streaming of transformation easier for implementors.

Streaming makes life easier whenever you have to process huge input streams of information to reduce memory footprint and deliver already processed chunks to the next consumer before the complete stream is successful processed.

On the other hand XSLT 2.0 which is a recommendation since 3 years is not widely adopted from implementors of XSLT engines -- Saxon, AltovaXML and few commercial ones-- so far.

Why? The main reason is complexity compared to the corresponding benefit. The implementation of a XSLT engine passes the 2.0 conformance tests is expensive and complex compared to the benefit for most users / use-cases. As in any commercial product each new feature / requirement should be qualified against the value and cost / effort it brings into the standard.

The interesting question would be -- how to measure benefits of a feature? How to interview the "users" of a standard? There is no trivial answer and that might be the main reason why standards tend to get over complex.

Monday, April 19, 2010

CMIS beta implementations

good overview of existing CMIS implementations (client and server) can be found here:

http://www-10.lotus.com/ldd/lqwiki.nsf/dx/11122008094143AMWEBK95.htm

because CMIS version 1.0 is not final approved yet (see also http://www.oasis-open.org/committees/tc_home.php?wg_abbrev=cmis) all of those implementations can be marked as "beta".  Hopefully the progress of the standard will end up in a state different to what happens to WebDAV where the initial hype never made substantial progress in major and public available products and infrastructure components. 

time will tell.....

Monday, March 08, 2010

CMS of what kind?

stumbled upon "Microsoft SharePoint: The CMS Killer" blog post. A personal view how MS SharePoint fits for ECMS use cases.
The results might be wrong or write depending on what your understanding of CMS is.

Problem

There is no common understanding of the the term CMS and even not for the derived term ECMS.

Cause

CMS means Content Management System. Based on this definition it is a application (system) to manage content. Thats trivial but what does content really means? Content is all and everything. Most of the content can be managed within IT systems as well.

Illustration

General categories of content maintained by IT applications:
  • Structured content
    Content maintained and structured as a collection of data (order or offer data). Content in this context means a collection of records with given structure
    Those type of data are typical maintained with applications called ERP or other kind of systems of this type (e.g. ALM systems).
  • Unstructured content
    Content maintained within documents. The content within the document is not addressable outside the application the document was created with (the semantic makes only sense in one specific usage scenario).
    Those type of data are typical maintained with applications called DMS or Web-CMS (maintaining HTML) systems.
The propagation of xml introduce a third categorie of content:
  • Semi-structured content
    Content used to create documents (information products) of some kind. The content within the document is addressable outside the application the document was created with.
    OR
    Content maintained as system independent instance of data (e.g. order or offer data).
    In both cases XML is the common format today.
Each CMS has a history in one of the mentioned areas and has its specific strength and weakness for usage in different domains.

The corresponding consultants knowing the mentioned issue and trying to create domain specific names for specific usage of systems
  • DBMS
    Main goal is to maintain relational data and used by dedicated applications on top.
  • DMS
    Main goal is to maintain documents
  • Web-CMS
    Main goal is to maintain intranet / extranet / internet sites
  • ECMS
    Enterprise-CMS
    Main goal is to provide enterprise ready workflow and records management on top of DMS feature set
  • CCMS
    Component-CMS
    Main goal is to maintain content stored in XML for single-source publishing
  • ?
    anything i missed, of course there are plenty of buzzwords / domains out there describing mixed-scenario usage.
Because non of the mentioned definitions are "formal approved" the specific vendors use the term with most customer attention in their target domain and each vendor tries to fulfill use-cases from other domains as well. Because each CMS has a particular implementation history the usage in most cases is limited even if the vendor claims to cover most of them.

Examples

  • a C-CMS vendor might support xml usage and publishing very well but does not scale if enterprise workflow or records management is required.
  • a ECMS vendor has enterprise BPM support build in but lacks of sophisticated xml semantic and functions.
  • a Web-CMS makes creating your Internet presence easy but lacks the usage of the same content for printed documents
  • .....
Summary

This reflects the current state of content management. I expect in the next few years that new and maybe existing systems will move into the semi-structured content area and some of them might succeed. They might reach the final goal that content can be create, maintain, re-purpose and publish based on different user communities from one single source.

Until that stage is reached....coming back to initial "Microsoft SharePoint: The CMS Killer" statement. Ask the author what kind of content use-case he has in mind and you can validate the statement.

Sunday, March 07, 2010

we live for our customers....

....really?

do you work with customer centric mindset?
do you work in a customer centric organization?

your first and quick answer might be yes, of course. my company does everything to make the customer happy, thats where our revenue comes from. of course.....

That was my first thought as well.

Than i looked at several websites incl. my own company one. you find service descriptions, white papers, success stories, awards -- everything focus to promote the own portfolio of services / products.

Have a look at your company's website. Looks it different?

To understand the customer means to understand their challenges, needs and questions they have. You have to answer the questions: what is the business value you can bring to your customer instead of leave them alone to select  a specific product or service you offer.

What is the reason for that?

It is much harder to define and understand the problems your customer have and how you can provide business value for them using their terms and definition. Instead showing what you did and do and let the potential customer decide if that might help them is much easier.

By the way you will be much more successful if you at least try to think and follow the problems your potential customer have to solve....