Friday, 27 March 2015

Nine Princes in Amber, by Roger Zelazny (Amber series book I)

Book cover Nine Princes in Amber describes a world in which only one place is "substance", the great city of Amber, while all others, including Earth, are just "shadow". Roger Zelazny doesn't really explain what is the difference, but we catch glimpses of the way Amber people can alter reality and that some of the things that happen in shadow cannot happen in Amber. Gun powder, for example, or perhaps even electricity. Great power can be channeled through tarot cards that hold the images of the children of Oberon or some other things, former king of Amber, while his numerous sons and daughters scheme and plan to take over the throne. The cards are drawn by a weird magician that seems to hold no political power, despite his amazing skill.

I took advantage of a flu that restricted me to bed in order to listen to the book in audio format, narrated by Roger Zelazny himself. While I enjoyed it, I didn't feel it taught me anything new. It was a pretty inconsistent story, as well. Princes and princesses and nobles, acting all smarty and aristocratic, dueling with swords and the occasional bow or crossbow, and some of the classic cliches encountered in this sort of context: like never killing your noble opponents, but imprisoning them or torturing/displaying them in order to demonstrate power.

In the end, the main protagonist remains a mystery. He starts with an amnesia, but soon he recovers his memory, making the entire memory loss kind of pointless. It is never quite spelled out who he is as a person, or what was it that he did on Earth. In truth, the book feels more like an introduction to the world of Amber, more like a teaser really, leaving the exposition of real character or description of the worlds to the following books. It is one of the first books from Zelazny, so maybe it will become better in the future. It is not explained why someone would want to rule Amber, either, as any other shadow world looks more appealing.

Even if I haven't fully enjoyed the story, I did promise a friend I would finish the entire Amber series, so we will see. After finishing it, I reserve the right to torture my friend in order to display my power by making him read something truly awesome and unsettling. I have not yet determined what yet, but a dark bird of my desire will carry my message and he will live in fear.

Zelazny himself died in 1995. I found a page written by George R. R. Martin, a beautiful remembrance of a mentor and friend. Apparently, he wrote the character Croyd (The Sleeper) in the Wild Cards series, one of the most interesting characters and appearing in many of the stories, even if very rarely as a main protagonist.

Tuesday, 24 March 2015

Blog Question: How to use a large image as a page background, without showing it loading slowly?

A blog reader asked me to help him get rid of the ugly effect of a large background image getting loaded. I thought of several solutions, all more complicated than the rest, but in the end settled on one that seems to be working well and doesn't require complicated libraries or difficult implementation: using the img onload event.

Let's assume that the background image is on the body element of the page. The solution involves setting a style on the body to hide it (style="display:none") then adding as child of the body an image that also is hidden and that, when completing loading, shows the body element. Here is the initial code:
<style>
body {
background: url(bg.jpg) no-repeat center center fixed;
}
</style>
<body>

And after:

<style>
body {
background: url(bg.jpg) no-repeat center center fixed;
}
</style>
<body style="display:none">
<img src="bg.jpg" onload="document.body.style.display=''" style="display:none;" />

This loads the image in a hidden img element and shows the body element when the image finished loading.

The solution might have some problems with Internet Explorer 9, as it seems the load event is not fired for images retrieved from the cache. In that case, a slightly more complex Javascript solution is needed as detailed in this blog post: How to Fix the IE9 Image Onload Bug. Also, in Internet Explorer 5-7 the load event fires for animated GIFs at every loop. I am sure you know it's a bad idea to have an animated GIF as a page background, though :)

Warning: While this hides the effect of slow loading background images, it also hides the page until the image is loaded. This makes the page appear blank until then. More complex solutions would show some simple html content while the page is loading rather than hiding the entire page, but this post is about the simplest solution for the question asked.

A more comprehensive analysis of image preloading, complete with a very nice Javascript code that covers a lot of cases, can be found at Preloading images using javascript, the right way and without frameworks

Monday, 23 March 2015

Goodbye SiteMeter, too!

In my opinion, when a software you have been using for a long time changes the way it works and intrudes on your already existing installations, not only it is disappointing and mean, but it should also be illegal. Today I noticed that the links from my blog went to intermediate sites (I apologize for not noticing it sooner) like vindicosuite. A quick Google search led me to this link: Goodbye Sitemeter. Apparently, SiteMeter, a software that I have been using to show a views counter on my blog, has been acquired by a crappy company called News Company. I mean, this is the actual name, I am not making fun of you; it's like displaying "Dr Doom's Evil Lair" on your house fence (and not kidding about it). Without the company saying anything, the SiteMeter script added these click and contextmenu handlers on my links, redirecting to other sites, maybe with ads on them (I have AdBlock Plus installed and so should you!, so I don't know). Anyway, the moment I realized this I removed the script from my blog. I have to apologize again for failing to notice this for so long.

Sunday, 22 March 2015

Parasyte - the maxim, an interesting anime

The main hero and his alien hand OK, I have no idea what most Japanese titles want to say. Is this about a parasite who is also a short, pithy statement expressing a general truth or rule of conduct? No, it is not. Parasyte is about a guy who gets infected by an alien metamorph, but somehow he manages to contain the infected area to his right arm. As a result, he maintains his personality, but now has a powerful alien as his right arm. It can change shape, it is very intelligent and it is generally useful when dealing with other afflicted, who usually have their brain infested, and thus are alien in their entirety.

Of course, being a Japanese anime, our hero is a highschool male student after which a number of girls are pining for no good reason and that he has to fight to protect. No scenes of using his versatile tentacle arm on these girls, though. There are also some discussions about the role of these parasites and/or humans in the world, a vague ecologist propaganda that really has nothing to do with the plot and lots and lots of gore. The interaction between the human highschooler and the amoral and fiercely individualistic alien makes for most of the fun in the anime.

The series is ongoing, but I just watched the first 23 episodes and I can safely say that they could have stopped there. Probably they can come with fresh ideas, but for me the story started and ended satisfactorily with episode 23. The animation is good, but nothing spectacular, the Japanese cliches are abundant, but only barely overused and the main character is someone you can easily like and understand.

As far as I can see the anime faithfully follows the manga and episode 23 ends where the manga chapter 62 ends. There are just two other manga chapters published, so the anime and mange are pretty much synchronized. You can read the Parasyte manga online.

Saturday, 21 March 2015

Blog question: Address validation using CRF models

I am starting a new blog series called Blog Question, due to the successful incorporation of a blog chat that works, is free, and does exactly what it should do: Chatango. All except letting me know in real time when a question has been posted on the chat :( . Because of that, many times I don't realize that someone is asking me things and so I fail to answer. As a solution, I will try to answer questions in blog posts, after I do my research. The new label associated with these posts is 'question'.

First off, some assumptions. I will assume that the person who said I'm working on this project of address validation. Using crf models is my concern. was talking about Conditional Random Fields and he meant postal addresses. If you are reading this, please let me know if that is correct. Also, since I am .NET developer, I will use concepts related to .NET.

I knew nothing about CRFs before writing this posts, so bear with me. The Wikipedia article about them is hard to understand by anyone without mathematical (specifically probabilities and statistics) training. However the first paragraph is pretty clear: Conditional random fields (CRFs) are a class of statistical modelling method often applied in pattern recognition and machine learning, where they are used for structured prediction. Whereas an ordinary classifier predicts a label for a single sample without regard to "neighboring" samples, a CRF can take context into account. It involves a process that classifies data by taking into account neighboring samples.

A blog post that clarified the concept much better was Introduction to Conditional Random Fields. It describes how one uses so called feature functions to extract a score from a data sample, then aggregates scores using weights. It also explains how those weights can be automatically computed (machine learning).

In the context of postal address parsing, one would create an interface for feature functions, implement a few of them based on domain specific knowledge, like "if it's an English or American address, the word before St. is a street name", then compute the weighting of the features by training the system using a manually tagged series of addresses. I guess the feature functions can ignore the neighboring words and also do stuff like "If this regular expression matches the address, then this fragment is a street name".

I find this concept really interesting (thanks for pointing it out to me) since it successfully combines feature extraction as defined by an expert and machine learning. Actually, the expert part is not that relevant, since the automated weighing will just give a score close to 0 to all stupid or buggy feature functions.

Of course, one doesn't need to do it from scratch, other people have done it in the past. One blog post that discusses this and also uses more probabilistic methods specifically to postal addresses can be found here: Probabilistic Postal Address Elementalization. From Hidden Markov Models, Maximum-Entropy Markov Models, Transformation-Based Learning and Conditional Random Fields, she found that the Maximum-Entropy Markov model and the Conditional Random Field taggers consistently had the highest overall accuracy of the group. Both consistently had accuracies over 98%, even on partial addresses. The code for this is public at GitHub, but it's written in Java.

When looking around for this post, I found a lot of references to a specific software called the Stanford Named Entity Recognizer, also written in Java, but which has a .NET port. I haven't used the software, but it seems as it is a very thorough implementation of a Named Entity Recognizer. Named Entity Recognition (NER) labels sequences of words in a text which are the names of things, such as person and company names, or gene and protein names. It comes with well-engineered feature extractors for Named Entity Recognition, and many options for defining feature extractors. Included with the download are good named entity recognizers for English, particularly for the 3 classes (PERSON, ORGANIZATION, LOCATION). Perhaps this would also come in handy.

This is as far as I am willing to go without discussing existing code or actually writing some. For more details, contact me and we can work on it.

More random stuff:
The primary advantage of CRFs over hidden Markov models is their conditional nature, resulting in the relaxation of the independence assumptions required by HMMs in order to ensure tractable inference. Additionally, CRFs avoid the label bias problem, a weakness exhibited by maximum entropy Markov models (MEMMs) and other conditional Markov models based on directed graphical models. CRFs outperform both MEMMs and HMMs on a number of real-world sequence labeling tasks. - from Conditional Random Fields: An Introduction

Tutorial on Conditional Random Fields for Sequence Prediction

CRFsuite - Documentation

Extracting named entities in C# using the Stanford NLP Parser

Tutorial: Conditional Random Field (CRF)

An appeal for generational concepts in Wikipedia

I often find a new thing that I haven't ever heard of, so I google it. A lot of the time, the first link returned is the Wikipedia article about that concept and I open it to get a general idea of what it is about. Most of the time I understand it immediately, but in some cases - mostly involving hard science like high level mathematics - that page is just a bunch of gibberish that means less to me than what I was looking to clarify in the first place. I mean, when I am searching for something, I usually use words, so there: much clearer.

However, that doesn't mean that I don't want to understand what is described on that page. One idea I had is that of "generational concepts", in other words the concepts that one needs to understand before tackling a new one. They are not "related concepts", they are not links to terms used in the description, they are the general concepts that you need to get first. I find it interesting and useful for several reasons:
  • I could open the links to those concepts and, if I understand them, I could come back and get the one that I wanted
  • If I don't understand the basic concepts, they would also have generational concepts to investigate
  • No one actually needs to create an entire chain of pages, like a teacher in a class, but just edit and existing page and link to the base concepts for it, yet the result is like a course that one can follow up and down
  • It would add context (and thus interest) to Wikipedia, which is now used as a collection of disparate tidbits
  • It would answer the question that I always ask myself when I open an incomprehensible page: what need I know in order to understand this crap?

So now I should put some time aside for fixing Wikipedia.

Thursday, 19 March 2015

Change all column default values with a certain value to another value in T-SQL

Just to remember this for future work. I wanted to replace GetDate() default column values with SysUtcDatetime(). This is the script used:
-- declare a string that will hold the actual SQL executed
DECLARE @SQL NVARCHAR(Max) = ''
SELECT @SQL=@SQL+
N'ALTER TABLE ['+t.name+'] DROP CONSTRAINT ['+o.name+'];
ALTER TABLE ['
+t.name+'] ADD DEFAULT SYSUTCDATETIME() FOR ['+c.name+'];
'
-- drop the default value constraint, then add another with SYSUTCDATETIME() as default value
FROM sys.all_columns c -- get the name of the columns
INNER JOIN sys.tables t -- get the name of the tables containing the columns
ON c.object_id=t.object_id
INNER JOIN sys.default_constraints o -- we are only interested in default value constraints
ON c.default_object_id=o.object_id
WHERE o.definition='(getdate())' -- only interested in the columns with getdate() as default value

-- execute generated SQL
EXEC sp_executesql @SQL