Friday, June 12, 2009

I'm GLAD I'm not programming in natural human language

At least one of the hard and undeniable facts of software development doesn't really swipe a developer across the face until he or she starts work: the difficulty of extracting good information from people. Someone may say "the bill is a set rate per unit" but a mere five minutes later, after further interrogation, is finally coaxed into saying "customers pay a flat rate for items in categories three and four". Similarly, the reliability of the words "never" or "always" should be closely scrutinized when spoken by people who aren't carefully literal. (On the other hand, xkcd has illustrated that literal responses are inappropriate in contexts like small talk...)

I'm convinced that this difficulty is partially caused by natural human language. It's too expressive. It isn't logically rigorous. By convention it supports multiple interpretations. While these attributes enable it to metaphorically branch out into new domains and handle ambiguous or incompletely-understood "analog" situations, the same attributes imply that it's too imprecise for ordering around a computing machine. Just as "I want a house with four bedrooms and two bathrooms" isn't a sufficient set of plans to build a house, "I want to track my inventory" isn't a sufficient set of software instructions to build a program (or even a basis on which to select one to buy).

Every time I perform analytical/requirements-gathering work, I'm reminded of why I doubt that natural human language will ever be practical for programming, and why I doubt that my job will become obsolete any time soon. I can envision what the programming would be like. In my head, the computer sounds like Majel Barrett.

Me: "Computer, I want to periodically download a file and compare the data it contains over time using a graph."
Computer: "Acknowledged. Download from where?"
Me: "I'll point at it with my mouse. There."
Computer: "Acknowledged. Define periodically."
Me: "Weekly."
Computer: "Acknowledged. What day of the week and what time of the day?"
Me: "Monday, 2 AM."
Computer: "Acknowledged. What if this time is missed?"
Me: "Download at the next available opportunity."
Computer: "Acknowledged."
Me: "No, wait, only download at the next available opportunity when the system load is below ___ ."
Computer: "Acknowledged. What if there is a failure to connect?"
Me: "Retry."
Computer: "Acknowledged. Retry until the connection succeeds?"
Me (getting huffy): "No! No more than three tries within an eight-hour interval."
Computer: "Acknowledged. Where is the file stored?"
Me: "Storage location ______ ."
Computer: "Acknowledged. What if the space is insufficient?"
Me: "Remove the least recent file."
Computer: "Acknowledged. What data is in the file?"
Me (now getting tired): "Here's a sample."
Computer: "Acknowledged. What is the time interval for the graph?"
Me: "The last thirty data points."
Computer: "Acknowledged. What is the color of the points? Does the graph contain gridlines? What are the graph dimensions? How will the graph be viewed?"
Me: "Oh, if only I had an analyst!"

Tuesday, June 09, 2009

your brain on CSPRNG

Not too long ago I used a cryptographically secure pseudo-random number generator API for assigning unique identifiers to users for requesting an unauthenticated but personalized semi-sensitive resource over the Internet. They already have the typical unique organizational IDs assigned to them in the internal databases, but these IDs are far from secret. As much as possible, the resource identifier/URL for a particular user had to be unpredictable and incalculable (and also had to represent sufficient bits such that a brute force attack is impractical, which is no big deal from a performance standpoint because these identifiers are only generated on request). Also, we've committed to the unauthenticated resource itself not containing any truly sensitive details - the user simply must log in to the actual web site to view those.

Anyhow, as my mind was drifting aimlessly before I got my coffee today, I wondered if the CSPRNG could be an analogy for the brain's creativity. Unlike a plain pseudo-random number generator that's seeded by a single value, the CSPRNG has many inputs of "entropy" or uncorrelated bits. The bits can come from any number of sources in the computer system itself, including the hardware, which is one reason why the CSPRNG is relatively slow at processing random numbers.

But like a CSPRNG algorithm the brain is connected to myriad "inputs" all the time, streaming from within and without the body. Meanwhile, the brain is renowned for its incredible feats of creativity. It's common for people to remark "that was a random thought" and wonder "where that came from". Given all these entropic bits and the stunning network effects of the cortex (an implementation of "feedback units" that outshines any CSPRNG), should we be surprised that the brain achieves a pseudorandom state? I hesitate to call it true randdomness in the same way as someone who believes the brain relies on quantum-mechanical effects; people who try to recite a "random sequence" tend to epically fail.

I'm not sure that this makes any real sense. I'm not suggesting that a CSPRNG is a suitable computational model for human thought. It's merely interesting to ponder. It's another reminder that just as software isn't a ghost-like nebulous presence that "inhabits" a computer - a microprocessor engineer would tell you that each software instruction is as real as the current that a pacemaker expels to regulate a heartbeat - our thoughts and minds are inseparable from the form or "hardware" of the brain, and the brain is inseparable from the nervous system, and the nervous system is inseparable from the rest of the body.

Tuesday, May 19, 2009

resist the temptation

Memo to the universe at large: resist the temptation to mention one or more of the words ["id","ego", "superego"] in discussions of the relationship among ["McCoy","Kirk","Spock"].

It's been done. To Death. Repeatedly.

Saturday, April 25, 2009

amarok 2 and "never played" playlists

Hey, Amarok, here's a tip: the 2.x series wouldn't be useless if I could order it to randomly append a never played track to the playlist...and have it work.

UPDATE
(May 6): After switching to a more recent beta, I can obtain the behavior I want. And the "playlist layout" customization allows me to view the actual play count. Now I just wish that I could search my collection based on play count or last played date...

Monday, February 09, 2009

Haskell comprehension measured through WTF/min

The top compliment I can give to Real World Haskell is that it manages to finally teach me the aspects of Haskell programming that I previously assumed to be both impenetrably complicated and useless. As I read I'm also reminded of what it was like when I first tried to comprehend Haskell code. I've concluded that the most noticeable sign of greater Haskell comprehension is a noticeable drop in my WTF/min when I'm figuring out a given code example.

WTF/min is "WTFs per minute". According to a highly-linked picture, this unit is "the only valid measurement of code quality" and it's determined through code reviews. My initial experiences of Haskell definitely exhibited high WTF/min. The following are some of the past Haskell-related thoughts I can recall having at one time or another.
  • Infinite lists like [1..]? WTF? Oh, lazy evaluation, right.
  • Functions defined more than once? WTF? Oh, each declaration pattern matches on a different set of parameters. It's like method overloading.
  • The underscore character has nothing to do with this problem domain but it's being matched against. WTF? Oh, it matches anything but discards the match.
  • Why is the scoping operator "::" strewn throughout? WTF? Oh, it's being used for types, not scopes.
  • Even a simple IO command like "putStrLn" has a type? WTF is "IO ()"? Oh, it's an expression with IO side-effects that evaluates to the value-that-is-not-a-value, ().
  • WTF? What is this ubiquitous 'a' or 't' type everywhere? Oh, it's like the type parameters of generics or templates.
  • Functions don't need "return" statements? WTF? Oh, all functions are expressions anyway.
  • WTF is going on with these functions not being passed all their parameters at once? Oh, applying a function once produces another function that only needs the rest of the parameters. That'll be helpful for reusing the function in different contexts.
  • This function definition doesn't have any parameters at all, and all it does is spit out the result from yet another function, a function that itself isn't being passed all the parameters it needs. WTF? Oh, "point-free" style.
  • Now I understand all those -> in the types. But WTF is this extra => in front? Oh, it's sorta like a list of interfaces that must be met by the included types, so the code is tied to a minimal "contract" instead of a particular set of explicit types. That's good, but how would I set those up?...
  • Ah, now this I'm sure I know. "class" and "instance" are easy. WTF?! How can that be it? Just more functions? Can't I store structured information anywhere? Oh, tuples or algebraic data types.
  • I like the look of these algebraic data types with the "|" that I know from regular expressions. Unions and enums in one swell foop. WTF? How do I instantiate it? Oh, what appear to be constituent data types are actually constructor functions.
  • After a value has been stuffed into the data type, how can my code possibly determine which data type constructor was used? WTF? Oh, just more pattern-matching.
  • WTF? Record syntax for a data type declaration results in automatic accessor functions for any value, but we use these same function names when we're creating a new record value? Oh.
  • I've acquainted with map and filter. WTF is foldr and zip and intercalate? Oh, I'll need to look over the standard list functions reference.
  • What's this "seq" sitting in the middle of the code and apparently doing jack? WTF for? Oh, to escape from laziness when needed.
  • WTF? How can a function name follow its first argument or a binary operator precede its first argument? Oh, `backticks` and (parentheses).
  • How come there's all these string escapes in the flow of code? W...T...F? Oh, lambda. Cute.
  • I've always been told that Haskell is heavily functional and pure. WTF are these do-blocks, then? Oh, monads. Wait, what?
  • Functor, Monoid, MonadPlus, WTF? Oh, more typeclasses whose definitions, like that of monads, enable highly generalized processing.
  • A way to gain the effects of several monads at once is to use "transformers"? WTF? Oh, when a transformer implements the monad functions it also reuses the monad functions of a passed monad.
  • Finally...I know that the ($) must be doing something. But what? Why use it? WTF? Oh, low-precedence function application (so one can put together the function and its arguments, then combine them).

Thursday, January 29, 2009

ravens at the intersection of logic and reality

The raven paradox presents an interesting question for anyone seeking to apply logic. It's also short and understandable. 1) The two implications "all ravens are black" and "anything that isn't black isn't a raven" are logically equivalent - either both false or both true. p -> q has the same truth table as ~q -> ~p. 2) A raven that is black is evidence for "all ravens are black". 3) Similar to 2, any object that isn't black and isn't a raven is evidence for "anything that isn't black isn't a raven". 4) Since the two propositions are logically equivalent (1), why wouldn't evidence for "anything that isn't black isn't a raven" (3) also be evidence for "all ravens are black" (2)? To summarize, how many green apples are required to convincingly support the proposition that all ravens are black? And isn't this question ridiculous?

In my opinion, the raven paradox is a matter of perspective. The Wikipedia article probably includes all of the following comments, stated differently. (If equations excite you, as usual the Wikipedia article won't disappoint in that department.)
  • Logic works best as a closed, limited system in which a truth neither "appreciates" nor "decays". I like the analogy of a microscope; it's good for ensuring that nothing is overlooked in a small fixed domain but it's unsuited for usefully observing a big unbounded area. We pick out pieces of reality and then apply logic to those pieces. The choice of axioms is vitally important. Logic's utility is tied to the universality of its rules and conclusions. The specific meanings of its "p"s and "q"s are irrelevant to its functioning.
  • Given any logical entity p, not-p (~p) is defined as the logical entity that is false whenever p is true and true whenever p is false. In the case of the raven paradox, p is "in the set of ravens" and ~p is "not in the set of ravens". q is "in the set of black" and ~q is "not in the set of black". If the "system" of these statements is all objects in the known universe, clearly ~p and ~q are huge sets in that system. But if the system of these statements is the collection of five doves and two ravens in a birdcage, isn't it more significant that five non-black birds aren't ravens than that two ravens are black? The (Bayesian) quantities matter. Some people downplay statistics because its formulas require assumptions about the source population and the randomness of samples, but it seems to me that a precise number calculated through known assumptions is still much better than an intuitive wide-ranging guess hampered by cognitive biases. When an entire population can't be measured, it's better to estimate and quantify the accompanying uncertainty probabilistically than to give up altogether.
  • Yet another factor in the perception of the raven paradox is difference in size not only between p and ~p (and q and ~q) but between the sets of p and q. There are many, many more members in "the set of black" than in "the set of ravens". Consider a more focused implication (regardless of its actual truth being rock-solid or not) like "grandfathers are older than 40". Here, the p is "grandfathers" and q is "older than 40". The overall system is people, not objects, and the sets are more stringent than colors and species. A person who is younger than 40 and not a grandfather makes one more likely to believe that all grandfathers are older than 40. For this implication, it feels more reasonable to think that evidence for ~q -> ~p is also evidence for p -> q.
  • Further probing the connection between p and q, some applications of logical implications are tighter than in the raven paradox. Laying aside sets and characteristics of objects, causes are commonly said to imply effects. Assign p to "I start a fire", and q to "the fuel is consumed (well, chemically converted)". When 1) the fuel is not consumed and 2) I haven't ignited a fire, it seems quite reasonable to accept these two facts as evidence that unconsumed fuel implies no fire-starting by me (~q -> ~p) and about as reasonable to advance these facts as evidence that my pyromaniacal actions would have led to the consumption of the fuel (p -> q). However, beware that cause and effect implication is susceptible to its own category of raven paradoxes, some of which are painfully woven into everyday life. After all, if 1) my friend isn't alive (~q) because of an auto accident and 2) I didn't tell him (~p) to avoid highway 30 on the way home, I shouldn't necessarily use these two facts to support the implication that if I had told him (p), then he would be alive (q).
  • A creative response to the raven paradox is to continue the example by pondering the unexaggerated multitude of statements that a green apple supports in addition to "all ravens are black". A green apple supports the statement that all roses are red (regardless of white roses...). A green apple supports the statement that all snow is white (again, regardless of yellow snow...). After tiring of that activity, someone could turn it around and name the statements that a black raven supports in addition to "all ravens are black". A black raven supports the statement that all watermelons are green. A black raven supports the statement that all basketballs are brown. Do this long enough and you'll realize that from logic's myopic and therefore unbiased definitions, contradictions are what matter because logic includes only true and false, is-raven and is-not-raven, is-black and is-not-black. This "binary" measure of truth results in there being no way for an implication to be progressively truthful as the evidence pile enlarges. When truth must be absolutely dependable, all-or-nothing, one contrarian member of a set trashes the implications that are blanket statements about all of the set's members.
The way I see it, the raven paradox isn't an argument that logic is useless in judging evidence or that the judgment of evidence proceeds illogically. It's an illustration of why logic is an excellent way to construct a truthful chain of reasoning out of existing truths but a flawed way to produce truths out of messy, endlessly astounding reality.

Tuesday, January 20, 2009

algorithms everywhere

I have a simple (*mumble* cheap) portable music player that allows file organization through one mechanism: unnested subdirectories listed alphabetically under the root with files listed alphabetically in each. I started pondering the best way to divvy out subdirectories and files to reduce searching time. The number of subdirectories ideally should be small, so it takes less time to select the desired subdirectory. But the number of files within each subdirectory ideally should be small as well, so it takes less time to select the desired file after selecting the subdirectory.

To cut the self-indulgent story short, I ended up reading about B-trees on Wikipedia. However, since this application has a maximum depth of one, a B-tree would be inappropriate. Yet I was sufficiently inspired to come up with my own set of insertion algorithms based on a subdirectory minimum of 5 files and a maximum of 10 files (in passing, note that these parameters meet the B-tree criterion that a full tree/subdirectory can split evenly into two acceptable trees/subdirectories).

Some might say that it's ludicrous to approach this task in this way, given that I "executed" the algorithms by hand inside a file manager in lieu of writing any code and the low-capacity music player contains less than 300 files. Thus, I lost time by analyzing the problem via a theoretical lens and formulating a general solution. I'm practical enough to acknowledge that reality.

My point is that algorithms and data structures are everywhere if one has the right perspective. And this is not strange compared to other specialties. Artists see lines and shapes and visual patterns that I wouldn't notice unless someone told me. Mathematicians see quantitative relationships (or more abstract stuff - as in abstract algebra sometimes drives me nuts). I could list numerous examples like doctors, lawyers, mechanics, architects, psychologists who all see aspects of their surroundings differently than me.

This is the part of vocational training that's hard to teach: to mold one's mind until the subject matter is a familiar mental frame or toolkit. The reason that professional software developers should study the "Computer Science-y stuff" is so that they can recognize and organize their thoughts, thereby avoiding the trap of attempting a solved problem or attacking it in a naive manner. They don't need to memorize what they can find on the Web or in a book, but they need to know enough to comprehend and adapt what they find!

Wednesday, December 03, 2008

gleaning good design advice from unit tests

I'm a beginner to the rigorous usage of xUnit unit tests, which enable easy unit test creation, execution, and repetition through coded tests that a computer can run non-interactively. I'm well-acquainted with the concept, of course, but mostly from blogs. At work, I'm alone in writing such tests. My guess is that the typical programmer categorizes up-to-date, comprehensive, effective test suites like good documentation: "nice but nonessential". The comparison is apt, since both must be maintained in parallel with the code. And the test source can function as sketchy documentation (open source projects, I'm scowling at you!).

But as so many others have observed in practice, good tests are of sufficient value to merit the developer's scarce attention. The upfront cost of writing or updating a test produces benefits repeatedly thereafter. When the code is first written, the test confirms that it meets its function. When the code changes later (it will), the test confirms that it still functions as expected and therefore won't create problems in other code that uses it according to those past expectations (this is even more important if any points of code interaction are resolved at run-time). When the code has a bug, the test that is written to check for the bug confirms that the bug is fixed and remains fixed. When the code is reorganized through the wisdom of hindsight, the test confirms that the transition hasn't accidentally abandoned established dependencies. You don't need to trust TDD or Agile wonks to enjoy these results. Just trust your tests.

The trickier aspect of unit tests, the facial mole that everybody notices and newcomers should acknowledge sooner rather than later, is the all-too-likely possibility that a typical OOP program's objects aren't the neat, independent, well-defined, composable units that facilitate unit testing. To some degree this is unavoidable, as no object is an island. Objects that are excellent individually will need to collaborate and delegate in order to perform their own useful tasks; when asked for its price including tax for a specific political domain, a sales item should need to ask yet another object for the tax rate, because tax rates are not one of an item's responsibilities. A horrendous level of difficulty of writing unit tests for an object indicates that its design or the overall design of the entire set of objects should be reexamined.

For considered as one more design constraint, greater "testability" encourages: 1) cohesion - limiting each object to a fixed and bounded purpose, 2) loose coupling - limiting brittle dependencies between objects, 3) law of Demeter - limiting the number of objects an object interacts with directly, 4) referential transparency - limiting the tendency for methods to rely on a tangled web of tedious-to-establish object states.

However, like other design constraints, testability can be detrimental when misinterpreted or carried to uncalled-for extremes. This list of downsides is more applicable to statically-typed source code that strictly enforces encapsulation (although dynamic typing is no excuse for shoddy object design, of course).
  • Exhibitionist getters and setters. The abuse of getter and setter methods is one of the evergreen blog debates. In the context of pursuing testability, inappropriate getters and setters happen because setters make object setup less of a hassle and getters make verification of a test result less of a hassle. One of the key guidelines to remember when writing unit tests is that, as much as is feasible, the test should be treating the object the same way as actual client code, so the test checks scenarios that matter. Would actual code reach down through the object's throat in order to get a drink? The true problem, as always, is that all access points that an object publicly exposes are by definition part of its (implicit) interface. Details that an object doesn't hide now have the potential to cause cascading maintenance headaches when the details change later. (Please realize that this isn't an attack against dependency injection. Getters and setters for systemic "service" collaborators, abstracted behind interfaces, is different than getters and setters for the object's internal data.)
  • Interface explosion. Since unit tests are meant for repetition, side effects are undesirable. A foolproof avoidance technique is to store objects with side effects as interface types and substitute fakes behind the interfaces during tests (this is also helpful for integration tests). The tradeoff is that interface types added purely for the sake of testing enlarge the code's size and complexity without enabling the code to do more (dynamic typing cheerleaders would say this is true of all static types). Although switching to an interface hypothetically prepares the code to work the same in the event that the class is swapped out for another, few programmers (not me) have the rare ability to exactly foretell just what common elements should be in an eternal interface. Moreover, including extraneous types, indirection, etc. conflicts with the principles of avoiding big up-front designs and stuff that's good but possibly inconsequential (gold plating). Programming to an interface instead of an implementation is good practice if an object is to accommodate frequent or dynamic replacements of its fellow objects, but this is not always the case. At this time my preference is a different strategy known as "extract and override": extract just the code that causes side effects, then override that code with a stub in a "testing version" subclass that matches the actual class as closely as feasible (similar to the Template Method OO design pattern).
  • Ugly factories. Where many interfaces are present, factory objects may be nearby, hence testability can also lead to more factories. I appreciate the decoupling that factories make possible. I recognize that factories are essential sometimes. For example, the code needs an unknown instance to fulfill an interface or an unfortunately-written object that is a chore to initialize. But I dislike factories. It feels icky and hacky to create a concrete object via a different object, not by a constructor. When objects don't have sufficient contextual information to create the right helper objects - since the more an object must know about its environment to function, the less easily it can be reused - my preference is handing those objects in (through a constructor or setter method). I'll admit there's a limit to pushing out the burden of object creation because the objects must be selected and instantiated somewhere. Centralizing object creation in an overarching Factory or Configuration singleton is a simple option but it also fails modularity. A few uses of the Abstract Factory OO design pattern could be a happy medium between factory-per-object and factory-for-all.
  • Bureaucratizing the simple. The balancing of design constraints is a cruel problem in which a compromise does what a compromise does: leaves everyone a little disappointed. After successfully dicing a jumble of concerns into testable object atoms that don't overreach, the way to accomplish a useful requirement is to...assemble an application from atoms. That's an exaggerated negative perspective, but bloggers have been mentioning or insinuating similar remarks about the standard Java APIs for a while (others might, but we understand the value of writing code with standardized stream and reader abstractions, as well as the necessity of not confusing bytes and characters in the Unicode world [Earth]). Like a UI, an API object design should have its knobs and buttons laid out understandably. Typical activities belong at front and center, with unobtrusive advanced flexibility/extensibility in the corners for heavy-duty users who know what they need. In terms of test coverage, convenience methods that straddle multiple objects are fine on the condition that the methods do nothing but delegate to fine-grained, unit-tested objects.
  • Library insulation. Libraries complicate unit tests. Closed-source libraries can't be modified in favor of greater testability, and the cost of modifying open-source libraries for testability is often significant. Insulating the code from the libraries with a lot of layers and interfaces, for faking the libraries during tests, seems like an overreaction. On the other hand, it's also wasteful writing unit tests that truth-be-told mostly exercise a mature, vetted library and just slightly exercise the new custom code that calls it. Once again, my inclination is extracting the code that needs testing from the library code that doesn't, and minimizing the untested "glue" code that actually bridges the two (like the rule that the Model and View of MVC can each be intricate but not the code in-between).
I've read that the intrusiveness of unit testing on design is meant to be embraced. But I'm sure no one is suggesting unit testing at the expense of an uncluttered, understandable, maintainable design.

Wednesday, November 12, 2008

functional languages and template engines: a StringTemplate docs quote

From the StringTemplate documentation (I started trying StringTemplate out just yesterday, when I was looking around for a currently-maintained template library for C#):

Just so you know, I've never been a big fan of functional languages and I laughed really hard when I realized (while writing the academic paper) that I had implemented a functional language. The nature of the problem simply dictated a particular solution. We are generating sentences in an output language so we should use something akin to a grammar. Output grammars are inconvenient so tool builders created template engines. Restricted template engines that enforce the universally-agreed-upon goal of strict model-view separation also look remarkably like output grammars as I have shown. So, the very nature of the language generation problem dictates the solution: a template engine that is restricted to support a mutually-recursive set of templates with side-effect-free and order-independent attribute references.

Sunday, November 02, 2008

stuff white people like, meta-midwest edition

After a recent conversation, I've discovered that white people based in the midwest U.S. enjoy Stuff White People Like. They appreciate reading wry observations about white people who differ from them. It allows them to feel better about not living like the people they see on TV, not raising awareness for popular causes, not buying the right devices, and not listening to public radio or appearing to enjoy classical music.

On the other hand, they fully relate to a fondness for shorts, dogs, apologies, living by the water (been to Michigan?), sweaters, scarves, outdoor performance clothes, etc.

Of course, midwest people in thriving urban areas or college towns just read Stuff White People Like because it is about them.