Intelligent information systems
Before we proceed further with our discussion of Intelligent IS, we need to define what we mean by the term “intelligence.”
We can regard “information” as being data + meaning. By extension, we can characterize “knowledge” as information + experience. Defining “intelligence,” however, is a rather more challenging task. In the AI world, intelligence is generally viewed as encompassing:
• Awareness of (knowledge about and the ability to interact with) the surrounding environment, and
• An ability to learn from experience and adapt accordingly.
The first of these criteria presupposes an efficient method of encoding, storing and retrieving knowledge. Several different methods exist for doing so, including if...then (or fuzzy) production rules, frames (schema), semantic networks, propositional/predicate logic, or Artificial Neural Networks. The second criteria raises the issue of “learning” from experience, incorporating such knowledge into a “Knowledge Base” (KB), and consulting this knowledge (wisdom?) when encountering new situations and circumstances, whether consciously or unconsciously (i.e., relying not just on reasoning but also on intuition). It also implies some pattern recognition ability, in order to extrapolate from known situations, to apply heuristics (rules-of-thumb), and to build upon existing knowledge.
The claim that IBM’s Deep Blue is intelligent because it defeated a human (World) Chess Champion is somewhat missing the point. The ability of the former to look a long
way ahead with sequences of next moves–up to 100 billion–is indicative of little more than a brute force approach, after all (i.e., un
intelligent).
In this article, we will restrict our focus to natural/biological systems that appear to exhibit “intelligence” (irregardless
of our specific definition of the latter term). Our premise is that by mimicking (or alternatively taking inspiration from) nature, we stand to develop systems which
naturally exhibit intelligence.” Table 1 compares and contrasts the attributes of such biologically-inspired (“soft computing”) approaches with that of logic and reasoning–the underpinning of conventional (algorithmic-based) computing.
REPRESENTATIVE INTELLIGENT
SYSTEMS
After briefly describing the underlying principles of each approach, we proceed to cite representative examples where researchers have applied “intelligent” techniques to solve real-world problems. We restrict ourselves to six biologically
Table 1. Classical vs. soft computing
Classical computing |
Soft computing |
2-valued(Boolean/crisp) logic |
Many-valued (Fuzzy) logic |
precise |
approximate |
deteministic |
Stochastic (i.e., incorporates some randomness/unpredictability) |
Exact/precise data |
ambiguous/approximate/inconsistent data |
Sequental processing |
parallel processing |
inspired soft computing methods here, these being Artificial Neural Networks, Genetic Algorithms, swarms, DNA immune- and membrane-based computing. We basically don’t consider Fuzzy Systems, reasoning systems orrule-based Expert Systems as such. However, we domention such systems in the context of combinations/hybrids of such soft computing techniques, which has become the province of the field of Computational Intelligence–CI–in recent times (Fulcher & Jain, 2008).
ARTIFICIAL NEURAL NETWORKS (ANNS)
Artificial Neural Networks (ANNs) are simplistic models of biological neural networks (brains), typically comprising dozens (but not billions) of neurons and hundreds (not tens of billions) of synapses (connections between neurons). The output (axon) of a biological neuron “fires” (produces a pulse train signal output) whenever the weighted sum of the signals from the inputs (synapses) exceeds some preset threshold. Now excitation (inhibition) of individual neurons is essentially an electrochemical process, involving different concentrations of potassium and sodium ions within and outside of the cell body; moreover, this is inherently an analog (linear) process. In the simplified neuron model commonly used in ANNs, neuron “firing” corresponds to a simple output level shift (0->1; 1->0). One characteristic of biological networks is their localized behaviour; in other words, certain areas of the brain are responsible for processing information incoming from our senses (although substantial preprocessing often takes place in the cerebral cortex prior to arriving in the brain proper), or for producing the necessary outputs (motor movement, speech, and so forth). Some ANN models reflect this localized behaviour, while others employ a more uniform, holistic architecture.
Another
characteristic of biological brains is their massively parallel
(analog) processing capability, such that despite the relative slow
processing capability of individual neurons (milliseconds), their
collective processing power far exceeds that of the fastest
supercomputers, at least for some tasks. Realization of parallel,
analog, neural network hardware in practice is by way of (sequential)
digital computer software simulation. The ANN models in common usage
are very much simplified versions of the biological networks from
which they derive their inspiration. The most popular ANN model
(Wong, Lai, & Lam, 2000) is the Multilayer Perceptron (MLP) of
Figure 2.
ANNs are not programmed in the traditional algorithmic sense, but rather learn by example, at least in the supervisedkind. Accordingly, supervised networks require numerous input-output training data pairs in order to learn the underlying “intelligence” of the system under study. Once trained, an ANN is capable of correctly recognizing input patterns not previously met during the training process; in other words, it exhibits generalization ability. Note that such a training process is an inherently data-driven, bottom-up approach (in contrast to conventional model-driven, top-down, algorithmic approaches). Furthermore, the training process can be quite time consuming; however, once trained, an ANN can respond almost instantaneously to new inputs
Figure 2. A (fully-connected) 3-layer MLP/BP
The Multilayer Perceptron (MLP) of Figure 2 is a fully connected, 3-layer, supervised, feedforward ANN, comprising input, hidden, and output layers, each of which contain n, p and m neurons,respectively. By “feedforward,” we mean that connections (weights) only exist in a forward direction, that is, from one of the n neurons in the Input Layer to one of the p neurons in the Hidden Layer (or from one of the pneurons in the Hidden Layer to one of the m neurons in the Output Layer). By contrast, no such restrictions apply in
brains. The MLP employs the so-called BackPropagation learning rule, which simply stated says that upon presentation of an input-output training exemplar pair, the actual output
produced by the network is compared with the desired output. During each successive training iteration, the weights are adjusted in proportion to this error (∆ or difference) signal: firstly, to adjust the weights connecting the Hidden Layer to
the Output Layer, then those between the Input Layer and the Hidden Layer. In this manner, the error signal “propagates” backward from the ANN output to its input, adjusting its weights in the process; hence its name (BP).
Presentation of all input-output training pairs (exem-plars)–one “epoch”–will see the weights change in many different (and incompatible/conflicting) directions. In practice many epochs will be necessary in order for the network to converge to an acceptable solution (which corresponds to the network having learned all I/O pattern associations).
It has been proven mathematically that the BP algorithm will eventually converge to an acceptable solution, although this might not be within a convenient timeframe from a user’s perspective! In practice, training of ANNs can take several hours, perhaps even overnight, even on top-of-the-range computers. Training ceases when the error (difference) signal falls below a certain level (say 0.1%), or alternatively after a certain predetermined number of training epochs.
Common practice is to divide the available training data in two, then use one half for training and the other half to test (verify) the network once trained. Now in practice such labelled I/O training data (exemplars) may not always be available, and hence some people prefer to use un
supervised neural networks. One has to exercise caution with the latter, however, because the resulting classes/clusters the network produces are often suspect. We restrict our current discus-
sion here to supervised ANNs, and indeed to only
one type (MLP/BP).
There is also the issue of how many I-O training exemplar pairs constitute a “minimum-yet-sufficient” set: too few will not lead to network convergence, whereas too many could
result in “overtraining” (akin to “overfitting” in mathematical function approximation/curve fitting).
ANNs are especially good at pattern recognition or pattern classification, irrespective of what the pattern actually represents. This means that in practice we need to be able to encode the pattern of interest (be this vision, speech, time series, or whatever) into an appropriate form. Indeed, preprocessing is often the most challenging aspect of applying ANNs to real-world problems. Typical preprocessing tasks
include the handling of missing, incomplete or noisy data, and most especially dimensionality reduction (because from what we have already seen, ANN training times are quite long; in fact, they increase exponentially as a function of the number of network weights; accordingly, any reduction in the dimensionality of the training data will have a dramatic effect on network convergence times).
Verma & Panchal (2006) used ANNs (supervised MLP/BP) in a standard pattern classification task, that of discriminating between malignant (cancerous) and benign
pap smears. Fyfe (2008) applied an unsupervised ANN–the self-organizing map–to data clustering and visualization. Likewise, Yin (2008) showed how the SOM–and variants thereof–could also be applied to vector quantization, image segmentation, density modeling, gene expression analysis,
and text mining. More sophisticated (Higher-Order, supervised) ANN models have been used for both satellite weather prediction (Zhang & Fulcher, 2004) and financial time series prediction (Fulcher, Zhang & Xu, 2006). Zeleznikow (2004) combined ANNs and rule-based reasoning in the develop-
ment of an Intelligent Legal Decision Support System. By contrast, Fu, Li, Wang, Ong, and Turner (2008) combined ANNs and multi-agents in order to predict network traffic
over media grids.
EVOLUTIONARY ALGORITHMS
As is the case with ANNs, evolutionary methods take their inspiration from Nature, in this case Darwinism and “survival-of-the-fittest.” Although there are other variants–most notably evolutionary programming and genetic programming–we will restrict our discussion here to that of Genetic Algorithms (GAs). We assume the simplified evolutionary model of Figure 3.
Prior to evolving a solution to the problem of interest, we must first be able to encode potential (candidate) solutions into (fixed-length) genetic string form. As with ANNs, in practice this preprocessing stage can often prove the most difficult part of the exercise. Commencing with random bit strings, we first select two “parent” strings from the available population on the basis of an objective (cost or fitness) function, and proceed to “mate” them. As in biological evolution, a “child” will inherit half of its genetic code (attributes, characteristics) from either parent. The aim is that over time stronger members will “evolve” more appropriate solutions to the problem at hand, while at the same time maintaining sufficient diversity among the population as a whole to ensure healthy future generations. As in nature, a certain degree of randomness (mutation) needs to be injected into this process, in order to prevent “inbreeding” and proceeding too far down evolutionary “blind alleys” (dead ends).
Figure 3. The steps in evolution
Not surprisingly, evolution of an acceptable solution can take a very long time, typically even longer than is the case with ANN training.
Mumford (2008) showed how GAs could be applied to set partitioning problems (such as graph colouring, bin packing and timetabling). Ishibuchi, Nojima and Kowajima (2008) evolved Fuzzy Classifiers using evolutionary techniques. Beale and Pryke (2006) combined GAs with interactive 3D dynamic visualization techniques in the realization of their Haiku Knowledge Discovery system. Tran, Abraham and Jain (2006) combined ANNs, EAs and Fuzzy inference methods in the development of intelligent Decision Support Systems (DSS).
SWARM INTELLIGENCE
The inspiration for this approach stems from the collective behaviour of bee/ant colonies, bird flocks, animal herds, and other social insects/animals. What we are attempting to exploit using such techniques is a system in which the whole is greater than the sum of the parts, in the sense that whereas individual members are relatively unintelligent, the collective behaviour of the colony/flock/herd exhibits “intelligent” behaviour.
Swarms differ from GAs in that there is no direct influence from one generation to the next; the focus is rather on how present-generation members affect the behaviour of others; “peer pressure” in a sense. Such influence is indirect, and often takes the form of general, “broadcast” messages, rather than “peer-to-peer” communication, as it were. A good example of this is the depositing of chemical (pheromone) trails which mark the path from a hive to a food source, say.
Apart from such reliance on indirect rather than direct communication between population members, a couple of other constraints apply to swarms, these being:
• Intragenerational learning only (i.e., no intergenerational learning), and
• Individual swarm members are assumed to have identical form, function, and status, and hence are interchangeable.
Actually there are several variants of swarms in com-mon usage, including Swarm Intelligence (SI) (Bonabeau, Dorigo, & Theraulaz, 1999), Particle Swarm Optimization (PSO) (Kennedy & Eberhardt, 1995), Ant Colony Optimi-
zation (ACO) (Dorigo, Maniezzo, & Colorni, 1996), and Autonomous Nanotechnology Swarms (ANTS), the latter having been proposed by NASA for future space exploration projects (Hinchey, Sterritt, & Rouff, 2007).
Hendtlass (2004) applied both Ant Colony and Particle Swarm Optimization (PSO) to the Travelling Salesman Problem (TSP). His ACO algorithm can be paraphrased as follows:
• Initialise pheromone levels on each path segment and randomly distribute Nants among C cities;
• Repeat
Repeat
- Each ant decides which city to move to next (provided it does not revisit a city)
Until it returns to its starting city
• Each ant calculates the length of this most recent tour and updates information about the shortest tour found to date;
• The pheromone levels on each path segment are updated (refreshed);
• All ants having completed a predefined maximum number of tours “die off” & are replaced by new ants at randomly selected cities;
• Untiltermination condition met (e.g., shortest path < predefined threshold, or maximum number of tours made).
Khosla, Kumar and Aggrawal (2006) combined PSO and the Taguchi Method to derive optimal Fuzzy Models of a Ni-Cd battery charger. Sharkey and Sharkey (2006) combined swarms and software agents in their work with collective/swarm robotics.
dna, Immune-Based, and
membrane-based computing
The three approaches considered so far, while being inspired by nature, are realized in practice by way of software simulations on (silicon-based) digital computers. With the emerging field of DNA computing, we turn our attention to carbon-based
computing, or so-called “wetware,” a “computer-in-a-test tube,” as it were. Classical algorithms are employed, rather than the data-driven approaches characteristic of ANNs, GAs and swarms.
The potential we are attempting to exploit using such an approach is the massive parallelism which results from the simultaneous reactions of large numbers of DNA molecules within a single test tube, despite the computation times of individual reactions being quite slow (a similar phenomenon to that previously encountered in relation to biological neural networks; in other words, fast overall behaviour results
from the relatively slow computations that take place within individual neurons).
Not surprisingly, one of the biggest challenges with DNA computing–just as with Quantum Computing, as it happens–is Input/Output. More specifically, how does one first encode the problem of interest into DNA strand form? Next, having done so, how does one decode the result of the chemical reaction(s) into an intelligible form (and moreover, one that relates back to the problem at hand)?
Watada (2008) demonstrated how DNA computing can be applied to scheduling problems, in particular synchronizing the movements of multiple elevators in a multistorey (high-rise) building.
Immune-based computing (IBC) utilizes “antibodies” to discriminate between “self” (good cells) and “nonself” (bad/cancerous cells) and to affect self-repair. Ishida (2008) applied IBC to the so-called “stable marriage problem,” and further showed how IBC could be extended to a general problem solver (the latter, by way of interest, was one of the traditional goals of AI).
The allied field of membrane-based computing is inspired by the so-called “reaction rules” between objects located within the compartments defined by a membrane structure. Not only are objects able to react with each other, they also on occasion pass through the membrane; also, the membrane itself can change shape, divide, dissolve or alter its permeability. Membranes are thus in a constant state of (nondeterministic) transition. Sequences of such transitions are the mechanism whereby we are able to realize parallel computations (Paun, 2002).
FUTURE TRENDS
Now, while much stands to be gained by employing the
(largely biology-inspired) “intelligent” techniques discussed above, currently many advances emanate from combinations (hybrids) of these, perhaps also incorporating statistical or Machine Learning methods more commonly encountered in Data Mining. Indeed, such hybrid approaches have become a growing concern within the discipline of Computational Intelligence, as previously mentioned (Fulcher & Jain, 2008).
A few representative examples have been mentioned. Considerable activity is currently underway in the research community in developing such hybrid systems, which no doubt will set the research agenda for the foreseeable future. Duch (2007) even suggests that Computational Intelligence might be capable of realizing a truly “intelligent” machine where AI has failed to do so during the past 50 years.
CONCLUSION
In this article, we have focussed our attention on Intelligent IS, emphasizing the incorporation of principles imitated from/inspired by nature, in attempts to realize more efficient and better performing systems. The primary areas in nature
which have served as inspiration to date include (i) Artificial Neural Networks (biological brains), (ii) Evolutionary Algorithms (Darwin’s theories of evolution and survival-of-the-fittest), (iii) swarms (i.e., ants, bees, and flocks of birds), and (iv) DNA, immune-based and membrane-based computing. Several different examples of IS which have employed such biologically-inspired approaches (and most especially hybrid techniques) were briefly described. It is the contention of the present author that there is yet more to be gained by adopting such approaches, and no doubt we will witness the development of many more Intelligent IS in the years (decades) to come.
KEY TERMS
Artificial Intelligence (AI):
The field of study devoted to building machines which exhibit “intelligence,” as commonly understood in relation to humans.
Artificial Neural Networks (ANNs):
Simplified models of the human brain (biological neural network) which are particularly adept at pattern recognition or classification.
Backpropagation (BP) Algorithm:
Used to train Multi-layer Perceptrons (supervised, feedforward neural networks); BP adjusts the weights connecting neurons in the various layers according to the error (or difference) between actual and desired outputs generated in response to presentation of input-output training pattern pairs (exemplars).
Computational Intelligence (CI):
Incorporates ANN, Fuzzy and Evolutionary approaches, and more especially hybrids of these (some authors extend this definition to include intelligent agents, stochastic reasoning and other techniques).
DNA Computing:
The implementation of classical computing algorithms by way of chemical reactions within a test tube (so called “wet computing”).
Evolutionary Algorithms (EAs):
An iterative procedure which involve the “mating” of suitable parents from a population of solutions to a problem of interest, in the hope that more suitable “offspring” (i.e., solutions) will evolve over time.
Expert (or Knowledge-Based) System (ES/KBS):
Comprise a (Graphical) User Interface, an Inference Engine and a Knowledge Base. The GUI accepts user queries (inputs), and presents the results to these queries (outputs) in a comprehensible manner, usually together with some justification (rules) or confidence level.
